Skip to content

Stop fetching every address on eth_getLogs when using RPC as a datasource #1141

Description

@ThallesP

Envio fetches every single registered address when calling eth_getLogs:
Image

While this approach makes sense during historical synchronization as it filters out unnecessary events from other addresses, it may not be optimal once we've reached the chain tip.

Issues with the current approach:

  1. It produces numerous extra requests
  2. It consumes substantial egress bandwidth, which is costly

I'm suggesting that when Envio reaches the chain tip, it should automatically switch to using eth_getLogs without any address filter, then internally discard logs from unregistered addresses.

I've created a proof of concept at notuslabs@643a719. I haven't tested it yet as I need to build and release it as a package for testing, but it could serve as inspiration (though it was generated with AI, so please take it with a grain of salt).

Additionally, it would be beneficial if we could configure this behavior via config.yml under the rpc section, allowing developers to set unfiltered addresses even during historical sync. The default setting should remain as live only.

Activity

  1. DZakh commented on Apr 23, 2026

    @DZakh
    Member

    You can use https://docs.envio.dev/docs/HyperIndex/wildcard-indexing to stop sending the requests during backfill

  2. DZakh commented on Apr 23, 2026

    @DZakh
    Member

    How many addresses do you index?

  3. ThallesP commented on Apr 23, 2026

    @ThallesP
    Author

    You can use https://docs.envio.dev/docs/HyperIndex/wildcard-indexing to stop sending the requests during backfill

    makes sense, then ignore my suggestion for backfill (I do not need it but thought others would want it).

  4. ThallesP commented on Apr 23, 2026

    @ThallesP
    Author

    How many addresses do you index?

    I register every Uniswap V3 pool, totaling approximately 37,000.

    Image

    So in theory, I would perform 7.4 times more requests than simply using eth_getLogs without an address filter, which would also provide a significant speed improvement.

  5. DZakh commented on Apr 24, 2026

    @DZakh
    Member

    Yeah, it makes sense. With HyperSync we have query caching layer in place, so it's not such a big problem, but definitely a huge egress cost with RPC :) We even almost got drained by AWS when in one of the old versions, a client indexed 2M of addresses 😅

  6. DZakh commented on Apr 24, 2026

    @DZakh
    Member

    Even though, it's solved with HyperSync, we still want to have it improved with RPC. Additionally having a single request is much smaller latency than 2+

  7. ThallesP commented on Apr 24, 2026

    @ThallesP
    Author

    Yep, hopefully that will be resolved soon. In the meantime, I implemented a better workaround.

    I built an RPC Proxy that cleanses the address parameter from eth_getLogs, forwards the request upstream, and then filters the response by address internally. It also caches the eth_getLogs response for approximately one minute to handle the additional seven requests.

    So far, everything seems to be working perfectly. I'll maintain this solution until HyperIndex resolves the issue.

    About HyperSync, that makes sense. We were previously using it, and after an update, the egress dropped massively.

  8. ThallesP commented on Apr 24, 2026

    @ThallesP
    Author

    In case someone needs it in the meanwhile, here's a gist: https://gist.github.com/ThallesP/10f558c53736f2b5de2f6a6676883923.

  9. self-assigned this
    on May 16, 2026
  10. sakulstra commented on Sep 30, 2026

    @sakulstra
    Contributor

    Very much looking forward for a solution to this.

    Just for the record, we solved it the other way around on our fork 😅
    Instead of removing the address filter for live indexing we just merged the buckets, so we potentially overfetch (e.g. receiving Transfers from tokens we only index for Approvals), but i guess we over fetch less than we would by omitting the address filter?
    But we also only index ~100 addresses, not 37,000 :D

    Not quite sure what is the better solution.

    edit: added our fork edits here #1673

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

No labels
No labels

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions