Repository navigation
Stop fetching every address on eth_getLogs when using RPC as a datasource #1141
Description
Activity
You can use https://docs.envio.dev/docs/HyperIndex/wildcard-indexing to stop sending the requests during backfill
How many addresses do you index?
You can use https://docs.envio.dev/docs/HyperIndex/wildcard-indexing to stop sending the requests during backfill
makes sense, then ignore my suggestion for backfill (I do not need it but thought others would want it).
Yeah, it makes sense. With HyperSync we have query caching layer in place, so it's not such a big problem, but definitely a huge egress cost with RPC :) We even almost got drained by AWS when in one of the old versions, a client indexed 2M of addresses 😅
Even though, it's solved with HyperSync, we still want to have it improved with RPC. Additionally having a single request is much smaller latency than 2+
Yep, hopefully that will be resolved soon. In the meantime, I implemented a better workaround.
I built an RPC Proxy that cleanses the
addressparameter frometh_getLogs, forwards the request upstream, and then filters the response by address internally. It also caches the eth_getLogs response for approximately one minute to handle the additional seven requests.So far, everything seems to be working perfectly. I'll maintain this solution until HyperIndex resolves the issue.
About HyperSync, that makes sense. We were previously using it, and after an update, the egress dropped massively.
In case someone needs it in the meanwhile, here's a gist: https://gist.github.com/ThallesP/10f558c53736f2b5de2f6a6676883923.
Reacted by Dmitry ZakharovVery much looking forward for a solution to this.
Just for the record, we solved it the other way around on our fork 😅
Instead of removing the address filter for live indexing we just merged the buckets, so we potentially overfetch (e.g. receiving Transfers from tokens we only index for Approvals), but i guess we over fetch less than we would by omitting the address filter?
But we also only index ~100 addresses, not37,000:DNot quite sure what is the better solution.
edit: added our fork edits here #1673

Envio fetches every single registered address when calling eth_getLogs:

While this approach makes sense during historical synchronization as it filters out unnecessary events from other addresses, it may not be optimal once we've reached the chain tip.
Issues with the current approach:
I'm suggesting that when Envio reaches the chain tip, it should automatically switch to using eth_getLogs without any address filter, then internally discard logs from unregistered addresses.
I've created a proof of concept at notuslabs@643a719. I haven't tested it yet as I need to build and release it as a package for testing, but it could serve as inspiration (though it was generated with AI, so please take it with a grain of salt).
Additionally, it would be beneficial if we could configure this behavior via
config.ymlunder therpcsection, allowing developers to set unfiltered addresses even during historical sync. The default setting should remain asliveonly.