Repository navigation
Conversation
…rs and expands matches to all referencing documents Signed-off-by: Pranshu-S <pranshushukla06@gmail.com>
Signed-off-by: Pranshu-S <pranshushukla06@gmail.com>
|
Interestingly, from my local runs it seems that there is a tip-off point where this new Reader/Writer pays off: When the number of unqiue vectors within a field are less - that is same vector is being repeated in the same field. Logically it can happen - such as when different records reference the same underlying item embedding. Here is what my results came out as: Total Unique = 100 Total Unique = 1000 |
|
Few optimisations that I can think off if I were to pull this to a state it could be used:
|
Signed-off-by: Pranshu-S <pranshushukla06@gmail.com>
|
Ran some benchmark suits, I was able to find meaningful results only when there close to 70-80% deduplication WITHIN A FIELD which is very unlikely in any dataset. Even in those cases, it would make sense to just run exact KNN on the dedup than build a Hnsw Graph. What might be good next points to explore would be IVF and tiered search similar to Vamana? |
Signed-off-by: Pranshu-S <pranshushukla06@gmail.com>
… worth Signed-off-by: Pranshu-S <pranshushukla06@gmail.com>
Signed-off-by: Pranshu-S <pranshushukla06@gmail.com>
The dedup vector formats in de-duplicate flat vector storage but delegate graph construction to Lucene99HnswVectorsWriter, which builds one HNSW node per document. When documents share vectors, that redundantly indexes the same point once per document.
This PR adds a de-duplication-aware HNSW writer/reader that build and search the graph over distinct vectors (group ordinals), and expand a matched distinct vector back to all referencing documents at search time. Storage savings from dedup are now reflected in the graph as well.