Skip to content

Add dedup-aware HNSW format that builds the graph over distinct vectors and expands matches to all referencing documents - #16731

Draft
Pranshu-S wants to merge 7 commits into
apache:mainfrom
Pranshu-S:dedup-hnsw
Draft

Pranshu-S wants to merge 7 commits into
apache:mainfrom
Pranshu-S:dedup-hnsw

Conversation

@Pranshu-S

Copy link
Copy Markdown
Contributor

The dedup vector formats in de-duplicate flat vector storage but delegate graph construction to Lucene99HnswVectorsWriter, which builds one HNSW node per document. When documents share vectors, that redundantly indexes the same point once per document.

This PR adds a de-duplication-aware HNSW writer/reader that build and search the graph over distinct vectors (group ordinals), and expand a matched distinct vector back to all referencing documents at search time. Storage savings from dedup are now reflected in the graph as well.

…rs and expands matches to all referencing documents

Signed-off-by: Pranshu-S <pranshushukla06@gmail.com>
Signed-off-by: Pranshu-S <pranshushukla06@gmail.com>
Signed-off-by: Pranshu-S <pranshushukla06@gmail.com>
@Pranshu-S

Copy link
Copy Markdown
Contributor Author

Interestingly, from my local runs it seems that there is a tip-off point where this new Reader/Writer pays off: When the number of unqiue vectors within a field are less - that is same vector is being repeated in the same field. Logically it can happen - such as when different records reference the same underlying item embedding.

Here is what my results came out as:

Total Unique = 100

==========================================================================
COMPARISON  (candidate vs baseline;  delta/pct = candidate - baseline)
            float32 vectors, filtered KNN search
metrics are mean±stddev over iterations (baseline n=5, candidate n=5)
==========================================================================
metric             baseline(n=5)      candidate(n=5)     delta      pct
-----------------  -----------------  -----------------  ---------  ------
recall             0.809±0.008        0.982±0.000        +0.173     +21.4%
latency(ms)        0.158±0.013        0.212±0.010        +0.055     +34.8%
netCPU             0.152±0.013        0.206±0.010        +0.055     +36.0%
avgCpuCount        0.962±0.003        0.971±0.003        +0.009     +0.9%
nDoc               10000              10000
searchType         KNN                KNN
topK               10                 10
fanout             100                100
resultSimilarity   N/A                N/A
decay              N/A                N/A
resultCount        10.000             10.000
maxConn            32                 32
beamWidth          200                200
quantized          7 bits             7 bits
indexEncoding      float32            float32
visited            230.200±7.155      100.000±0.000      -130.200   -56.6%
index(s)           1.568±0.038        0.890±0.019        -0.678     -43.2%
index_docs/s       6379.994±161.224   11232.210±231.098  +4852.216  +76.1%
merge(s)           0.000±0.000        0.000±0.000        +0.000
force_merge(s)     2.180±0.055        1.060±0.045        -1.120     -51.4%
num_segments       1.000±0.000        1.000±0.000        +0.000     +0.0%
index_size(MB)     0.658±0.004        0.620±0.000        -0.038     -5.8%
filterStrategy     index-time-filter  index-time-filter
filterSelectivity  0.50               0.50
overSample         1.000              1.000
vec_disk(MB)       48.981±0.000       48.981±0.000       +0.000     +0.0%
vec_RAM(MB)        9.918±0.000        9.918±0.000        +0.000     +0.0%
bp-reorder         false              false
indexType          HNSW               HNSW
rerank             no                 no

Total Unique = 1000

==========================================================================
COMPARISON  (candidate vs baseline;  delta/pct = candidate - baseline)
            float32 vectors, filtered KNN search
metrics are mean±stddev over iterations (baseline n=5, candidate n=5)
==========================================================================
metric             baseline(n=5)      candidate(n=5)     delta      pct
-----------------  -----------------  -----------------  ---------  ------
recall             0.964±0.002        0.965±0.002        +0.001     +0.1%
latency(ms)        0.225±0.004        0.224±0.028        -0.001     -0.3%
netCPU             0.217±0.003        0.217±0.028        +0.000     +0.2%
avgCpuCount        0.966±0.006        0.970±0.007        +0.004     +0.4%
nDoc               10000              10000
searchType         KNN                KNN
topK               10                 10
fanout             100                100
resultSimilarity   N/A                N/A
decay              N/A                N/A
resultCount        10.000             10.000
maxConn            32                 32
beamWidth          200                200
quantized          7 bits             7 bits
indexEncoding      float32            float32
visited            496.000±8.746      422.400±2.302      -73.600    -14.8%
index(s)           2.406±0.065        1.920±0.166        -0.486     -20.2%
index_docs/s       4159.980±109.441   5241.906±499.526   +1081.926  +26.0%
merge(s)           0.000±0.000        0.000±0.000        +0.000
force_merge(s)     3.450±0.071        1.666±0.047        -1.784     -51.7%
num_segments       1.000±0.000        1.000±0.000        +0.000     +0.0%
index_size(MB)     5.140±0.000        5.100±0.000        -0.040     -0.8%
filterStrategy     index-time-filter  index-time-filter
filterSelectivity  0.50               0.50
overSample         1.000              1.000
vec_disk(MB)       48.981±0.000       48.981±0.000       +0.000     +0.0%
vec_RAM(MB)        9.918±0.000        9.918±0.000        +0.000     +0.0%
bp-reorder         false              false
indexType          HNSW               HNSW
rerank             no                 no

@Pranshu-S

Copy link
Copy Markdown
Contributor Author

Few optimisations that I can think off if I were to pull this to a state it could be used:

  1. Decide dynamically when to use this new Writer/Reader
  2. Performing post-filter on results on the Hnsw results since if we go with (1) - high change that we are doing something similar to IVF (Grouping related (same in our case) vectors together and reducing the search set: Wow!) - so better do it after we have reduced the working set?

Signed-off-by: Pranshu-S <pranshushukla06@gmail.com>
@Pranshu-S

Copy link
Copy Markdown
Contributor Author

Ran some benchmark suits, I was able to find meaningful results only when there close to 70-80% deduplication WITHIN A FIELD which is very unlikely in any dataset. Even in those cases, it would make sense to just run exact KNN on the dedup than build a Hnsw Graph.

What might be good next points to explore would be IVF and tiered search similar to Vamana?

Signed-off-by: Pranshu-S <pranshushukla06@gmail.com>
… worth

Signed-off-by: Pranshu-S <pranshushukla06@gmail.com>
@kaivalnp kaivalnp linked an issue Oct 5, 2026 that may be closed by this pull request
1 of 6 tasks
Signed-off-by: Pranshu-S <pranshushukla06@gmail.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Improvements to the sandboxed de-duplicating vector format

1 participant