Repository navigation
Conversation
SegmentIVF is a sandbox KnnVectorsFormat that builds a per-segment IVF index. Each segment clusters its vectors into nlist cells with spherical k-means, spilling vectors near cell boundaries into a second cell, and stores two tiers per vector: a 2-bit Nitrox2 coarse code scanned with SIMD Hamming kernels, and an INT8 or FP32 fine record used to rerank a shortlist of coarse candidates. Indexing: every non-empty segment trains exactly nlist cells, so any segment of the same configuration can seed another. Flushes and merges warm-start from the largest segment this process has written (never from segments merely found on disk), averaging same-lineage centroids by their live populations, and merges carry existing cell assignments so they converge quickly without cold re-clustering. Search: SegmentIVFKnnQuery merges every segment's deduplicated coarse shortlist into one index-wide shortlist of 7 candidates per requested neighbor (at least 100) and fine-reranks only those, reading records in place, so fine reads per query do not grow with the segment count. Filtered searches take the same path. Memory: coarse codes and the slot-to-document section are copied off-heap on a background thread, within a budget derived from the cgroup memory limit and the heap, so the page cache can never evict them. Fine records are read mapped while the fine tiers fit in memory; when they do not, or the cgroup reports memory pressure, reranks read them with one batched buffered io_uring submission so page-cache misses run in parallel.
gsmiller
left a comment
There was a problem hiding this comment.
Thanks for this new iteration! It looks super interesting. I'm going to attempt to go "layer by layer" starting at the top and leave feedback as I go instead of all at once. I hope that's OK. I've looked through the new query implementation as a starting point to understand the top-level logic and left a few questions/thoughts. Thanks again!
| /** Intersects the filter with the field and pre-creates its weight for segment search. */ | ||
| @Override | ||
| public Query rewrite(IndexSearcher searcher) throws IOException { | ||
| if (reranked != null || filterWeight != null) return super.rewrite(searcher); |
There was a problem hiding this comment.
OK, if I'm reading this right (I may not be), we're essentially relying on the super rewrite logic to manage a glorified merge sort of the final/global ranked results. Is that right? I'm talking specifically about the case where reranked gets populated by the logic in this query. There's probably a lighter-weight way to handle this. Thoughts? Maybe there's some other reason I'm overlooking to delegate to the super#rewrite logic here?
| : sivf.candidates(field, target, k, strategy, accept); | ||
| }); | ||
| } else { | ||
| return null; // Some segment is not SegmentIVF. |
There was a problem hiding this comment.
When would this happen? I'm sure there's an obvious case I'm overlooking, but it seems like the field specified for vector search would have to have different types/formats across segments for this to occur. Is that right?
| ranked[at++] = (long) distances[i] << 40 | (long) l << 20 | i; | ||
| } | ||
| } | ||
| Arrays.sort(ranked); |
There was a problem hiding this comment.
We don't need a full sort here do we? Would a partial sort be good enough? If so, we could use a quick select partial-sort algo here.
| * Returns the segment's live documents, intersected with the filter as a bit set when there is | ||
| * one, or null when nothing in the segment is accepted. | ||
| */ | ||
| private static AcceptDocs accepted(LeafReaderContext context, Weight filter) throws IOException { |
There was a problem hiding this comment.
There are a couple places in here where you go re-fetch live docs from the segment. I think you already have live docs in all the call sites. Would it make sense to pass it along?
gsmiller
left a comment
There was a problem hiding this comment.
Spent a little more time on this, so leaving a few more comments. I was planning to get through the Reader and Centroids implementations before leaving more feedback, but one thing that jumps out to me is a combination of, (1) very compact code (lots of one-liners, embedded ternary statements, etc.), and (2) a lack of comments making a lot of the core logic pretty cumbersome to work through—for me at least. I have to ask... after the previous PR was 13k+ lines of code, did you prompt your agent to target fewer lines? Is this highly compact code the result? :)
I don't want to be too picky about this, but let's aim to keep this code base readable by other contributors. Even if it inflates the line count :)
| return out; | ||
| } | ||
|
|
||
| int search(byte[] qCode, int ef, int[] out, int[] outDist) { |
There was a problem hiding this comment.
I found this method quite difficult to read. Not sure if more descriptive variable names would help. Or a descriptive comment of the algorithm. Or both? I don't know if it's just me or if others will also snuggle through this, but you might consider ways to make these core algorithms a little more readable by humans.
| if (visited.length <= nlist) VISITED.set(visited = new int[nlist + 1]); | ||
| int gen = ++visited[0], nOut = 0, frontierN = 0, bestN = 0; | ||
| int cap = Math.min(Math.max(ef, 1), nlist); | ||
| long[] frontier = new long[Math.min(nlist, Math.max(64, cap * 4))], best = new long[cap]; |
There was a problem hiding this comment.
I think you could use LongHeap off-the-shelf for frontier and best?
| int deg = | ||
| adj != null ? adj.length : (nodes[degOff] & 0xFF) | (nodes[degOff + 1] & 0xFF) << 8; | ||
| fan = 0; | ||
| for (int i = 0, off = degOff + 2; i < deg; i++, off += ORD_BYTES) { |
There was a problem hiding this comment.
I wonder if this loop would be more readable if you had separate logic for the case when building != null and when building == null...
| got = graph.search(qCode, ef, candidates, coarse); | ||
| int cap = Math.max(VERIFY_MIN, probe * VERIFY_MULTIPLIER); | ||
| if (cap < got) { | ||
| int[] counts = new int[bins]; |
There was a problem hiding this comment.
As a general bit of feedback, I think some comments would go a long way, especially in these more core algorithm bits of your work. For example, it took me a minute to realize this is a counting sort. A comment describing the fact that you use a counting sort when you get back more scored centroids than the cap would go a long way for readers.
| got = n; | ||
| } | ||
| } | ||
| long[] ranked = rank(candidates, got); |
There was a problem hiding this comment.
Is this another place where we don't actually need a full rank, but rather a partially ordered top-k list? If so, maybe another candidate for a quick-select style algo?
Fine records are read only to rerank a shortlist deduplicated by document, so spill copies now store only a coarse code. Records are written in ordinal order and reranks address them by ordinal. Format version 1; version 0 segments, with a record per slot, still read and merge.
CoarseTier.BBQ codes each slot with Lucene's OptimizedScalarQuantizer: one bit per dimension of the vector's offset from its cell's centroid, with interval corrections, scored by AND and popcount against a 4-bit query's bit planes. The query is quantized once, uncentered, so each probed cell costs one exact dot product with its centroid rather than a quantization. The centroid graph and clustering's routing use the same tier, anchored on the mean centroid, so a BBQ field runs no Nitrox2 code. NITROX2 stays the default. Format version 2 records the coarse tier; older versions imply Nitrox2.
SegmentIVF: a fast(er) Lucene-native inverted-file vector codec
SegmentIVF is a Lucene KNN codec designed to make vector search behave like Lucene full-text search architecture. It is a further development of IVFaster #16567 with some cool new features. Before going into the implementation, it is important to describe the architectural philosophy behind the codec and the restrictions that make traditional vector indexes difficult to use inside Lucene.
Lucene's segment architecture gives us a simple concurrency model, immutable files after flush, natural near-real-time publication, and straightforward distribution across nodes. Vector search methods are often at odds with this model. Graph structures such as HNSW must be rebuilt during segment merges, while traditional k-means IVF implementations can impose a large independent training cost on every new segment. Quantizers trained from segment-local data also require full-precision vectors to survive so that every merge can train a new quantizer and encode the vectors again.
SegmentIVF starts from the opposite constraint: a vector's code must remain valid regardless of which segment holds it. Flushes and merges should reuse earlier clustering work, merging should copy existing encoded rows instead of returning to float32, and search work should scale with the documents searched rather than multiplying expensive reranking by the number of segments.
Benchmark
This benchmark uses 1M Cohere Embed multilingual-v3 Wikipedia-en vectors, dimension 1024, unit-normalized with dot-product similarity, and 1,000 held-out queries. Recall is recall@100 against exact nearest neighbors. The index was force-merged to one segment. The machine is a 64-core Arm server
m7g.16xlargewith 256-bit SVE running Corretto 25.0.4. Search latency is warm, single-threaded, and the minimum of three independent JVM runs through luceneutil.SegmentIVF uses its defaults:
nlist=1000(sqrt(1M)),nprobe=32,spillBits=1,spillMargin=1.05, a0.75adaptive probe margin, Nitrox2 coarse codes, INT8 fine codes, and a global rerank ofmax(100, 7 * k)candidates. Recall is dialed withnprobealone.The Lucene HNSW measurements below are the numbers from the previous benchmark on this dataset and machine. They use Lucene's 7-bit scalar-quantized HNSW path with a Lucene99 graph,
M=16, andbeamWidth = 100.nprobe=10)fanout=0)nprobe=14)fanout=25)nprobe=18)fanout=50)nprobe=28)fanout=100)nprobe=48)fanout=200)nprobe=128)fanout=300)Indexing
SegmentIVF builds the initial index 8.6x faster, force-merges 8.3x faster, and produces an index 61.6% smaller. The median end-to-end time to the final one-segment index was 31.575 seconds, compared with 264.1 seconds for the previous HNSW measurement, an 8.4x improvement.
The size comparison needs one qualification because the indexes retain different data. Lucene's scalar-quantized HNSW format wraps a raw float32 delegate so that vectors can be requantized during merge. Most of its 4965 MB is therefore full-precision vector data. SegmentIVF keeps no float32 copy of each document vector. Its fine representation is one signed byte per dimension plus a per-vector scale, and its coarse representation is a 2-bit Nitrox2 code. Spill copies those compressed records only for vectors close enough to a cell boundary.
The indexing speed comes from the same architectural choice. SegmentIVF has no document graph to rebuild. Build rows are staged sequentially on disk, encoding and clustering use all available build workers, compatible merges copy existing fine and coarse codes, and HotStart carries useful centroid and assignment state into the next clustering run.
900M-vector scaling (on a single
m7g.16xlarge)This run indexed the first 900M vectors from a 901,180,094-vector FineWeb corpus. The vectors are 768-dimensional FP16 input with cosine similarity. The codec used
nlist=3873, which is approximatelysqrt(15M)for the target maximum segment population,spillBits=1,spillMargin=1.05, and the INT8 fine tier.SegmentIVF is designed to scale. The complete build took 10 hours 18 minutes 47 seconds, including waiting for already-running background merges at the end. It was not force-merged: the result is a normal Lucene index with 104 immutable segments averaging 8.65M vectors each. Background merges ran concurrently with ingestion; the separate 156-second figure is only the final drain after the last document was committed.
Global, segment-independent codes
Quantization is critical ANN infrastructure: it reduces storage and memory traffic while retaining production-quality recall. In Lucene, quantizer design also determines merge behavior.
SegmentIVF's transform and encodings do not depend on segment statistics. Every vector is L2-normalized and rotated by
where
Sapplies deterministic random plus/minus-one signs,Pis a Fisher-Yates permutation, andFis a normalized fast Walsh-Hadamard transform applied block-diagonally over the set bits of the dimension. The seed is derived from the dimension, so every segment for a compatible field uses the same rotation.The rotation spreads variance across dimensions, preserves similarity, and puts unit vectors on a shared grid whose coordinate scale is determined by the dimension rather than by a training set. The INT8 fine code then uses a per-vector scale, and the Nitrox2 coarse code uses fixed thresholds derived from
1 / sqrt(dim). Moving a vector between segments does not change either code.This gives merges ordinary Lucene semantics: copy the encoded row, update its document and cell metadata, and place it in the new posting lists. There is no segment-level requantization pass and no need to retain the original float32 vector.
Nitrox2: A quantizer built for vector search
The coarse tier has a narrower job than the fine tier. It does not need to produce the final ranking; it only needs to keep true neighbors in a bounded shortlist. That permits a much cheaper representation.
Nitrox2 stores two threshold planes for every rotated coordinate. At dimension
d, the thresholds are-0.5 / sqrt(d)and+0.5 / sqrt(d), producing three ordered levels. Because the planes are cumulative, Hamming distance between two codes is the sum of the per-coordinate level differences:The query and document use the same code. There is no document correction term and no asymmetric expression that a scan path can accidentally omit. At 1024 dimensions the complete coarse code is 256 bytes.
The scan kernel is a fused XOR and popcount over contiguous rows. The Panama Vector API implementation processes native memory directly, scores several rows together so query chunks are reused, and falls back to scalar Java when the vector module is unavailable. The result is a bandwidth-oriented posting-list scan with no native library dependency.
The fine tier defaults to INT8. Each rotated vector uses one signed byte per dimension and one scale. The query is encoded once, and reranking uses SIMD byte dot products before applying Lucene's similarity-score transformation. FP32 is also supported when a caller prefers exact fine records over the smaller default.
Indexing: Segment-aware clustering and HotStart
How cold segment build works:
nlistcentroids from a deterministic reservoir sample.The expensive mistake in a segment architecture is treating every flush and merge as an unrelated clustering problem.
Clustering.HotStartretains recent clustering state and reuses it in two places.For a flush, SegmentIVF finds the largest live compatible segment written by this process. Its centroids seed the new segment, and compatible segments from the same lineage contribute membership-weighted centroid state. This makes a stream of near-real-time flushes progressively warmer instead of repeatedly paying the full cold-start cost.
For a merge, SegmentIVF selects the largest compatible source segment as the donor. Its centroid layout seeds the merged segment. Sources from the same HotStart lineage contribute their live primary-cell populations to a weighted seed, and their rows carry primary and runner-up assignments into the first iteration. Rows whose lineage or assignment is unknown are safely routed from scratch.
Exact movement pruning
Rerouting every vector after every centroid update is mostly wasted work because only vectors near a Voronoi boundary can change assignment. SegmentIVF tracks the best and runner-up centroids for each vector and accumulates the movement of those centroids.
A vector can only keep its current assignment when
where
d1andd2are the distances to its current best and runner-up cells. When the runner-up is not known, the implementation uses a conservative bound derived from the maximum centroid movement. This is a sufficient condition, so skipped vectors cannot have changed their nearest cell.The bound follows the k-means pruning line from Elkan, Hamerly, and Yinyang, applied here to donor-warm-started segment clustering. As centroids settle, almost all vectors stop participating in later routing passes.
Indexing: Spill
Traditional IVF probes the cells whose centroids are closest to the query. In high dimensions, pairwise distances concentrate and cell boundaries become unreliable: a true neighbor can sit just across a Voronoi boundary in a cell that is not among the query's first probes.
Spill attacks that failure directly. A vector close to a boundary is written to its primary cell and
spillBitssecondary cells, so a query can find it from either side. The defaultspillBits=1permits one secondary copy, andspillMargin=1.05limits duplication to vectors whose two closest cells are competitive within the margin.The secondary cell is selected with SOAR, Spilling with Orthogonality-Amplified Residuals. Choosing the second-nearest centroid often duplicates the same failure direction as the primary. SOAR instead penalizes a candidate whose residual is parallel to the primary residual:
The default
lambda=1favors a complementary cell whose residual covers a different direction. That makes each spill copy more useful and lets the codec keep the spill count and index size low.Spill means one document may appear in several probed posting lists. A fixed slot shortlist can therefore collapse to far fewer distinct documents. SegmentIVF admits up to
shortlist * (1 + spillBits)slots, orders them by coarse distance, deduplicates by document, and keeps the requested number of distinct candidates. The writer's spill cap consequently gives the reader a hard bound on how much extra admission is needed.Search: Query-Time Cell selection
Choosing the cells to scan is itself a nearest-neighbor problem. A brute-force scan across all
nlistcentroids is reasonable while indexing because many documents can share a tiled, cache-resident centroid scan. A single search query cannot amortize that work, andnlistmust grow with the corpus to keep posting lists small.SegmentIVF therefore builds a compact graph over centroid codes. Each node stores the centroid's 2-bit Nitrox2 code next to up to 16 neighbor ordinals in a cache-aligned record inspired by USearch. At dimension 1024 the code is 256 bytes instead of the centroid's 4096-byte float32 vector.
The query starts with a maximum
nprobe, then applies an adaptive margin to the exactly ranked centroid candidates. A cell is discarded once its centroid distance exceeds the configured fraction of the best cell's distance. Easy queries with one dominant cell stop early; ambiguous queries retain more cells. The default margin is0.75, and a margin of1.0disables the cutoff.Search: Index-wide rerank
Lucene searches segments independently, which creates a cost for a two-tier vector index. If every segment fine-reranks its own fixed shortlist, fine scoring and random fine-record reads grow linearly with the number of segments. That is especially undesirable for near-real-time indexes, where segment count is expected to fluctuate.
How
SegmentIVFKnnQueryworks:max(100, 7 * k)candidates.For a top-100 query, the entire index fine-reranks 700 records whether it has one segment or many. A plain
KnnFloatVectorQueryremains supported through the codec interface, but it performs the normal per-segment rerank.SegmentIVFKnnQueryis the path that provides the index-wide bound.Note: 7 * k was chosen empirically using Cohere Wikipedia v3-1024 and may not be optimal for your dataset.
Search: Automatic mapped reads and io_uring
The coarse and fine tiers have different memory access patterns, so SegmentIVF manages them differently.
Coarse posting runs are large, contiguous, and scanned repeatedly. On reader open, SegmentIVF queues a background off-heap copy of the coarse codes and slot-to-document mapping when the memory budget allows it. At most three quarters of that remaining space is reserved for pinned coarse data. Searches never wait for the copy; they use the mapped view until the pinned copy is published.
Fine records are accessed sparsely after the global shortlist is known. Mapped reads are fastest while the working set fits in memory, but a page miss turns each scattered record into a synchronous fault. SegmentIVF tracks the total fine-tier bytes across open readers and reevaluates its read path at most once per second:
some avg10pressure and back below 0.1%, to avoid oscillation.The io_uring path submits the rerank's scattered fine-record reads as one batch. Cached records still come from the page cache, while page misses can run concurrently instead of serializing as faults on the search thread. Rings and buffers are per search thread; each segment reader owns only its file descriptor. If io_uring is unavailable or a batch fails, the reader falls back to the mapped input.
Filtered search
Filters are applied before vector scoring.
SegmentIVFKnnQuerymaterializes a dense filter as a per-segment bit set and passes it into the codec.The reader chooses between two paths based on filter cost:
Segments whose slots fit within the coarse admission pool scan all slots instead of spending probes on mostly empty cells. Sufficiently selective filters visit their accepted documents directly. Filtered candidates enter the same coarse ordering and global rerank as unfiltered candidates, preserving the index-wide fine-read bound.
Final Note
Hi everyone, this PR is a culmination of a lot of trial and error and a collaborative effort between many people. Thank you for reading this very long PR description. I would appreciate any feedback on the method or general thoughts of this IVF implementation in Lucene.