Skip to content

feat(examples): batch RaCo-ALIKED extraction and LightGlue matching for out-of-process consumers - #24

Merged
edgarriba merged 8 commits into
mainfrom
feat/batch-extract-and-match
Aug 11, 2026
Merged

edgarriba merged 8 commits into
mainfrom
feat/batch-extract-and-match

Conversation

@edgarriba

Copy link
Copy Markdown
Member

Two file-based bridges so a crate that cannot link against this one can still use the pipeline. The
motivating consumer pins a different kornia, so the two sides meet on disk.

aliked_batch <engine> <img_dir> <out_dir> — one .vrtk per frame: magic, n, dim, then
n * (x, y, desc[dim]) f32.

lightglue_batch <extractor> <matcher> <img_dir> <pairs.txt> <out.vrtm> — matches a pair LIST,
one file for the whole run.

Three things worth reviewing:

  • Keypoints come back in SOURCE-image pixels. Both tools undo the extractor's internal 32 px fit
    and their own downscale. A coordinate left in model space is wrong by the resize ratio in a way
    the consumer cannot detect.
  • lightglue_batch keeps an LRU of extraction results rather than extracting per pair. Each
    frame appears in ~17 pairs of a windowed schedule, so per-pair extraction runs the backbone ~34x
    per frame, while caching every frame holds n*K*128*4 bytes — 700 MB at 459 frames on an 8 GB
    Jetson.
  • Both downscale to fit the shipped engines' shape profile (256x256 .. 640x640). Without it the
    extractor rejects any frame with a side over 640 — a 480x853 phone preview fails on the long axis,
    which is how the first run produced zero files.

Neither sorts nor truncates keypoints: a detector's output order is not a quality ranking, and
truncating raster-ordered output crops the image rather than keeping the strongest points.

Measured with these on a 459-keyframe indoor map: LightGlue matched at ~75% geometric survival
(RANSAC on a fundamental matrix) against SIFT's ~13% on the same pairs.

🤖 Generated with Claude Code

edgarriba and others added 8 commits August 11, 2026 05:11
…or out-of-process consumers

Two file-based bridges so a crate that cannot link against this one can still use the pipeline.
flux-map pins a different kornia, so the two sides meet on disk.

`aliked_batch <engine> <img_dir> <out_dir>` writes one `.vrtk` per frame: magic, n, dim, then
`n * (x, y, desc[dim])` f32. Keypoints are returned in SOURCE-image pixels — the tool undoes both
the extractor's internal 32 px fit and its own downscale, because a coordinate left in model space
is wrong by the resize ratio in a way the consumer cannot detect.

`lightglue_batch <extractor> <matcher> <img_dir> <pairs.txt> <out.vrtm>` matches a pair LIST and
writes one file for the whole run. An LRU of extraction results, not per-pair extraction: each frame
appears in ~17 pairs of a windowed schedule, so extracting per pair runs the backbone ~34x per
frame, while caching every frame holds n*K*128*4 bytes — 700 MB at 459 frames on a board that is
usually also holding a reconstruction.

Both downscale to fit the shipped engines' shape profile (`256x256 .. 640x640`). Without it the
extractor rejects any frame with a side over 640 — a 480x853 phone preview fails on the long axis,
which is how the first run produced zero files.

Neither sorts nor truncates keypoints: a detector's output order is not a quality ranking, and
truncating raster-ordered output crops the image rather than keeping the strongest points.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…d fail loudly

Three defects found in review, the first two silent and landing directly in a consumer's map.

SIZING WAS CAPPED AT 640 REGARDLESS OF THE ENGINE. The cap was justified by the shipped engines'
profile (`256x256 .. 640x640`), but the engine is argv[1] and a caller may pass one built larger —
the motivating consumer passes a `256x256 .. 1088x1920` build and stages full-resolution frames
specifically to get that precision. Capping anyway reproduced the exact regression it was staging
those frames to avoid: keypoints found on a 352x640 image and multiplied by ~3 to reach source
coordinates, which multiplies their localisation error by ~3 too, while looking like a
full-resolution run. Now the natural (floor-32) size is submitted FIRST and 640 is a fallback taken
only when the engine answers `ShapeRejected` — asking the engine what it accepts instead of assuming
a profile it may not have. Verified: a 1080x1920 frame through the large-profile engine now yields
keypoints spanning x[1,1082] y[0,1921]; it previously capped near 352x640 before rescale.

THE RESIZE INVERSE OMITTED THE HALF-PIXEL TERM. kornia's bilinear forward map is
`dst = s*src + (s-1)/2` (`resize/bilinear.rs`), so inverting it as a bare `x / s` leaves a
`(s-1)/(2s)` bias: -0.18 px on a preview-sized fit, but **-1.03 px at 1080x1920 -> 352x640**. That is
a ~1 px principal-point shift against a ~2 px reprojection target, and the consumer takes intrinsics
from video metadata without freeing them, so it lands in the map. This repo already documents the
hazard in `vrt-lightglue/examples/common/mod.rs`: a bare `diag(s, s, 1)` "silently biases every
rescaled intrinsic and homography".

SILENT ZERO-FILE SUCCESS. Every skip path `continue`d and main returned Ok, so a run that wrote
nothing exited 0. The consumer checks only `status.success()` and reads each absent `.vrtk` as "this
frame has no features" — building a map from empty feature sets. Both tools now exit non-zero when
they produce nothing, or (aliked_batch) fewer files than frames. `lightglue_batch` additionally
writes its `.vrtm` UNCONDITIONALLY before any error return: a transient failure at pair 7000 of 7800
used to propagate out of main and discard the whole run's GPU work.

Also: `lightglue_batch` now remembers frames that failed, instead of re-decoding and re-uploading
them on each of their ~17 pairs — at full resolution that was a 1080x1920 decode every time.

Both sizing policies are kept deliberately identical between the two tools and the reason is stated
in both: the `.vrtm` indices address the keypoints `aliked_batch` wrote, and the two extract
independently, so same engine PLUS same resize is what makes those orderings the same set.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The two bridge tools carried near-identical private resize helpers. That was
not a tidiness problem: `aliked_batch` writes a `.vrtk` of keypoints and
`lightglue_batch` writes a `.vrtm` of index pairs addressing that same
ordering, and the two extract independently. Identical sizing is what makes
those orderings the same set, so two copies that agree today and drift
tomorrow produce a well-formed file addressing the wrong points -- a failure
the consumer already measured once from a different cause (4,062,208
correspondences fed, 62 inliers per surviving pair, a map with a quarter of
the expected points). A comment asking them to match is not an enforcement
mechanism; one function is.

Promote the reviewed `resize_to_fit` out of `vrt-lightglue/examples/common/`
into `vrt-raco-aliked` as `fit_to_engine`, alongside `DIM_DIVISOR` and
`FALLBACK_MAX_SIDE` which it depends on. `vrt-lightglue` already depends on
that crate, so all three call sites now share one implementation; `common`'s
`resize_to_fit` becomes a wrapper holding only the XFeat-grid assertion the
library cannot know about.

This also fixes the anisotropy the duplicated version had. Flooring each axis
to a multiple of 32 independently squashed 1080x1920 by 2.3% horizontally
(sx 0.32593 against sy 0.33333) -- a systematic shear applied to every
keypoint before the geometry ever sees it, and up to 2.9% on Oxford `bark`.
`fit_to_engine` applies one scale to both axes and crops the remainder from
the right/bottom, which leaves the origin and every kept coordinate untouched.
At natural size it now crops rather than resampling at all, so a full-
resolution frame costs no resize and its scale stays exactly 1.

The inverse map moves with it, as `Scaled::to_source`, keeping the half-pixel
term next to the forward map that motivates it rather than in a comment at
each call site.

Verified on 5 keyframes plus 5 pairs through both tools: keypoints land inside
the cropped extent with no overflow past the source width, and the `.vrtm`
indices resolve against the `.vrtk` keypoints into a coherent flow field that
composes -- 0->1 at 19.61 px, 1->2 at 7.07 px, 0->2 at 26.66 px -- which is
only possible if both tools found the same keypoints in the same order.
Four unit tests cover the isotropy, the natural-size crop, the half-pixel
term, and the too-small rejection.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
`written == 0` was the only failure gate, so 7,799 of 7,800 pairs matched
exited 0 and flux-map's `status.success()` check saw a clean run. The two
things that can go wrong are not the same failure and should not share a
verdict.

A DEAD FRAME is a disagreement, not a degraded result. `aliked_batch` exits
non-zero unless it wrote a `.vrtk` for every frame, so if the feature pass
succeeded then every frame has keypoints -- and a frame this tool cannot
extract means the two tools no longer share a frame set. The indices in this
file address the other tool's keypoint ordering, so there is no safe partial
answer: exit non-zero naming the frames.

A FAILED PAIR is a degraded result -- the map loses one graph edge and carries
on. One transient failure should not discard a 40-minute build, so this
tolerates 1% and fails on anything systemic, warning loudly when any pair is
lost at all. The `.vrtm` is written before either check, so a caller that
disagrees with the threshold still has the run's work.

Both counts now appear in the summary line, which previously reported only
matched pairs and could not distinguish a clean run from a lossy one.

Verified: 5 clean pairs exit 0; a pair naming an absent frame exits 1 with
"the two tools disagree on the frame set" and still leaves a 26,048-byte
`.vrtm` on disk.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…ed output

Closes the remaining review findings on the batch bridge.

ENGINE IDENTITY (finding 8). Neither file recorded which engine produced it,
so running `aliked_batch` with the k1024 engine and `lightglue_batch` with
k3072 succeeded and emitted indices up to 3071 into files holding 1024
keypoints -- detectable only by noticing the map came out wrong. A new shared
`engine_fingerprint` (sampled FNV over size plus head and tail, not a full
digest of a 400 MB blob) now goes into the `.vrtk`'s previously-zero reserved
word and into the `.vrtm` header. Verified: the k3072 run stamps 2763df9a on
both files, the k1024 run stamps 96155e77, and the two no longer agree.

The `.vrtm` magic moves `VRTM` -> `VRM2` for this. Appending fields under the
old magic would make a stale reader misparse the fingerprint as its first pair
record; a new magic makes it stop instead.

INPUT VALIDATION (finding 9, nit 11c). `pairs.txt` parsing used
`filter_map(..ok()?)`, so a comma-separated file yielded an empty pair list, a
well-formed 8-byte output and exit 0 -- a total no-op indistinguishable from a
run with no work. Every non-blank, non-comment line must now parse or the tool
names the file and line. Pairs are normalised to `a < b` and deduped so `(a,b)`
and `(b,a)` cannot both be emitted, and a self-pair is rejected outright: it
contributes K zero-parallax correspondences that constrain nothing.

ERROR POLICY (finding 5). The loop mixed two policies -- some failures skipped
a pair silently, while `?` on a sync or readback discarded the entire run's GPU
work at the point of failure. Now every per-pair failure is counted, and the
whole loop is wrapped so nothing can return before the header count is patched
and the file flushed. Output-write failures still abort, because an unwritable
destination leaves nothing to salvage.

STREAMED OUTPUT (finding 10b). The `.vrtm` was accumulated in a `Vec` -- ~112
MB at 7800 pairs, with doubling transients near 224 MB on a 7.4 GB board. It
now streams through a `BufWriter` and seeks back to patch the count.

TRUE LRU (finding 10a). `order` was never refreshed on a cache hit, making the
eviction FIFO despite its name. Harmless at CACHE=24 against a ~13-frame
window, and silently not harmless the moment either number moves.

Also: `lightglue_batch` accepted only `.jpg` while `aliked_batch` accepted
jpg/jpeg/png, making it the one tool that could find nothing in a directory the
other read fine; `aliked_batch`'s extension filter was case-sensitive, so a
`.JPG` directory matched nothing; and its unreachable short-descriptor guard
is now an error rather than a skip, since if the library's buffer contract ever
changes the alternative is a panic mid-directory.

Verified end to end: malformed pairs exit 1 naming the line, a self-pair exits
1, four lines with a duplicate and a transpose collapse to two pairs, and the
healthy 5-pair run is unchanged at 10,307 correspondences.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
CI's `cargo fmt --all --check` was failing on both files. The drift predates
this branch in `aliked_batch` and was introduced by the closure wrapper in
`lightglue_batch`; no logic changes.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
`fit_to_engine_cuda` does the same fit on the GPU: the resize is the SAME
`resize_fast_u8_aa` call, which kornia routes to its CUDA u8 kernels when both
operands are device-resident and documents as bit-identical ("the
coordinate/weight tables come from the same host builders the CPU uses"), and
the crop is `warp_affine_u8` under an identity transform, which carries the
same residency dispatch and the same guarantee.

The geometry is now a shared `Plan`, computed once and used by both paths.
Two copies of that arithmetic is precisely the drift this branch has spent
four commits removing -- a build mixing the two would place keypoints from one
geometry into coordinates computed for the other.

The crop was first written as one `memcpy_dtod` per row. At ~5 us of launch
overhead and 1920 rows that was 11.6 ms against the host path's entire 2.1 ms,
so the launches were the whole cost. It now takes one of two single-call
forms: a contiguous prefix copy when the width is unchanged (every frame
already on the 32 px grid, including all 640 fallbacks), and one warp launch
otherwise. That brought it to 2.24 ms, and the resize+crop case to 0.53 ms
against the host's 1.13 ms.

The batch tools deliberately keep the HOST fit. Their frames arrive as JPEGs,
so the device path would upload the raw 6.2 MB frame where the host path
uploads the fitted 0.68 MB one, and measured end to end that loses:

                                    host fit + upload   device fit + upload
  1080x1920 natural (crop only)          1.95 ms          3.33 ms
  1080x1920 -> 640 (resize + crop)       1.13 ms          1.62 ms

The CUDA path is for callers whose frame is ALREADY on the device, where that
upload does not exist and the 2x faster resize is the whole story. Both
numbers are a few percent of the ~79 ms the extractor takes, so this is a
choice of which resource to spend rather than a bottleneck either way. Written
down in the doc comment so the next reader does not have to re-measure it.

New `tests/gpu_fit.rs` (#[ignore]d, needs a GPU but no engine) asserts the two
paths agree byte-for-byte across crop-only, >2x downscale, landscape, and the
exact-2x pyrdown kernel, on noise with pixel-scale structure so a half-pixel
disagreement cannot hide. Verified separately that a full `aliked_batch` /
`lightglue_batch` run through the CUDA fit produced `.vrtk` and `.vrtm` files
byte-identical to the host fit.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…re the extension set

Cleanup pass. No behaviour change: the `.vrtk` output is byte-identical to the
previous binary, and the host/device equivalence tests still pass.

- The host crop was a hand-rolled row loop; `kornia_imgproc::crop::crop_image`
  is the same operation and brings the NEON strided-row copy and rayon split a
  scalar loop here does not. On the crop-only path that copy IS the cost of the
  fit. Measured unchanged-to-slightly-better despite the extra zero-fill:
  1.95 -> 1.93 ms natural, 1.13 -> 1.01 ms at the 640 cap.

- `copy_rows_dtod`'s doc claimed it did "`ch` stream-ordered D2D row copies…
  row-wise rather than one copy", which is what the body USED to do and what
  the comment two lines below explains was measured at 11.6 ms and rejected. A
  reader trusting the header would have reintroduced that. Renamed
  `crop_top_left_device` and the header now describes the two real branches.

- That function took `src_w`, `cw`, `ch` alongside the buffers they describe,
  so it was possible to hand it a geometry disagreeing with the images — in the
  one function whose whole purpose is that geometry cannot diverge. Read from
  the buffers instead.

- `scale_x`/`scale_y` were computed at both `fit_to_engine` return sites;
  hoisted once, mirroring the CUDA twin.

- `FitError::NotDeviceResident` carried a `&'static str` that only ever took
  "source" or "destination", and "destination" was unreachable (the caller
  allocates it two lines earlier). Payload dropped.

- The accepted image extensions were written twice in two different forms —
  a case-sensitive `matches!` in one tool, a hardcoded `.jpg` probe list in the
  other. That is the same divergence class the shared resize exists to prevent,
  and its failure mode here is the dead-frame hard error. Now one exported
  `FRAME_EXTS`, compared case-insensitively on both sides.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@edgarriba
edgarriba merged commit 9b73859 into main Aug 11, 2026
3 checks passed
edgarriba added a commit that referenced this pull request Aug 13, 2026
…or out-of-process consumers (#25)

Completes the file bridge that `aliked_batch` and `lightglue_batch` started (#24).
flux-map spawns five external tools; those two answered "what is in this frame" and
"which keypoints match". These two answer "which frames look alike" (retrieval) and
"how far away is each pixel" (metric depth), and both lived out of tree.

Same reason for existing as the merged pair: the consumer pins a different kornia than
these crates, so the two sides cannot be linked and meet on files instead.

dino_batch  <engine> <img_dir> <out.vrtb>   one L2-normed CLS descriptor per image
depth_batch <engine> <img_dir> <out_dir>    one f32 metric raster per image

Ported to the conventions the merged pair established, which the out-of-tree versions
predated:

- Case-INSENSITIVE extension filter. A `.JPG` directory matched nothing before, and the
  tool then reported success over zero frames.
- Zero frames in, or a partial run out, is a non-zero exit. The caller checks only
  `status.success()` and cannot see stderr; a silent empty success reads downstream as
  "retrieval found nothing" / "these frames have no depth", both of which degrade the
  map rather than failing it.
- Buffer-contract violations are errors, not truncation: a ragged descriptor row, or a
  raster whose length disagrees with its declared grid.
- `depth_batch` takes the grid from the RESULT it is wrapping, not from the model.

`.vrtb` additionally grows a name table after the descriptor block, and word 12 - a
reserved zero - now holds its length. The rows were addressed by position alone, which
is the caller's frame index only if the caller's inputs are gapless; flux-map writes
`kfNNNN.jpg` only for keyframes it holds a thumbnail for, so one gap shifted every
later descriptor onto the wrong keyframe with nothing in the file able to say so.
Measured on device with a deliberate gap: row 4 reads as keyframe 5 positionally, and
the table says kf0006. The descriptor block stays at offset 16 in the same layout, so
a reader that only wants descriptors is unaffected.

Verified on the Orin against the shipped engines: 7 frames (one `.JPG`, one gap) ->
10817-byte `.vrtb` = 16 + 7*384*4 + 49 name bytes, every row unit-norm, names in row
order with the gap preserved; 7 x 614672-byte `.vrtd` = 16 + 392*392*4, median depth
4.12 m. Empty directory exits 1 and writes no file.

10 tests over the two file formats, run in CI (`--examples`, the same way the
vrt-lightglue harness tests run). Against the pre-port implementations 7 of the 10
fail; the 3 that pass on both are the layout-compatibility assertions, which are
supposed to.

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant