Production-style card pose tracking for video: a stable perspective quad with consistent corner identity, occlusion-aware homography, and a compact export the browser can render at 60fps.
The hard problem is not “find a rectangle in an easy frame.” It is keeping a persistent geometric object while the card rotates, the perspective changes, and fingers cross the surface.
This repository is the implementation of that pipeline, plus a trial harness that exposes the failure modes of naive per-frame corner detection.
A card is a planar object. Its image is a homography from a canonical plane whose corners are always TL, TR, BR, BL. Pose is estimated from many interior correspondences with RANSAC, not from four independently detected corners.
video frame
│
├─ initialize (contour / ArUco) → lock corner identity once
│
├─ track interior points (Lucas–Kanade; CoTracker optional)
├─ segment fingers/hands (skin / MediaPipe / SAM2)
├─ drop occluded & failed tracks
├─ RANSAC homography (canonical plane → image)
├─ reject impossible jumps; hold last good pose
├─ One-Euro temporal smoothing
│
└─ export track.json + packed occlusion masks
└── browser applies CSS/WebGL perspective at display refresh
Corner identity is never assigned by sorting points on x/y. That ordering flips as soon as the card rotates past a perspective configuration. Identity is carried by the homography and by cyclic assignment to the previous quad.
cd "C:\Users\ebima\Desktop\Object Tracking"
python -m venv .venv
.\.venv\Scripts\Activate.ps1
pip install -e ".[dev]"
cardtrack trial --pattern dots --out outputs/trialThis generates a rotating card with simulated finger occlusions, runs the tracker, writes:
| file | purpose |
|---|---|
outputs/trial/clip/video.mp4 |
synthetic source |
outputs/trial/clip/gt.json |
ground-truth quads |
outputs/trial/run/track.json |
compact browser payload |
outputs/trial/run/preview.mp4 |
debug overlay |
outputs/trial/trial_report.json |
IoU, identity-flip rate, hold/lost stats |
Then open web/viewer.html, load the video and track.json. The client does no detection or tracking.
Measured on the synthetic trial (60 frames, rotation + fingers):
| pattern | mean IoU | identity flips | hold fraction | mean corner error |
|---|---|---|---|---|
dots |
0.96 | 0 | 0 | 5.3 px |
aruco |
0.91 | 0 | 0 | 12.9 px |
blank |
0.69 | 0 | 0.82 | 62.6 px |
A textureless white card does not lose identity — the hold policy prevents flips — but it cannot keep a tight pose. The micro-dot grid is the production default: it looks almost blank and supplies the interior correspondences RANSAC needs. Four ArUco markers are an explicit upper bound for initialization, not automatically a better homography than a dense subtle grid.
To compare a blank card against a trackable pattern and against naive / CSRT baselines:
cardtrack benchmark --out outputs/benchmark --frames 90Drop a capture into data/raw/ and run:
cardtrack process "data/raw/your_clip.mp4" --out outputs/real --config configs/trial.yamlBatch:
cardtrack batch data/raw --out outputs/batchIf you generate the blank cards, a subtle interior pattern removes most downstream fragility. A white rectangle has almost no trackable texture; optical flow and RANSAC then depend on edges that fingers occlude.
cardtrack generate-card --pattern dots --out data/processed/card_dots.png
cardtrack generate-card --pattern aruco --out data/processed/card_aruco.png| pattern | look | tracking |
|---|---|---|
blank |
clean white | worst case; documents failure |
dots |
low-contrast micro-dot grid | recommended production default |
edge |
faint border ticks | better than blank, weak in the interior |
aruco |
four small corner markers | maximum reliability / debug |
charuco |
calibration board | pose-lab / evaluation |
A small upstream print change is cheaper than an ever-more-complex reconstruction stack on a textureless surface.
src/cardtrack/
pipeline.py persistent tracker + hold/reinit policy
detection/ contour quad + ArUco init
tracking/ LK points, OpenCV bbox baselines, CoTracker adapter
geometry/ identity, homography, One-Euro smoothing
occlusion/ skin / MediaPipe Hands / SAM2 adapter
patterns/ printable card generators
export/ track.json + packed 1-bit masks
eval/ IoU, identity flips, backend comparison
synth/ rotating-card clips with ground truth
io/ video + batch
web/ 60fps overlay viewer
- Confidence from inlier ratio, point count, occlusion fraction, and homography condition.
- If confidence drops: hold the last reliable quad (no overlay jump).
- After
hold_max_frames, re-detect and re-seed points, matching the new quad to the locked identity. - Temporal smoothing is applied only to accepted poses so jitter is not mixed with stalls.
CSRT / KCF / MOSSE are evaluated as baselines. They track a box, not a perspective quad, and they cannot keep corner identity. They are useful as a “still in view” prior, not as the pose source. The production pose is always a homography from tracked interior points.
CoTracker and SAM2 are real backends (tracking/cotracker.py, occlusion/ + models/).
Default occlusion is auto: YOLO11n-seg on CPU, SAM2 tiny when CUDA is available, unioned with a skin fallback so synthetic fingers still mask. Enable CoTracker with --set tracking.backend=cotracker.
Weights download once into models/weights/ (yolo11n-seg.pt, sam2.1_t.pt).
track.json is intentionally small:
{
"v": 1,
"w": 960,
"h": 540,
"fps": 30,
"order": ["tl", "tr", "br", "bl"],
"frames": [
{ "i": 0, "s": "tracking", "q": [0.12, 0.20, 0.48, 0.18, 0.50, 0.62, 0.10, 0.64], "c": 0.91, "o": 0.14, "n": 47 }
]
}Coordinates are normalized. The viewer maps them through a CSS matrix3d (or your existing WebGL layer) so the overlay stays in the browser’s render loop.
Optional masks.bin stores packed 1-bit occlusion masks at 1/8 resolution for finger cutouts without shipping a PNG sequence.
configs/default.yaml is production. configs/trial.yaml is slightly less aggressive on smoothing so failure modes stay visible. Override any key:
cardtrack process clip.mp4 --out outputs/run --set occlusion.backend=mediapipe --set tracking.max_points=200MediaPipe Hands: pip install -e ".[hands]".
pytestGeometry tests lock the invariant this whole project is built on: identity survives rotation; x/y sorting does not.
- Architecture — design decisions, backends, export
- Trial protocol — how to score a real clip and what to log