Curate billion-scale image-text datasets for vision-language pretraining with Backblaze B2 as the sole storage layer. This is a working implementation of the DataComp filtering workflow: raw web-scraped image-text pairs live in B2 as WebDataset .tar shards, a configurable pipeline streams those shards from B2, scores image-text alignment with CLIP (via open_clip), re-packs the passing pairs into new shards under filtered/, and writes per-shard quality metrics as JSON under metrics/ — all over the S3-compatible API, no local staging, no second API key.
The point: you filter a noisy pool down to high-quality shards without pulling it locally, and the filtered output streams straight back to training jobs from B2.
- Full-stack: Next.js 16 + React 19 + Tailwind v4 + shadcn/ui frontend, FastAPI backend with a strict layered architecture.
- B2 is the only store — the raw pool, filtered output, run manifests, and metrics are all objects in one bucket. No database.
- Runs on-device — CLIP scoring runs locally (CPU by default, CUDA/Apple MPS auto-detected). $0 in external API cost; only your B2 credentials.
Dashboard — filter-run totals, pairs kept, and average storage reduction, over a table of the most recent runs.
Filter Runs — every CLIP filter job with its strategy, status, kept/in counts, and reduction, plus the full create/run/delete lifecycle.
Run detail — a completed run's configuration and results alongside per-shard metrics and the per-pair CLIP scores with kept/dropped decisions.
Pool Explorer — open a WebDataset shard to inspect the image-text pairs inside it: thumbnail, caption, and resolution.
Ingest — upload raw WebDataset shards or image-text assets straight into B2 with a presigned, direct-to-bucket PUT.
You need: Node.js ≥ 20, pnpm ≥ 9, Python ≥ 3.12, and a free Backblaze B2 account.
# 1. Install deps, create the Python venv, copy .env.example -> .env
pnpm run setup
# 2. Fill in your B2 credentials (see the env table below)
# edit .env
# 3. Seed a small synthetic image-text pool into B2 (keyless, no download)
services/api/.venv/bin/python services/api/scripts/seed_pool.py
# 4. Install the heavy ML stack (torch/open_clip/webdataset) — needed to RUN a
# filter. It is deliberately excluded from `pnpm run setup` and CI.
services/api/.venv/bin/pip install -r services/api/requirements-ml.txt
# 5. Start the app, then create + run a Filter Run in the UI
pnpm devOpen http://localhost:3000, go to Runs → New run, accept the defaults, Start the run, then open the Pool Explorer to see which pairs were kept vs dropped and why.
Use
http://localhost:3000(not127.0.0.1) — dev CORS is scoped to thelocalhostorigin.
- Ingest — raw image-text pairs land in B2 as WebDataset
.tarshards underpool/(the seed script generates a synthetic pool; the Ingest page uploads your own via presigned direct-to-B2 PUT). - Filter — a Filter Run streams each
pool/shard from B2 and scores every image-text pair with CLIP. - Write — passing pairs are re-packed into
filtered/<run>/*.tarand written back to B2. - Score — per-shard quality metrics (kept/dropped counts, mean CLIP score, threshold) are written as
metrics/<run>/*.json, and the run manifest atruns/<id>/manifest.json. - Serve — filtered shards stream straight back to training jobs from B2, in the same WebDataset format.
- DataComp-style CLIP-score filtering — stream shards from B2, score image-text alignment with CLIP, keep the top-percentile pairs (DataComp's
clip_scorebaseline). - Configurable baseline filters — CLIP-score percentile, min image resolution, caption-length bounds, and near-duplicate removal — the
clip_score/basic/image_based/text_basedfamilies from DataCompbaselines.py, chosen on the Run form. - Filter Runs, full lifecycle in the UI — create, read, edit (pending), delete, and run. B2 manifests are the sole store; no database.
- Pool Explorer — open a shard to inspect the image-text pairs inside it: thumbnail + caption + CLIP score + kept/dropped. Scoped to
pool/andfiltered/. - Bucket Explorer — browse every object in the bucket (pool, filtered, manifests, metrics) with preview/download/delete.
See docs/features/ for per-feature detail.
pnpm run setup copies .env.example to .env. Fill in these (from your B2 bucket + application key):
| Variable | Required | What it is |
|---|---|---|
B2_APPLICATION_KEY_ID |
yes | B2 application key ID (the S3 access key ID) |
B2_APPLICATION_KEY |
yes | B2 application key (the S3 secret) |
B2_BUCKET_NAME |
yes | Bucket that holds the pool + filtered output |
B2_REGION |
yes | Region slug, e.g. us-west-004. The S3 endpoint is built at runtime as https://s3.<B2_REGION>.backblazeb2.com — no region is hardcoded anywhere |
B2_PUBLIC_URL_BASE |
no | Public base URL for a public bucket (builds direct object URLs). Leave unset for a private bucket |
The application key needs listBuckets, listFiles, readFiles, writeFiles, and deleteFiles.
- Model:
open_clipOpenAI CLIPViT-B-32(pretrainedopenai) by default, selectable up toViT-L-14. Weights are ungated and keyless — downloaded once (~350 MB for ViT-B-32) from open_clip's own hosting. No Hugging Face token, no gated terms. - Device: auto-detected at runtime — CUDA → Apple MPS → CPU, defaulting to CPU. Never hard-requires a GPU.
- Dependency split: the heavy stack (
torch/torchvision/open_clip_torch/webdataset) lives inservices/api/requirements-ml.txt, excluded frompnpm run setupand CI so the app boots and every static gate passes ML-free.service/filtering.pylazy-imports it; if it is absent, a run is persisted asfailedwith an actionable message (the POST never 500s).
pnpm run setup # idempotent cold-start setup (.env copy, deps, venv)
pnpm dev # start frontend + backend
pnpm contract:export # export the FastAPI OpenAPI contract
pnpm contract:check # verify the OpenAPI artifact + frontend route registry agree
pnpm check:agent-docs # agent instruction/documentation drift check
pnpm verify # credential-free pre-PR suite
pnpm verify:api # backend half (lint, tests, structure)
pnpm verify:web # frontend half (lint, unit tests, typecheck + build)
pnpm verify:full # doctor + verify + Playwright E2EThe API runs CLIP locally, so a serverless deploy runs the UI, CRUD, and explorers but not the filter engine (which needs the ML stack and on-device compute). Deploy the full app as one Vercel project (web + api services in vercel.json, one origin):
After deploying a browser-upload origin, add a bucket CORS rule for it:
services/api/.venv/bin/python services/api/scripts/setup_b2_cors.py --origin https://your-app.vercel.app --applySee infra/vercel/README.md and infra/railway/README.md for the full delivery contracts.
Backend layers: types/ → config/ → repo/ → service/ → runtime/, with boto3 confined to repo/ and the ML stack confined to service/filtering.py — both enforced by structural tests. See ARCHITECTURE.md and AGENTS.md.
When should I use this? As a reference for building a DataComp-style dataset-curation pipeline on object storage, or to evaluate B2 as the storage layer for VLM-pretraining data curation.
When should I not use this? As-is at petabyte scale — the demo holds each shard in memory (documented simplification); production streams via WebDataset's S3 reader and runs the engine on a GPU fleet. The UI is unauthenticated and bucket-wide (single-tenant demo).
Does it cost anything to run? No external API cost — CLIP runs locally. You pay only for B2 storage/egress. A free B2 account is enough to run everything here.
Do I need a GPU? No. CLIP scoring auto-detects CUDA → Apple MPS → CPU and defaults to CPU.
MIT — see LICENSE.




