ML researcher & engineer. ECE + math at Rutgers ('27) — not your average CS major (I'm Engineering). Previously: Google (Ads AI/ML), NASA, Goldman Sachs, Columbia AI. First author of arXiv:2601.18710. Founder of the Grey Matter Society at Yale School of Medicine.
The thread through everything I build: measurements that survive scrutiny — leave-one-subject-out splits, false-discovery-rate control, confidence intervals on everything, and reporting the null results most people bury.
- interpretability & evals — interp asks why a model made a prediction (logit lens, activation patching, SAE features) from inside your coding agent · evalkit is evals done right — bootstrap CIs, unbiased pass@k, judge-bias controls · steering-audit found that only 31–50% of "successful" activation steers stay coherent — the classifier can't tell degenerate text from a steered one · cross-sae matches sparse-autoencoder features between vision models and human visual cortex, with the false-discovery rate actually controlled
- ML for medicine — thaakat: endometriosis detection from pelvic MRI, AUC 0.961 on a patient-level split (the honest kind) · vigil: pilot cognitive state from EEG/ECG, where fixing the leaky eval cost 0.19 AUROC — I kept the honest number · aurelis: an LLM judge for clinical notes, validated against human graders · neuroloop: closed-loop neuromodulation graphs that prove their safety envelope before they run
- biological time — the niche I keep coming back to: circa · rhythmrx · fitra — the right intervention at the right circadian phase
- carbonium — with @RyanRana, exploring whether an abductive engine over fleet-maintenance data can name a failure and its evidence chain — and be evaluated with confounder traps so it can't bluff
I'd rather fix the ruler than argue about the measurement. PRs to the tools researchers actually run —
- merged · TransformerLens — direct logit attribution · inspect_evals (UK AI Security Institute) — MedCalc-Bench · nnsight — multi-invoker
.backward()fix · MOABB — Nadeau–Bengio corrected t-test, so EEG benchmarks stop reporting optimistic p-values · braindecode — spec-compliant zarr JSON in Hub stores - in review · HELM — a regex that scored correct answers 0, ×2: opt-in dedup for MedHELM's ~33%-duplicate dialog test set · Medplum — GraphQL error semantics · MONAI — invertible
NormalizeIntensity· SAELens — covariance whitening · garak — NonEngagement detector · pyvene — ELECTRA support · lifelines — RMST was returning the wrong variance (~34× inflated); now Greenwood, validated against survRM2 · PyHealth — evaluation loss was batch-size-dependent, quietly steering checkpoint selection; now example-weighted and pooled · TorchIO — Motion's "rigid" transform moved content along the wrong axis by the wrong amount; now convention-aligned with Affine · dynamo ×2 · SkyPilot · FlashInfer - in discussion · lm-evaluation-harness #3080 — design for per-instance repeats aggregation with unbiased pass@k
azra-bano.com · linkedin · arxiv · email

