I'm a Data Scientist working across Applied AI and Machine Learning, with an engineering background and a particular interest in what happens after “the model works.”
Does it still work when the data gets messy?
What happens when the assumptions change?
Is the retrieval system actually finding the right evidence?
Should every workload be going to the same model?
What happens in the tail, not just at the average?
A suspicious number of my projects start with “I wonder what happens if...” and end with several more experiments than originally planned.
I’m interested in evaluating AI systems beyond a single performance metric. That means looking at quality, reliability, latency, cost, robustness, calibration, and failure modes, and understanding the tradeoffs between them.
I experiment with RAG, information retrieval, and multi-hop reasoning, including whether retrieving smaller reasoning units such as sentences, claims, or evidence can outperform traditional document chunking.
Not every workload needs the same model. I’m exploring how AI systems can route workloads across models and tools based on quality, cost, latency, reliability, and operational constraints.
I benchmark emerging model architectures to understand where their advantages hold, where they disappear, and what happens when the data becomes less cooperative.
I use simulation, stress testing, and sensitivity analysis to understand how systems behave under uncertainty, particularly when the interesting behavior lives in the tails.
Applied AI & Machine Learning · AI Evaluation & Reliability · Retrieval & Reasoning · Foundation Models · Experimentation · Causal Inference · Robustness · Decision-Making Under Uncertainty

