Measurement scientist who builds. I work where human judgment meets automated scoring — model evaluation, people analytics, and assessment systems that have to be defensible, not just accurate.
My background is a master's in industrial-organizational psychology with an advanced quantitative methods focus, plus hands-on engineering with machine learning and large language models. The short version: I can train the model and tell you whether its scores actually mean anything.
Open to roles in AI evaluation, people analytics, and psychometric data science.
Does the Model Rate Like a Human? — an LLM scorer for open-ended responses,
validated like an assessment: inter-rater reliability (ICC, weighted kappa, Krippendorff's
alpha) and a fairness audit (Cohen's d, four-fifths adverse-impact rule).
LLM eval · psychometrics · fairness
Predicting Attrition Without the Black Box — turnover prediction built to be
explained and audited: logistic odds ratios tied to turnover theory, permutation importance,
and per-group error-rate parity.
applied ML · interpretation · fairness
What Employees Actually Say — open-ended survey comments turned into themes, then
linked to engagement so the output is what to act on, not a word cloud.
NLP · topic modeling · constructs
Measurement reliability & validity · factor analysis · IRT · adverse-impact analysis
ML scikit-learn · model evaluation & interpretation · experiment design
LLM/AI Anthropic & OpenAI APIs · LLM-as-evaluator pipelines · rubric design · RAG
Eng Python (pandas, NumPy) · Git · reproducible analysis · R