Version en español: README.es.md
I built this after noticing how brittle a fixed dashboard is for the kind of question people actually ask about a mining operation. A dashboard answers exactly the questions its screens were designed for — this month's flotation recovery, this week's maintenance alerts, this applicant's risk score — but the moment someone asks something slightly off that shape ("was September's recovery normal, and is any of the equipment that's been down a lot also flagged as anomalous?"), you either need a new screen or someone runs an ad hoc query by hand. Neither scales, and the second one is exactly where numbers get misremembered or fudged under time pressure.
What I wanted to test here is whether an LLM tool-calling loop gives you something better than either option: a natural-language front end that can answer a wider range of ad hoc questions than any fixed dashboard, while staying grounded — every number in its answer has to come from an actual tool call (a real DuckDB query, a real model prediction), not from the model's own guess at what a plausible-sounding number would be. That's the whole point of routing through tools=[...] in the chat-completions API instead of just asking the model to answer from context: the model can request a tool, but it cannot fabricate a tool_result.
Everything a tool touches — the operational data and the credit-risk model — is synthetic and generated inside this repo, with a fixed random seed (numpy.random.default_rng(42)), so the whole pipeline (schema, data, model, tools, tests) is reproducible from a clean clone with python -m src.setup_data.
flowchart LR
U[User question] --> A[MiningOpsAgent loop]
A -->|tool_use requested| D{Dispatcher}
D --> T1[warehouse_query_tool<br/>DuckDB]
D --> T2[credit_risk_tool<br/>LogisticRegression]
D --> T3[anomaly_check_tool<br/>IsolationForest]
T1 --> R[tool_result]
T2 --> R
T3 --> R
R --> A
A -->|end_turn| F[Final text answer]
The loop lives in src/agent.py (MiningOpsAgent). It sends the user message and the tool schemas to the model; if the response asks for tool_use, it dispatches to the matching Python function, wraps the result (or the exception, if the tool raised one) as a tool_result, and sends it back — up to a bounded number of iterations (default 5) so a misbehaving model can't loop forever.
Two ways to drive it: agent.run(question) is stateless — one question in, one answer out, no memory of anything before or after. agent.chat(question) is the multi-turn version: it appends every turn (including the intermediate tool_use/tool_result exchanges) to agent.history and resends the growing history on each call, so a follow-up like "and what about August?" resolves against the prior answer without the caller re-stating the question. agent.reset() clears it. src/cli.py exposes both: python -m src.cli "question" for one-shot, python -m src.cli (no args) for an interactive multi-turn REPL.
All three are plain, synchronous Python functions with a companion JSON schema (TOOL_DEFINITIONS) in the shape the tool-use API expects ({"name", "description", "parameters"}), wrapped into the provider's function-tool envelope in src/tools/__init__.py. Each one runs, and is tested, with no LLM and no API key involved.
| Tool | File | What it does |
|---|---|---|
get_flotation_summary, get_maintenance_alerts, get_procurement_summary |
warehouse_query_tool.py |
Fixed, named, parameterized queries against a local DuckDB (data/ops.duckdb) — not arbitrary SQL from the user, by design. |
score_credit_risk |
credit_risk_tool.py |
Scores a credit applicant profile with a small self-contained LogisticRegression trained on synthetic applicants (data/credit_risk_model.joblib). |
check_maintenance_anomalies |
anomaly_check_tool.py |
Runs an IsolationForest over recent maintenance events to flag equipment with unusual downtime/severity patterns. |
The DuckDB warehouse has three synthetic tables: flotation_batches, maintenance_events, procurement_orders, all generated by src/setup_data.py with numpy.random.default_rng(42).
| Technique | Where | What it's for |
|---|---|---|
OpenAI tool-use API (openai Python SDK, chat.completions.create(tools=...)) |
src/cli.py |
Lets the model decide which tool to call and with what arguments, instead of hand-coded intent matching. |
Bounded agent loop (MiningOpsAgent.run) |
src/agent.py |
Dispatches tool_calls to real Python functions, feeds tool messages back to the model, and caps iterations (default 5) so a stuck model can't loop forever. |
| DuckDB, parameterized queries | src/tools/warehouse_query_tool.py |
Fast local OLAP queries over the synthetic operations warehouse; parameters are bound (? placeholders), never string-interpolated SQL. |
LogisticRegression (scikit-learn) |
src/tools/credit_risk_tool.py, src/setup_data.py |
A small, self-contained credit-risk scoring model trained on synthetic applicant data; returns a probability of default and a risk tier. |
IsolationForest (scikit-learn) |
src/tools/anomaly_check_tool.py |
Unsupervised outlier detection over per-equipment maintenance features (event count, total downtime, share of critical/high-severity events) to flag anomalous equipment without a hand-set threshold. |
Oracle-AUC ceiling check (GradientBoostingClassifier vs. the true generating probability) |
src/model_ceiling_check.py |
Verifies a weak AUC is the label's own noise, not a fixable model choice, by scoring the true probability behind each synthetic label against the held-out set and checking no model — deployed or more flexible — can beat it. |
unittest.mock fake OpenAI client |
tests/test_agent.py, tests/test_cli.py, tests/test_eval_harness.py |
Verifies the dispatcher's routing logic (correct tool, correct args, correct tool message shape, malformed-argument handling, iteration cap, multi-turn history growth), the CLI's REPL, and the eval harness's own scoring logic — all without a real API call. |
Multi-turn conversation state (MiningOpsAgent.chat) |
src/agent.py |
Accumulates turns (including intermediate tool calls) into agent.history so a follow-up question resolves against prior context, instead of every question starting from a blank slate. |
| Tool-selection + groundedness eval harness | src/eval_harness.py |
A labeled query set (EVAL_QUERIES) checks whether the model calls the right tool(s) for a question, and a number-extraction check (check_groundedness) verifies every number in its final answer traces back to a real tool result rather than a plausible-sounding guess. |
Every number and figure below comes from python -m src.generate_report (src/visualization/plots.py) calling the exact same tool functions the agent dispatches at runtime and plotting their real outputs — nothing here is hardcoded or re-simulated separately from the tools. The full pipeline (setup_data.py → generate_report.py → pytest, 42/42 passing, CI-enforced on every push) is reproducible from a clean clone, and the three results below tell three different stories about what happens when you actually check a tool's output instead of trusting the number it returns:
score_credit_risklooks weak (AUC 0.586) until you compute the theoretical ceiling and find it's already capturing 96% of the AUC that's actually available — the model isn't underperforming, the label is just noisy by design.check_maintenance_anomaliesflags equipment a simple downtime ranking would miss entirely, including one flagged for having too little downtime — evidence the Isolation Forest is using more than one signal, not just proof it runs.get_flotation_summary/get_procurement_summaryground the other two tools in an operational baseline: a real 12-month recovery trend and a real spend breakdown the agent can quote instead of guess.
| Tool | Metric | Value |
|---|---|---|
score_credit_risk |
ROC-AUC (held-out test, n=400) | 0.586 |
score_credit_risk |
PR-AUC (base rate 0.245) | 0.335 |
score_credit_risk |
Test accuracy | 0.755 |
check_maintenance_anomalies |
Equipment flagged (60d window) | 3 / 24 |
get_flotation_summary |
Months of data | 12 |
Three panels, all from the same held-out test set (n=400): the ROC curve (left) hugs the no-skill diagonal — it never gets far above it, which is what a ROC-AUC of 0.586 looks like when you actually plot it instead of just reporting the number. The PR curve (center) spikes near recall=0 (a handful of confident, correct high-probability predictions) and then decays fast toward the 0.245 base rate, which is the realistic floor once recall increases. The histogram (right) makes the reason visible: the predicted-probability distributions for "default" and "no default" overlap almost completely.
That overlap used to just be asserted as "the model is deliberately simple." It's now measured. setup_data.py computes a true_prob_default for every synthetic applicant before flipping the coin that decides their default label — so the true probability behind each label is known, not estimated. Scoring that probability against the real held-out labels (src/model_ceiling_check.py) gives the best AUC any model could ever achieve here, because the only randomness left once you know the true probability is the coin flip itself:
| Model | Held-out AUC | % of ceiling captured |
|---|---|---|
| Oracle (true probability, not fit from data) | 0.611 | 100% (by definition) |
LogisticRegression (deployed in score_credit_risk) |
0.586 | 96.0% |
GradientBoostingClassifier (comparison only, not deployed) |
0.607 | 99.4% |
The deployed logistic model is already capturing 96% of the AUC that is theoretically available — and swapping in a considerably more flexible gradient-boosted model only closes the remaining gap to 99.4%, it doesn't blow past the ceiling. That rules out the obvious alternative explanation for a weak AUC (an underpowered model choice): the label itself is this noisy by construction (a Bernoulli draw around a probability that only ranges roughly 0.04–0.88 across applicants), and no model — however expressive — can out-predict a coin flip it doesn't get to see. tests/test_model_ceiling_check.py locks this in: it asserts no model's AUC can exceed the oracle's, and that the deployed model captures at least 85% of it.
Interactive version: predicted P(default) vs. debt-to-income, all 400 held-out applicants, hover for the full profile — opens a live Plotly chart (self-contained HTML, generated by plot_credit_risk_interactive() in src/visualization/plots.py) rather than a static image.
Two real applicant profiles scored with score_credit_risk (called directly, no LLM in the loop):
| Profile | Age | Income (CLP/mo) | Debt-to-income | Months employed | Late payments | Requested (CLP) | probability_default |
risk_tier |
|---|---|---|---|---|---|---|---|---|
| Low-risk example | 42 | 1,400,000 | 0.15 | 96 | 0 | 1,500,000 | 0.0823 | low |
| High-risk example | 24 | 380,000 | 0.92 | 3 | 6 | 5,500,000 | 0.7892 | critical |
Both rows are one real call to score_credit_risk(...) each, run in this session against the model committed at data/credit_risk_model.joblib — not hand-picked to look clean, just two profiles at opposite ends of the input ranges the model was trained on.
3 of 24 pieces of equipment get flagged by the IsolationForest, and the flagged ones are not simply the three with the most downtime — that's the interesting part of this chart. EQ-007 has more total downtime (23.2h) than the flagged EQ-009 (22.6h) but isn't flagged; EQ-023 gets flagged at a mid-range 10.4h, well below several normal bars; and EQ-024 is flagged despite having almost no downtime at all (0.5h) — anomalous for being unusually low, not high. That's expected from an Isolation Forest run over multiple maintenance-event features (frequency, severity, downtime) rather than a single downtime threshold, and it's a more realistic signal than "flag whatever has the biggest bar."
The GIF above animates the flotation recovery trend across the 12 months; the PNG below is the static reference for detailed reading.
Left: monthly average flotation recovery over the synthetic 12-month window — it oscillates in a fairly tight 87–89.5% band, with a visible dip to 86.7% in April 2026 before recovering to a 12-month high of 89.4% in August 2026. Right: procurement spend by category, ranked services > fuel > safety_equipment > reagents > spare_parts — services and fuel alone account for roughly 45% of total procurement spend in this synthetic warehouse, ahead of consumables like reagents and spare parts.
Three more warehouse tool calls, run directly in this session (no LLM involved):
>>> get_flotation_summary("2025-09")
{'n_batches': 18, 'avg_feed_grade_pct': 0.83, 'avg_recovery_pct': 88.87,
'avg_concentrate_grade_pct': 28.28, 'total_tonnage_processed': 26282.7, 'month': '2025-09'}
>>> get_procurement_summary("delayed")
{'status_filter': 'delayed', 'n_orders': 66, 'total_amount_usd': 514709.88, 'avg_amount_usd': 7798.63}
>>> check_maintenance_anomalies(days=60)
{'window_days': 60, 'n_equipment_evaluated': 24, 'n_equipment_flagged': 3,
'flagged': [
{'equipment_id': 'EQ-009', 'n_events': 7, 'total_downtime_hours': 22.53, 'critical_share': 0.286, 'anomaly_score': -0.0365},
{'equipment_id': 'EQ-024', 'n_events': 2, 'total_downtime_hours': 0.6, 'critical_share': 0.5, 'anomaly_score': -0.0174},
{'equipment_id': 'EQ-023', 'n_events': 3, 'total_downtime_hours': 10.46, 'critical_share': 0.667, 'anomaly_score': -0.007}]}
The system prompt says "never invent numbers that a tool could return" — src/eval_harness.py turns that from an instruction into something checkable. It runs a labeled set of queries (EVAL_QUERIES, including a multi-tool query and an out-of-scope one where no tool should fire) through the real agent and real tools, and scores two things per query:
- Tool-selection accuracy — did the model call exactly the tool(s) the query actually needs?
- Groundedness — does every number in the final answer trace back to a real tool result (
check_groundedness), or did the model state something that no tool call actually returned?
OPENAI_API_KEY=sk-... python -m src.eval_harnessHonest limitation: no OPENAI_API_KEY was available on the machine this was built on, so this repo does not claim a live accuracy number — that would mean reporting a metric nobody actually measured, exactly what the rest of this README argues against. What is verified, in tests/test_eval_harness.py, is that the scoring logic itself is correct: extract_numbers/check_groundedness are tested against hand-built cases (a grounded answer, a fabricated number, no tool calls at all), and run_eval is tested against scripted mock responses for both a correct and an incorrect tool selection. Run the harness with a real key to get real numbers.
python -m venv .venv
.venv\Scripts\activate # Windows PowerShell: .venv\Scripts\Activate.ps1
pip install -r requirements.txt-
Generate the synthetic data and train the risk model:
python -m src.setup_data
Writes
data/ops.duckdbanddata/credit_risk_model.joblib. -
Generate the result figures and metrics summary:
python -m src.generate_report
Writes
reports/figures/*.png,reports/figures/*.gif,reports/metrics.json, andoutputs/interactive/credit_risk_scores.html. -
Run the test suite (works fully offline, no API key needed):
pytest
-
Talk to the agent (requires a real
OPENAI_API_KEY):# One-shot, stateless: python -m src.cli "What was the flotation recovery in September 2025?" # Interactive, multi-turn (remembers context across questions until you type 'reset'): python -m src.cli
Without a key set, either mode exits with a clear message instead of a traceback.
-
Run the evaluation harness (also requires a real
OPENAI_API_KEY— see Evaluation harness above):python -m src.eval_harness
data/ops.duckdb, data/credit_risk_model.joblib, reports/figures/*.png + reports/metrics.json, and outputs/interactive/credit_risk_scores.html are committed to the repo rather than gitignored. They are small, fully synthetic/derived, deterministically regenerable (python -m src.setup_data && python -m src.generate_report), and versioning them means the tools (and the results above) are visible immediately after a clone without a mandatory setup step. .gitignore still excludes the virtual environment and caches.
Pablo Reyes — Data Scientist, Santiago, Chile.




