Skip to content

Latest commit

 

History

13 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Chile Mining Ops Agent

Version en español: README.es.md

tests

Overview

I built this after noticing how brittle a fixed dashboard is for the kind of question people actually ask about a mining operation. A dashboard answers exactly the questions its screens were designed for — this month's flotation recovery, this week's maintenance alerts, this applicant's risk score — but the moment someone asks something slightly off that shape ("was September's recovery normal, and is any of the equipment that's been down a lot also flagged as anomalous?"), you either need a new screen or someone runs an ad hoc query by hand. Neither scales, and the second one is exactly where numbers get misremembered or fudged under time pressure.

What I wanted to test here is whether an LLM tool-calling loop gives you something better than either option: a natural-language front end that can answer a wider range of ad hoc questions than any fixed dashboard, while staying grounded — every number in its answer has to come from an actual tool call (a real DuckDB query, a real model prediction), not from the model's own guess at what a plausible-sounding number would be. That's the whole point of routing through tools=[...] in the chat-completions API instead of just asking the model to answer from context: the model can request a tool, but it cannot fabricate a tool_result.

Everything a tool touches — the operational data and the credit-risk model — is synthetic and generated inside this repo, with a fixed random seed (numpy.random.default_rng(42)), so the whole pipeline (schema, data, model, tools, tests) is reproducible from a clean clone with python -m src.setup_data.

Architecture

flowchart LR
    U[User question] --> A[MiningOpsAgent loop]
    A -->|tool_use requested| D{Dispatcher}
    D --> T1[warehouse_query_tool<br/>DuckDB]
    D --> T2[credit_risk_tool<br/>LogisticRegression]
    D --> T3[anomaly_check_tool<br/>IsolationForest]
    T1 --> R[tool_result]
    T2 --> R
    T3 --> R
    R --> A
    A -->|end_turn| F[Final text answer]
Loading

The loop lives in src/agent.py (MiningOpsAgent). It sends the user message and the tool schemas to the model; if the response asks for tool_use, it dispatches to the matching Python function, wraps the result (or the exception, if the tool raised one) as a tool_result, and sends it back — up to a bounded number of iterations (default 5) so a misbehaving model can't loop forever.

Two ways to drive it: agent.run(question) is stateless — one question in, one answer out, no memory of anything before or after. agent.chat(question) is the multi-turn version: it appends every turn (including the intermediate tool_use/tool_result exchanges) to agent.history and resends the growing history on each call, so a follow-up like "and what about August?" resolves against the prior answer without the caller re-stating the question. agent.reset() clears it. src/cli.py exposes both: python -m src.cli "question" for one-shot, python -m src.cli (no args) for an interactive multi-turn REPL.

Tools exposed (src/tools/)

All three are plain, synchronous Python functions with a companion JSON schema (TOOL_DEFINITIONS) in the shape the tool-use API expects ({"name", "description", "parameters"}), wrapped into the provider's function-tool envelope in src/tools/__init__.py. Each one runs, and is tested, with no LLM and no API key involved.

Tool File What it does
get_flotation_summary, get_maintenance_alerts, get_procurement_summary warehouse_query_tool.py Fixed, named, parameterized queries against a local DuckDB (data/ops.duckdb) — not arbitrary SQL from the user, by design.
score_credit_risk credit_risk_tool.py Scores a credit applicant profile with a small self-contained LogisticRegression trained on synthetic applicants (data/credit_risk_model.joblib).
check_maintenance_anomalies anomaly_check_tool.py Runs an IsolationForest over recent maintenance events to flag equipment with unusual downtime/severity patterns.

The DuckDB warehouse has three synthetic tables: flotation_batches, maintenance_events, procurement_orders, all generated by src/setup_data.py with numpy.random.default_rng(42).

Techniques used

Technique Where What it's for
OpenAI tool-use API (openai Python SDK, chat.completions.create(tools=...)) src/cli.py Lets the model decide which tool to call and with what arguments, instead of hand-coded intent matching.
Bounded agent loop (MiningOpsAgent.run) src/agent.py Dispatches tool_calls to real Python functions, feeds tool messages back to the model, and caps iterations (default 5) so a stuck model can't loop forever.
DuckDB, parameterized queries src/tools/warehouse_query_tool.py Fast local OLAP queries over the synthetic operations warehouse; parameters are bound (? placeholders), never string-interpolated SQL.
LogisticRegression (scikit-learn) src/tools/credit_risk_tool.py, src/setup_data.py A small, self-contained credit-risk scoring model trained on synthetic applicant data; returns a probability of default and a risk tier.
IsolationForest (scikit-learn) src/tools/anomaly_check_tool.py Unsupervised outlier detection over per-equipment maintenance features (event count, total downtime, share of critical/high-severity events) to flag anomalous equipment without a hand-set threshold.
Oracle-AUC ceiling check (GradientBoostingClassifier vs. the true generating probability) src/model_ceiling_check.py Verifies a weak AUC is the label's own noise, not a fixable model choice, by scoring the true probability behind each synthetic label against the held-out set and checking no model — deployed or more flexible — can beat it.
unittest.mock fake OpenAI client tests/test_agent.py, tests/test_cli.py, tests/test_eval_harness.py Verifies the dispatcher's routing logic (correct tool, correct args, correct tool message shape, malformed-argument handling, iteration cap, multi-turn history growth), the CLI's REPL, and the eval harness's own scoring logic — all without a real API call.
Multi-turn conversation state (MiningOpsAgent.chat) src/agent.py Accumulates turns (including intermediate tool calls) into agent.history so a follow-up question resolves against prior context, instead of every question starting from a blank slate.
Tool-selection + groundedness eval harness src/eval_harness.py A labeled query set (EVAL_QUERIES) checks whether the model calls the right tool(s) for a question, and a number-extraction check (check_groundedness) verifies every number in its final answer traces back to a real tool result rather than a plausible-sounding guess.

Results

Every number and figure below comes from python -m src.generate_report (src/visualization/plots.py) calling the exact same tool functions the agent dispatches at runtime and plotting their real outputs — nothing here is hardcoded or re-simulated separately from the tools. The full pipeline (setup_data.py → generate_report.py → pytest, 42/42 passing, CI-enforced on every push) is reproducible from a clean clone, and the three results below tell three different stories about what happens when you actually check a tool's output instead of trusting the number it returns:

  • score_credit_risk looks weak (AUC 0.586) until you compute the theoretical ceiling and find it's already capturing 96% of the AUC that's actually available — the model isn't underperforming, the label is just noisy by design.
  • check_maintenance_anomalies flags equipment a simple downtime ranking would miss entirely, including one flagged for having too little downtime — evidence the Isolation Forest is using more than one signal, not just proof it runs.
  • get_flotation_summary / get_procurement_summary ground the other two tools in an operational baseline: a real 12-month recovery trend and a real spend breakdown the agent can quote instead of guess.
Tool Metric Value
score_credit_risk ROC-AUC (held-out test, n=400) 0.586
score_credit_risk PR-AUC (base rate 0.245) 0.335
score_credit_risk Test accuracy 0.755
check_maintenance_anomalies Equipment flagged (60d window) 3 / 24
get_flotation_summary Months of data 12

score_credit_risk — weak, and now verified to be near-optimal, not just "honest"

Credit risk evaluation

Three panels, all from the same held-out test set (n=400): the ROC curve (left) hugs the no-skill diagonal — it never gets far above it, which is what a ROC-AUC of 0.586 looks like when you actually plot it instead of just reporting the number. The PR curve (center) spikes near recall=0 (a handful of confident, correct high-probability predictions) and then decays fast toward the 0.245 base rate, which is the realistic floor once recall increases. The histogram (right) makes the reason visible: the predicted-probability distributions for "default" and "no default" overlap almost completely.

That overlap used to just be asserted as "the model is deliberately simple." It's now measured. setup_data.py computes a true_prob_default for every synthetic applicant before flipping the coin that decides their default label — so the true probability behind each label is known, not estimated. Scoring that probability against the real held-out labels (src/model_ceiling_check.py) gives the best AUC any model could ever achieve here, because the only randomness left once you know the true probability is the coin flip itself:

Ceiling comparison

Model Held-out AUC % of ceiling captured
Oracle (true probability, not fit from data) 0.611 100% (by definition)
LogisticRegression (deployed in score_credit_risk) 0.586 96.0%
GradientBoostingClassifier (comparison only, not deployed) 0.607 99.4%

The deployed logistic model is already capturing 96% of the AUC that is theoretically available — and swapping in a considerably more flexible gradient-boosted model only closes the remaining gap to 99.4%, it doesn't blow past the ceiling. That rules out the obvious alternative explanation for a weak AUC (an underpowered model choice): the label itself is this noisy by construction (a Bernoulli draw around a probability that only ranges roughly 0.04–0.88 across applicants), and no model — however expressive — can out-predict a coin flip it doesn't get to see. tests/test_model_ceiling_check.py locks this in: it asserts no model's AUC can exceed the oracle's, and that the deployed model captures at least 85% of it.

Interactive version: predicted P(default) vs. debt-to-income, all 400 held-out applicants, hover for the full profile — opens a live Plotly chart (self-contained HTML, generated by plot_credit_risk_interactive() in src/visualization/plots.py) rather than a static image.

Two real applicant profiles scored with score_credit_risk (called directly, no LLM in the loop):

Profile Age Income (CLP/mo) Debt-to-income Months employed Late payments Requested (CLP) probability_default risk_tier
Low-risk example 42 1,400,000 0.15 96 0 1,500,000 0.0823 low
High-risk example 24 380,000 0.92 3 6 5,500,000 0.7892 critical

Both rows are one real call to score_credit_risk(...) each, run in this session against the model committed at data/credit_risk_model.joblib — not hand-picked to look clean, just two profiles at opposite ends of the input ranges the model was trained on.

check_maintenance_anomalies — downtime isn't the whole story

Anomaly scores by equipment

3 of 24 pieces of equipment get flagged by the IsolationForest, and the flagged ones are not simply the three with the most downtime — that's the interesting part of this chart. EQ-007 has more total downtime (23.2h) than the flagged EQ-009 (22.6h) but isn't flagged; EQ-023 gets flagged at a mid-range 10.4h, well below several normal bars; and EQ-024 is flagged despite having almost no downtime at all (0.5h) — anomalous for being unusually low, not high. That's expected from an Isolation Forest run over multiple maintenance-event features (frequency, severity, downtime) rather than a single downtime threshold, and it's a more realistic signal than "flag whatever has the biggest bar."

get_flotation_summary / get_procurement_summary — the operational backdrop

Warehouse overview animated Warehouse overview

The GIF above animates the flotation recovery trend across the 12 months; the PNG below is the static reference for detailed reading.

Left: monthly average flotation recovery over the synthetic 12-month window — it oscillates in a fairly tight 87–89.5% band, with a visible dip to 86.7% in April 2026 before recovering to a 12-month high of 89.4% in August 2026. Right: procurement spend by category, ranked services > fuel > safety_equipment > reagents > spare_parts — services and fuel alone account for roughly 45% of total procurement spend in this synthetic warehouse, ahead of consumables like reagents and spare parts.

Three more warehouse tool calls, run directly in this session (no LLM involved):

>>> get_flotation_summary("2025-09")
{'n_batches': 18, 'avg_feed_grade_pct': 0.83, 'avg_recovery_pct': 88.87,
 'avg_concentrate_grade_pct': 28.28, 'total_tonnage_processed': 26282.7, 'month': '2025-09'}

>>> get_procurement_summary("delayed")
{'status_filter': 'delayed', 'n_orders': 66, 'total_amount_usd': 514709.88, 'avg_amount_usd': 7798.63}

>>> check_maintenance_anomalies(days=60)
{'window_days': 60, 'n_equipment_evaluated': 24, 'n_equipment_flagged': 3,
 'flagged': [
   {'equipment_id': 'EQ-009', 'n_events': 7, 'total_downtime_hours': 22.53, 'critical_share': 0.286, 'anomaly_score': -0.0365},
   {'equipment_id': 'EQ-024', 'n_events': 2, 'total_downtime_hours': 0.6,  'critical_share': 0.5,   'anomaly_score': -0.0174},
   {'equipment_id': 'EQ-023', 'n_events': 3, 'total_downtime_hours': 10.46, 'critical_share': 0.667, 'anomaly_score': -0.007}]}

Evaluation harness: does the agent pick the right tool, and does it stay grounded?

The system prompt says "never invent numbers that a tool could return" — src/eval_harness.py turns that from an instruction into something checkable. It runs a labeled set of queries (EVAL_QUERIES, including a multi-tool query and an out-of-scope one where no tool should fire) through the real agent and real tools, and scores two things per query:

  • Tool-selection accuracy — did the model call exactly the tool(s) the query actually needs?
  • Groundedness — does every number in the final answer trace back to a real tool result (check_groundedness), or did the model state something that no tool call actually returned?
OPENAI_API_KEY=sk-... python -m src.eval_harness

Honest limitation: no OPENAI_API_KEY was available on the machine this was built on, so this repo does not claim a live accuracy number — that would mean reporting a metric nobody actually measured, exactly what the rest of this README argues against. What is verified, in tests/test_eval_harness.py, is that the scoring logic itself is correct: extract_numbers/check_groundedness are tested against hand-built cases (a grounded answer, a fabricated number, no tool calls at all), and run_eval is tested against scripted mock responses for both a correct and an incorrect tool selection. Run the harness with a real key to get real numbers.

Installation

python -m venv .venv
.venv\Scripts\activate       # Windows PowerShell: .venv\Scripts\Activate.ps1
pip install -r requirements.txt

Usage

  1. Generate the synthetic data and train the risk model:

    python -m src.setup_data

    Writes data/ops.duckdb and data/credit_risk_model.joblib.

  2. Generate the result figures and metrics summary:

    python -m src.generate_report

    Writes reports/figures/*.png, reports/figures/*.gif, reports/metrics.json, and outputs/interactive/credit_risk_scores.html.

  3. Run the test suite (works fully offline, no API key needed):

    pytest
  4. Talk to the agent (requires a real OPENAI_API_KEY):

    # One-shot, stateless:
    python -m src.cli "What was the flotation recovery in September 2025?"
    
    # Interactive, multi-turn (remembers context across questions until you type 'reset'):
    python -m src.cli

    Without a key set, either mode exits with a clear message instead of a traceback.

  5. Run the evaluation harness (also requires a real OPENAI_API_KEY — see Evaluation harness above):

    python -m src.eval_harness

Design decision: versioning the generated data

data/ops.duckdb, data/credit_risk_model.joblib, reports/figures/*.png + reports/metrics.json, and outputs/interactive/credit_risk_scores.html are committed to the repo rather than gitignored. They are small, fully synthetic/derived, deterministically regenerable (python -m src.setup_data && python -m src.generate_report), and versioning them means the tools (and the results above) are visible immediately after a clone without a mandatory setup step. .gitignore still excludes the virtual environment and caches.

Author

Pablo Reyes — Data Scientist, Santiago, Chile.

About

Agente con tool-calling real (API de OpenAI) sobre operaciones mineras: consultas SQL a DuckDB, scoring de riesgo crediticio y detección de anomalías — responde llamando herramientas, no alucinando números.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages