Skip to content
View retinapeg's full-sized avatar

Block or report retinapeg

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
retinapeg/README.md

Leo Aarons-Ditson

UCL-trained physicist (BSc, PGCert with Distinction) · London · CV

I build agentic AI systems and then test whether they can be trusted. The pattern across my recent work: build the system, evaluate it, read the traces when the headline score looks too good, diagnose what failed, add controls, and measure whether the controls actually help. Negative results stay in the write-up.

Featured work

The numbers below are the committed counts; each linked README has an evidence table pointing at the files that hold them.

agentic-physics-bench · completed, frozen

Does an offered tool change how a model solves a problem, and can the trace prove that the system scored was the system declared? Correctness hit the ceiling in every condition (54/54 × 3), so the score said nothing. A per-call trace audit during development found that a server-side advisor that consults a second model had been active in 31 of the 48 unique calls across the two development stages run before it was disabled, all made with the CLI's tools switched off. Those two stages were voided, the pathway was disabled, per-call model-identity checks were added, and all 324 held-out calls came back clean. Correct outputs alone did not prove that the declared system was the one being evaluated.

agent_reliability_lab · first run complete

Does an independent model review catch coding defects that deterministic tests miss, and what does that cost? On 12 tasks, every generated solution passed its visible tests and 2 still failed hidden tests. A separate reviewer flagged both, plus one flag the hidden tests could not confirm; one bounded revision per flag fixed one of the two confirmed defects, and the other returned byte-identical code. Review took 1.36× the coding time. In this one run the oversight caught what the tests missed and also had its own cost and failure modes, which is why it is measured rather than assumed.

agent-failure-analysis · v0.1.1 release candidate

Give it a recorded agent run plus the task and success criteria; get back an evidence-linked account of what happened, what is established, what is still a hypothesis, what evidence is missing and what regression test to add. Deterministic code verifies that every cited excerpt exists in the trace; the model interprets; the checker says in print that it does not verify the interpretation. Evaluated on synthetic traces only, and labelled as such.

institutional-workbench · prototype, negative result kept

Claude and Codex take a small repo change from assessment to tested patch; does giving each model an expert role make it a better specialist? In a predeclared 96-call experiment, 33 calls failed on timeouts, provider errors or malformed output, all on the Claude Code CLI side (Codex 0/48), and 14 of the 16 malformed answers were JSON wrapped in Markdown fences. Almost all of the role prompts' apparent gain came from fewer format and completion failures; on schema-valid answers the three conditions were within 0.06 of each other, so no reasoning benefit was established. Protocol robustness dominated specialisation.

Also

  • agent-context-router: bounded context loading for agent sessions with hash-checked edits. The keyword picker lost to plain BM25 (21/31 vs 24/31), and the README says so.
  • careerops-ai: a durable, graph-controlled CV workflow where Python owns state and transitions, models return bounded structured proposals, and independent red/blue reviews feed a bounded revision loop. No autonomous submission.
  • agent-workflow-orchestrator: Codex and Claude build the same task and review each other's diffs; one recorded contest in which a review caught a Unicode bug.
  • institutional-ai: AI specialists report independently, challenge each other and keep dissent on the record. Frozen prototype with scripted agent replies; no model is called.
  • fleetcast: NYC taxi-pickup forecasting; boosted trees cut MAE by 24% against persistence on a held-out fortnight.

Earlier and smaller: dronewatch · YLOOKUP · before-coffee · support-triage-agent · canton-collateral-optimizer · uk-orbit-guard · schrodinger-harmonic-demo

How I work

Freeze the protocol before the scored run. Keep invalid episodes in the denominator, and keep voided runs on record. Treat each call's trace, not the aggregate score, as the primary evidence. Say what a result does not show.

Pinned Loading

  1. agentic-physics-bench agentic-physics-bench Public

    Does Claude use a tool it's offered? Answers hit the ceiling either way; a trace audit found a hidden second model.

    Python

  2. agent-workflow-orchestrator agent-workflow-orchestrator Public

    Codex and Claude build the same task, review each other's diffs, and plain code decides what passes.

    Python

  3. YLOOKUP YLOOKUP Public

    Capital-call checks: a model may read the notice, code does the arithmetic, a person clears every break.

    Python

  4. support-triage-agent support-triage-agent Public

    Support-ticket agent that resolves only with evidence and otherwise escalates. Offline mock demo.

    Python

  5. agent-context-router agent-context-router Public

    Hands a fresh agent session only the notes it needs, with hash-checked edits. A keyword picker lost to BM25.

    Python

  6. fleetcast fleetcast Public

    Forecasting NYC taxi pickups 30 minutes ahead: boosted trees cut MAE 24% vs persistence on held-out data.

    Python