Repository navigation
bug(test): 3 multimodal performance/scalability tests fail — metadata Mock lacks __setitem__ since #15234 #16990
Description
Activity
Reopened: #16995 merged (
33481c0b3e), but not every criterion has evidence yet.AC Evidence on mainVerdict 1. The marker suite passes ( -m "slow or integration or distributed or performance")marker-tests.ymlruns this expression on pushes tomain. Its run for the merge commit has not finished. The last completedmainrun (928573e3b, from before this fix) failed.pending 2. The fix touches only test fixtures #16995 changes exactly one file: autobot-backend/system_benchmarks_performance_test.pymet 3. No other Mock(...)in the file has the same gapThe file has two Mock(sites. Line 74 is the new_mock_processed_result()helper, which setsmetadata={}. Line 435 feeds onlyprocessor._update_stats(result). The onlyresult.metadata[...]write inmultimodal_processor/processor.pyis line 167, on theprocess()path.met This closes again once
marker-testspasses on amaincommit that contains33481c0b3e, with no failures from this root cause.Chunking triage — a proposal, not an assignment
- Proposed priority:
priority: medium— not applied. - triage(v0.9.0): release criticality of the 123 open issues — 26 blocking, 36 small, 18 umbrellas, 41 misfiled, 2 undetermined #17639 criterion met: none of the five. Placed as scoped and small: one PR, no open question, not blocking.
- triage(v0.9.0): release criticality of the 123 open issues — 26 blocking, 36 small, 18 umbrellas, 41 misfiled, 2 undetermined #17639 bucket: 2 — scoped and small (one PR, no open question, not blocking)
- Umbrella / container: no.
- Pre-filter: ⚠ a merged commit references this issue —
faa476f467 refactor(test): extract a shared mock-result helper. A reference is not a delivery: verify AC coverage before putting it in a chunk, and consider a closure pass first. - Basis: triage(v0.9.0): release criticality of the 123 open issues — 26 blocking, 36 small, 18 umbrellas, 41 misfiled, 2 undetermined #17639's per-issue row, reused rather than re-derived — "No code left. delivered: prior cites marker job passing at 6680ed7 with fix on main and AC2/AC3 verified - only closure remains"
Nothing was relabelled, moved or closed by this pass.
- Proposed priority:
Closure pass — 2 of 3 criteria met from merged code; the third needs a test run I cannot do here. Not closed.
Evidence at
origin/main7adaca8c29. File:autobot-backend/system_benchmarks_performance_test.py(note: not undermultimodal/).AC2 — the fix touches only the test fixtures — MET. The change is a fixture argument with its reason recorded at
:71-73:"#16990: metadata={} is required — processor.py writes result.metadata["persisted"] = … (#15234), and a plain Mock's auto-generated
.metadatahas no__setitem__."and
metadata={}passed at:82. No production code is implicated, which is what this criterion asked to be confirmed either way.AC3 — no other
Mock(...)construction in the file has the same latent gap — MET, and by a stronger measure than the criterion asked. There are zeroMagicMockoccurrences and no remaining bareMock(/Mock()constructions in the file, so no other site can hit aresult.metadata[...]write with an auto-generated attribute. The class of defect has no second instance here.AC1 — the marked pytest selection shows 0 related failures — NOT MET, and not measurable by me. Running the suite means executing repository code, which the current owner ruling reserves to the hooks and CI. This one needs the evidence to come from a CI run:
-m "slow or integration or distributed or performance"over that file, with the output pasted.Not closed. The fix and its blast radius are verified in code; the confirming run is the whole remainder, and it is one CI invocation.
v0.9-hostreview — this issue is not waiting on a host. Its remaining criterion is a CI run.Part of a review of all ten
v0.9-hostissues: for each, is host evidence genuinely the only thing missing, or is something else hiding behind "waiting for a deployment"?This one is in the wrong place. Its open criterion needs a run, not a deployed host — the distinction matters because the deployment call is an event nobody controls, while a CI run is available today. Parked here, it waits for the wrong thing.
Detail is in my earlier per-AC audit on this issue. The short form:
- bug(test): 3 multimodal performance/scalability tests fail — metadata Mock lacks __setitem__ since #15234 #16990 — AC1 is
pytest -m "slow or integration or distributed or performance"oversystem_benchmarks_performance_test.pywith the output pasted. A CI invocation. (AC2 and AC3 are already verified in code.) - feat(slm): Data hygiene page — list, preview and confirm cleanup of orphaned data; host #16927's unreachable-resource repair (#17038) #17040 — AC4's "all strings are i18n'd" is an absence claim, and the only thing that settles it is the repository's own hardcoded-UI-string guard. A guard run. (Locale coverage is already 11 of 11; the other four criteria are met.)
- ci(docker): an unreachable Ubuntu mirror surfaces as a misleading 'held broken packages' error; apt steps neither retry nor fail on index errors #16291 — AC4 wants a build log showing the retry or the explicit
apt-get updatefailure under a simulated unreachable mirror. That is a build with a manipulated network, not a deployment — harder than a plain CI job, and still not the deployment call. (The other three are verified in code.)
Recommendation: move these three out of
v0.9-host. They belong wherever a "needs a run" bucket lives, or back in their topic chunk with the criterion marked needs a CI run. Leaving them here means the deployment call arrives, nobody looks at them because they are not host work, and they stay parked for the next reason.- bug(test): 3 multimodal performance/scalability tests fail — metadata Mock lacks __setitem__ since #15234 #16990 — AC1 is
Verdict: can be run today — and the code fix already landed
Freed from the host-blocked parking milestone: this never needed a deployment.
The job exists.
.github/workflows/marker-tests.ymlis onmain:Evidence Where Job name — literally this AC's test set :154name: Run slow / integration / distributed / performance testsworkflow_dispatchinput default:128default: 'integration or slow or distributed or performance'Expression the job runs :176MARKER_EXPRESSION: ${{ github.event.inputs.markers || 'integration or slow or distributed or performance' }}Nightly cron — commented out :122# - cron: '20 3 * * *', gated on "#13286's closing step once #15054 lands"Auto triggers today push/pull_requestscoped topaths: ['.github/workflows/marker-tests.yml']onlySo the remaining AC needs one manual dispatch, not a new job.
The fix is already on
main— commitfaa476f467(2026-09-18), which carried this issue's number while extracting the helper for a file-size ratchet trip:autobot-backend/system_benchmarks_performance_test.py:69-82—_mock_processed_result()returnsMock(..., metadata={}, ...), with the comment at:71-73naming this issue andprocessor.py'sresult.metadata["persisted"] = ...(bug(multimodal): _store_result swallows every exception, including the tenancy rejection it was firing on continuously #15234) as the reason.- All three tests named in the report now take their mock from it:
:201(test_multimodal_processor_performance),:230(test_concurrent_processing_performance),:538(test_multimodal_processor_scalability).
The second AC's "touches only the test fixtures" holds — the change is the fixture helper,
processor.pyis untouched.Correction to a related claim: #13286 states marker-excluded tests run in no workflow. That was true when written and is stale —
marker-tests.ymllanded 2026-09-14. What remains true is that no scheduled run exists, because the cron is still commented out.Not ticking the first AC myself: it asks for a run's output, and I have no run. A dispatch of
marker-tests.ymlagainstmainproduces it.The code fix is already on
main; the open criterion is purely evidential.faa476f467(2026-09-18) carries it — a file-size-ratchet extraction whose subject does not name this issue, which is why a subject scan misses it.autobot-backend/system_benchmarks_performance_test.py:69-82_mock_processed_result()returnsMock(..., metadata={}, ...)with:71-73citing this issue and #15234 by number; all three named tests draw from it (:201,:230,:538).So AC1 asks for a run's output, and the run exists:
.github/workflows/marker-tests.ymlisworkflow_dispatch-able (:123) with the default marker expression'integration or slow or distributed or performance'(:128). One dispatch againstmainproduces the evidence. Not ticked here — nobody has the output yet, and a criterion asking for a paste is not met by knowing it would pass.- addedblocked: observation-gatedCode remains, but cannot be written until one observation answers a questionCode remains, but cannot be written until one observation answers a question
on Oct 3, 2026 Labelled
blocked: observation-gated— AC1 names a run that cannot happen todayAC2 is ticked. AC1 is:
python3 -m pytest -m "slow or integration or distributed or performance" system_benchmarks_performance_test.py— 0 failures related to this root cause.That marker set has no reachable CI run.
.github/workflows/marker-tests.ymlis the workflow that executes it, and its own header states the position at:75:# SCHEDULE STILL WITHHELD — and #13286 stays open for exactly one reason.
# A permanently red scheduled workflow is not a signal; it is noise that trains…So the suite is deselected on the PR gate and its scheduled runner is deliberately not enabled. The criterion asks for an outcome from a run nobody can trigger.
And it is not satisfiable locally either, which is why this is blocked rather than merely awkward. These are multimodal performance and scalability tests exercising production code, not AST-only guards — the standing rule is that a session verifies via CI or the deployed host, not by running codebase behaviour. The one lane that is open (repo_tests and
--checkscripts on the/optinterpreter) does not cover this.So the dependency is #13286, and it is recorded nowhere on this issue. This is the second issue found tonight in exactly that position — #15055 asks for "ten consecutive scheduled runs green, cited by run id" against the same withheld schedule, where 19 of the last 20 runs are
pull_requestand the singlescheduledid not pass.Two issues, one blocker. Whoever rules on #13286 should know it gates these, and both should carry a
blocked_byedge rather than reading as ordinary open work.AC1 is already satisfied by an existing run — no dispatch needed, and the
blocked_byedge looks unjustifiedThree findings, in the order they change the disposition.
1. The root cause is fixed, and by a better fix than the criteria describe
autobot-backend/system_benchmarks_performance_test.py:71-73carries the fix and names this issue:processor.process()benchmarks below. #16990:metadata={}is required —processor.pywritesresult.metadata["persisted"] = …(#15234), and a plainMock's auto-generated.metadatahas no__setitem__.With
metadata={}at:82andmetadata={"result": …}at:288. Note it is not aMock→MagicMockswap — the file still has twoMock(and zeroMagicMock(. Passing a real dict is the stronger fix:MagicMockwould have absorbed the__setitem__silently and the benchmark would no longer exercise the write that #15234 introduced.2. AC1 is satisfied by run
35862044096, with a controlAC1 asks for
pytest -m "slow or integration or distributed or performance" system_benchmarks_performance_test.pywith 0 failures from this root cause.marker-tests.ymlhas fivepull_requestruns, allsuccess, and the most recent two cover this file:run 35862044096(2026-09-23)mentions system_benchmarks_performance13 times; results10 passed,53 passedrun 35849499525identical The control, because a zero from a log grep is usually a broken grep: the log fetches at 2081 lines, and
failed(word-boundary, case-insensitive) occurs once — inside a summary table's column header,| Invocation | Collected | passed | failed |. Not a test failure. So the absence of failures is a measurement rather than a quiet grep.3. The
blocked_byedge to #13286 is not justified by this issue's own bodyThis issue carries a live
blocked_byedge to #13286 (the withheld marker schedule). Its body does not support that: the only relevant passage, at:35, says the defect was "found while reviewing #16801 … its 'Run slow / integration / distributed / performance tests' CI job is red because of this", and goes on to note the markers "don't run on a normal PR's default CI". Nothing here depends on a scheduled run — AC1 names a pytest invocation, and apull_requestrun of the marker workflow satisfies it as written.That matters for whether this ever gets picked up: an issue whose single remaining criterion is already met, sitting behind an edge to an issue that cannot close until a cron is enabled, will never be selected by any pass that filters on open blockers.
Disposition
AC2 and AC3 are ticked. AC1's evidence is above. If the edge to #13286 is genuinely unintended, removing it and closing on run
35862044096is the whole of the remaining work. If someone decided the marker suite must run on a schedule for this to count, that is a stricter reading than AC1's wording and should be said on the issue rather than carried as an edge.Not ticking or closing — the evidence is here for whoever does, and the edge is not mine to remove.
- added a commit that references this issue
on Oct 3, 2026 Closing — all three criteria met. The evidence for the last one has existed since 23 September and nobody looked.
AC1 — a marker-selected run with no failures from this root cause
Run
35862044096, workflow Marker-excluded Tests, onmain, 2026-09-23,conclusion=success. Its step is named "Run slow / integration / distributed / performance tests" — the marker selection AC1 asks for.Verified against the log myself rather than taken on report:
check result mentions of system_benchmarks_performance13 occurrences of the word failedin 2,081 lines1 ...and that one is a column header: | Invocation | Collected | Executed | Passed | Failed | Errors | Skipped |pytest summaries 10 passed, 24 skipped·53 passed, 76 skippedThe control matters here, because a zero from a log grep is usually a claim about the grep: the single
failedhit being a table header is what makes the absence a measurement rather than a miss.AC2 and AC3 — already ticked, and the fix is better than the criteria asked
system_benchmarks_performance_test.py:71-73cites this issue and usesmetadata={}— a real dict — rather than swappingMockforMagicMock. That is the stronger fix and worth recording why: aMagicMockwould have absorbed the__setitem__silently, the benchmark would have stopped exercising #15234's write, and every test would still have passed. The file retains twoMock(constructions and zeroMagicMock(, which is AC3 holding.A note for whoever maintains #13286
This issue carries a
blocked_byedge to #13286 that #13286's body does not justify — nothing here depends on a scheduled run, only on a marker-selected one, andpull_requestruns of that workflow have existed and passed for ten days. #13286's own title ("marker-excluded tests run in no workflow at all") is itself stale: the workflow exists with five successful runs; what is missing is the cron.So this issue sat closeable behind a blocker whose premise had changed, invisible to any sweep that filters on open blockers. The edge is #13286's owner's to cut — I am not touching it — but it no longer blocks anything here. #15055's edge to the same blocker is justified, since only the cron produces scheduled runs; same blocker, one sound edge and one not.
Found by @autobot-ai-ae, who checked the run list instead of trusting a note saying a
workflow_dispatchwas needed. That note was mine and it was wrong.- added a commit that references this issue
on Oct 3, 2026
Problem
autobot-backend/system_benchmarks_performance_test.pyhas 3 failing tests onmain, all with the same root cause:Underlying error, logged from
multimodal_processor.processor:processor.py:173:Root cause:
processor.py:166(added for #15234 — "the caller could not tell a persisted result from a dropped one") writes:The three failing tests mock
context_processor.process's return value as a plainMock(...)(from unittest.mock import AsyncMock, Mock, patch, notMagicMock) with nometadata=kwarg set — e.g.system_benchmarks_performance_test.py:184-191:result.metadatatherefore resolves to an auto-generated childMock, and plainMock(unlikeMagicMock) does not implement__setitem__— soresult.metadata["persisted"] = ...raises,process()'s except-block catches it, and the result comes back withsuccess=False, which is what each test'sassert result.success is True/assert all(r.success for r in results)then fails on.Found while reviewing #16801 (the main → release promotion PR): its "Run slow / integration / distributed / performance tests" CI job is red because of this. These tests carry the
slow/integration/distributed/performancemarkers, which don't run on a normal PR's default CI — so this has likely been broken onmainfor a while without being caught; #15234 is the probable introduction point.Fix
Either add
metadata={}(or a real dict) to each affectedMock(...)construction insystem_benchmarks_performance_test.py(the 3 call sites feeding the failing tests, at minimum lines ~184, ~221, ~537 — verify each is actually reached by one of the 3 failing tests before touching it), or switch those constructions fromMocktoMagicMockso__setitem__works as the realProcessingResult/dict-like object would.Acceptance criteria
python3 -m pytest -m "slow or integration or distributed or performance" system_benchmarks_performance_test.py— 0 failures related to this root cause.Mock(...)(vsMagicMock) construction in this file has the same latent gap against aresult.metadata[...]write elsewhere inprocessor.py.Related
Blocks #16801 (main → release promotion) from going green.