Repository navigation
FEAT Add role-aware Scenario target-attempt accounting #3043
Description
Activity
- addedhelp wantedExtra attention is neededExtra attention is needed
on Oct 8, 2026 Hi Roman Lutz (@romanlutz) , I would love to take this up!
I've reviewed the issue description and the architectural guidelines from
doc/code/framework.md.My plan is to introduce the
producer-rolepredicate into the target-attempt rollups across the specified components, strictly following the separation of concerns:pyrit/memory/memory_interface.py— Update the aggregation logic to filter/group by the producer role.pyrit/analytics/scenario_statistics.py— Adjust the counting policy to use the new role-aware predicate.pyrit/backend/services/scenario_progress_read_model.py— Ensure the backend properly surfaces this updated accounting.
I will also review the upstream reference (
58a93534...) for context and ensure no GUI changes are included, keeping it strictly logic-focused as requested.Could you please assign this to me? I have my local environment ready and can start immediately!
romanlutz commented
on Oct 9, 2026 ContributorAuthorMore actionsNimit Jain (@Nimit3418) thanks for picking this up! I see you're already assigned. Your component split makes sense.
One coordination point before you start: #3059 now contains the shared
OutcomeStatistics/compute_outcome_statisticscalculation, including bothsuccess_rate_decidedandsuccess_rate_allfor attack and Scenario analytics. It is still unmerged. Please reuse that shared calculation and coordinate overlapping changes with that PR, rather than building another success-rate implementation here. Sharing the calculation does not make raw saved-result counts interchangeable with Scenario's latest-attempt-per-logical-unit counts.For this issue, please add separate role-aware rollups, not filter or redefine the existing counts. Preserve planned logical progress, latest-attempt selection, saved-plan denominators, cache inclusion, and raw history. Memory should aggregate the role categories in SQL, analytics should own the selection/counting policy, and backend/SDK/report consumers should present that same policy. Use
AttackResultMetadataas the role contract and keep unknown/malformed roles visible rather than guessing or discarding them.For example, two target-facing child results plus one Sequential envelope are 2 target-facing producer attempts and 1 orchestration row, but can still complete 1 outer planned unit. A target-facing preparation failure still belongs to that producer category; it is not evidence of an actual target call.
Please cover those distinctions with the issue's real SQLite parity cases and Azure SQL query-compilation checks. This work does not need to wait for dataset UX or add hierarchy endpoints.
Hi Roman Lutz (@romanlutz), thanks for the guidance and the heads-up!
I will take a look at PR #3059 right now. I'll update my PR to reuse their
OutcomeStatistics/compute_outcome_statisticscalculation so the math is fully shared and everything coordinates perfectly. I will push the updates shortly!
Is your feature request related to a problem? Please describe.
Current main records producer roles but does not use them for target-attempt
rollups. The shared analytics helper, progress, and history SQL count attempts
without a producer-role predicate. Thus their existing
logical-unit/raw-history counts must not be presented as target-attempt counts.
For example, two target-facing child results plus one Sequential envelope
are three historical result rows, but only two target-facing producer attempts.
The envelope can still complete one outer planned logical unit. A preparation
error may also have a target-facing role, so this role alone does not prove a
target request occurred. This check inspected code; it did not run an attack.
Describe the solution you'd like
Add a consumed typed target-facing accounting projection using the existing
AttackResultMetadatacontract and shared Scenario analytics/identity policy.Keep planned logical progress, raw historical attempts, and target-facing
attempts distinct. Memory owns database aggregation; backend/SDK/report
consumers present the shared policy rather than independently infer roles.
Acceptance criteria:
attempts and their target-facing error/retry totals.
latest-attempt identity/order, raw history, and FEAT: Improve adversarial benchmark dataset and scoring #2551/FIX Share scenario success statistics across SDK, backend and reports #2820 cache inclusion.
Add the new rollup without silently redefining existing unit/raw fields.
never infer a role from a conversation ID, class name, display text, or index.
Keep orchestration and unknown-role totals visible alongside target-facing
totals so the producer categories reconcile with raw history.
progress, detail, SDK/report consumers, and history.
Sequential/Adaptive, retry/resume, recovered and unrecovered errors,
preparation failure, negative retries, unknown roles, invalid plans,
tied timestamps, and cache hits.
role filtering. Do not add full-result hydration, per-attempt queries, or
Python grouping as a replacement for the new SQL aggregate.
confirmed target invocations; no new request-tracing feature is required.
Dependencies: merged #2997 and #2820; coordinate indexed analytics #2792 and
history-query extractions #2768/#2769. The separate child-link fix #3039 and
dataset/technique UX are not hard dependencies for counting persisted roles.
Describe alternatives you've considered, if relevant
Do not redo shared success statistics, change logical completion to child
counts, or remove historical rows. Exclude schema/duplicate hierarchy stores,
child-link repair, hierarchy endpoints/UI, dataset/estimate changes, request
tracing, scorer semantics, and new attack algorithms.
Additional context
Roadmap: scope 8.
Design: original result/accounting references.
Evidence: upstream
58a93534838ed58e766276c060c4f0d9149d6483,pyrit/analytics/scenario_statistics.py,pyrit/backend/services/scenario_progress_read_model.py, andpyrit/memory/memory_interface.py.Follow
doc/code/framework.md: models own types, analytics owns countingpolicy, memory persists/aggregates, and backend/output present it.
No GUI change or unrelated screenshots are requested.