Skip to content

FEAT Add role-aware Scenario target-attempt accounting #3043

Description

@romanlutz

Is your feature request related to a problem? Please describe.

Current main records producer roles but does not use them for target-attempt
rollups. The shared analytics helper, progress, and history SQL count attempts
without a producer-role predicate. Thus their existing
logical-unit/raw-history counts must not be presented as target-attempt counts.

For example, two target-facing child results plus one Sequential envelope
are three historical result rows, but only two target-facing producer attempts.
The envelope can still complete one outer planned logical unit. A preparation
error may also have a target-facing role, so this role alone does not prove a
target request occurred. This check inspected code; it did not run an attack.

Describe the solution you'd like

Add a consumed typed target-facing accounting projection using the existing
AttackResultMetadata contract and shared Scenario analytics/identity policy.
Keep planned logical progress, raw historical attempts, and target-facing
attempts distinct. Memory owns database aggregation; backend/SDK/report
consumers present the shared policy rather than independently infer roles.

Acceptance criteria:

  • Orchestration envelopes remain observable but contribute zero to target-facing
    attempts and their target-facing error/retry totals.
  • Preserve exact saved-plan denominators, outer-unit completion, existing
    latest-attempt identity/order, raw history, and FEAT: Improve adversarial benchmark dataset and scoring #2551/FIX Share scenario success statistics across SDK, backend and reports #2820 cache inclusion.
    Add the new rollup without silently redefining existing unit/raw fields.
  • Legacy/malformed roles remain explicitly unknown using conservative decoding;
    never infer a role from a conversation ID, class name, display text, or index.
    Keep orchestration and unknown-role totals visible alongside target-facing
    totals so the producer categories reconcile with raw history.
  • Report unmatched attempts separately, with the same matching policy across
    progress, detail, SDK/report consumers, and history.
  • Real-domain/SQLite parity cases cover ordinary/baseline, flat and nested
    Sequential/Adaptive, retry/resume, recovered and unrecovered errors,
    preparation failure, negative retries, unknown roles, invalid plans,
    tied timestamps, and cache hits.
  • SQLite runtime and Azure SQL compilation cover equivalent database-side
    role filtering. Do not add full-result hydration, per-attempt queries, or
    Python grouping as a replacement for the new SQL aggregate.
  • Field names and documentation say target-facing producer attempts, not
    confirmed target invocations; no new request-tracing feature is required.

Dependencies: merged #2997 and #2820; coordinate indexed analytics #2792 and
history-query extractions #2768/#2769. The separate child-link fix #3039 and
dataset/technique UX are not hard dependencies for counting persisted roles.

Describe alternatives you've considered, if relevant

Do not redo shared success statistics, change logical completion to child
counts, or remove historical rows. Exclude schema/duplicate hierarchy stores,
child-link repair, hierarchy endpoints/UI, dataset/estimate changes, request
tracing, scorer semantics, and new attack algorithms.

Additional context

Roadmap: scope 8.
Design: original result/accounting references.
Evidence: upstream 58a93534838ed58e766276c060c4f0d9149d6483,
pyrit/analytics/scenario_statistics.py,
pyrit/backend/services/scenario_progress_read_model.py, and
pyrit/memory/memory_interface.py.
Follow doc/code/framework.md: models own types, analytics owns counting
policy, memory persists/aggregates, and backend/output present it.
No GUI change or unrelated screenshots are requested.

Activity

  1. Nimit3418 commented on Oct 8, 2026

    @Nimit3418

    Hi Roman Lutz (@romanlutz) , I would love to take this up!

    I've reviewed the issue description and the architectural guidelines from doc/code/framework.md.

    My plan is to introduce the producer-role predicate into the target-attempt rollups across the specified components, strictly following the separation of concerns:

    1. pyrit/memory/memory_interface.py — Update the aggregation logic to filter/group by the producer role.
    2. pyrit/analytics/scenario_statistics.py — Adjust the counting policy to use the new role-aware predicate.
    3. pyrit/backend/services/scenario_progress_read_model.py — Ensure the backend properly surfaces this updated accounting.

    I will also review the upstream reference (58a93534...) for context and ensure no GUI changes are included, keeping it strictly logic-focused as requested.

    Could you please assign this to me? I have my local environment ready and can start immediately!

  2. romanlutz commented on Oct 9, 2026

    @romanlutz
    ContributorAuthor

    Nimit Jain (@Nimit3418) thanks for picking this up! I see you're already assigned. Your component split makes sense.

    One coordination point before you start: #3059 now contains the shared OutcomeStatistics / compute_outcome_statistics calculation, including both success_rate_decided and success_rate_all for attack and Scenario analytics. It is still unmerged. Please reuse that shared calculation and coordinate overlapping changes with that PR, rather than building another success-rate implementation here. Sharing the calculation does not make raw saved-result counts interchangeable with Scenario's latest-attempt-per-logical-unit counts.

    For this issue, please add separate role-aware rollups, not filter or redefine the existing counts. Preserve planned logical progress, latest-attempt selection, saved-plan denominators, cache inclusion, and raw history. Memory should aggregate the role categories in SQL, analytics should own the selection/counting policy, and backend/SDK/report consumers should present that same policy. Use AttackResultMetadata as the role contract and keep unknown/malformed roles visible rather than guessing or discarding them.

    For example, two target-facing child results plus one Sequential envelope are 2 target-facing producer attempts and 1 orchestration row, but can still complete 1 outer planned unit. A target-facing preparation failure still belongs to that producer category; it is not evidence of an actual target call.

    Please cover those distinctions with the issue's real SQLite parity cases and Azure SQL query-compilation checks. This work does not need to wait for dataset UX or add hierarchy endpoints.

  3. Nimit3418 commented on Oct 9, 2026

    @Nimit3418

    Hi Roman Lutz (@romanlutz), thanks for the guidance and the heads-up!

    I will take a look at PR #3059 right now. I'll update my PR to reuse their OutcomeStatistics / compute_outcome_statistics calculation so the math is fully shared and everything coordinates perfectly. I will push the updates shortly!

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions