Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion .claude-plugin/plugin.json
Original file line number Diff line number Diff line change
@@ -1,7 +1,7 @@
{
"name": "fdeops",
"description": "Skills for forward deployed engineers across strategy, architecture and engineering. Use individual tasks or @fde coordination; local customer memory supports continuity.",
"version": "5.2.2",
"version": "5.2.3",
"category": "productivity",
"tags": [
"community-managed"
Expand Down
6 changes: 6 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
@@ -1,5 +1,11 @@
# Changelog

## 5.2.3 - 2026-10-09

- Connect observed AI failures to specific checks using existing evaluation reports.
- Distinguish targeted test coverage from production failure-rate estimates.
- Check code evaluators against passing, failing and boundary cases before relying on their results.

## 5.2.2 - 2026-10-08

- Add task-focused UI guidance for existing design systems, interaction states, accessibility and browser verification.
Expand Down
2 changes: 1 addition & 1 deletion mcp/fdeops-ingest/package.json
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
{
"name": "fdeops-ingest-mcp",
"version": "5.2.2",
"version": "5.2.3",
"private": true,
"description": "Thin stdio MCP sink for FDEOps ingest (stage \u2192 propose \u2192 apply). Zero runtime dependencies.",
"bin": {
Expand Down
2 changes: 1 addition & 1 deletion package.json
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
{
"name": "fdeops",
"version": "5.2.2",
"version": "5.2.3",
"description": "Skills for forward deployed engineers across strategy, architecture and engineering. Use individual tasks or @fde coordination; local customer memory supports continuity.",
"bin": {
"fdeops": "bin/install.js",
Expand Down
2 changes: 1 addition & 1 deletion plugin.json
Original file line number Diff line number Diff line change
@@ -1,7 +1,7 @@
{
"$schema": "https://agent-plugins.org/schemas/1.0.0/plugin.schema.json",
"name": "fdeops",
"version": "5.2.2",
"version": "5.2.3",
"description": "Skills for forward deployed engineers across strategy, architecture and engineering. Use individual tasks or @fde coordination; local customer memory supports continuity.",
"author": {
"name": "Subash Natarajan",
Expand Down
2 changes: 1 addition & 1 deletion skills/build/.fde-generated.json
Original file line number Diff line number Diff line change
Expand Up @@ -6,7 +6,7 @@
"agents/openai.yaml": "abd0a33ab4efa25ea038932d37eae2faa2527baa96ebae8732ea90f1ebc3865e",
"references/build.md": "daaa7be30ce6e4c42715dc708f877ab1d2270df0dffdcc33680085450b9f0b2f",
"references/debug.md": "bdc09abdd7a075574fc2daf91cecfd2858e67d75a964146a118bf782328192f6",
"references/eval-pack.md": "f883ea0be4777c9c3893db0a616b473f172140d5a3a1be4c449570c8608e4d5b",
"references/eval-pack.md": "40ca29fbee82ac1846d3a31ddf92ef20c6d7906ade53571decf81d2f79ee936c",
"references/integrate.md": "1cb7a60d7545b0bf224fce678a04ce6ccdf368c47877d9c8e4dc4272bb0d5b0c",
"references/qa.md": "41de4d70827c83291efa217e97d777f62ec2849827687fbba7e4b1d17484b87e",
"references/review.md": "36b77ed44989a125a36e9948ce19c2ced57d529bee06643e32dda96f896d9d5b",
Expand Down
6 changes: 3 additions & 3 deletions skills/build/references/eval-pack.md
Original file line number Diff line number Diff line change
Expand Up @@ -6,9 +6,9 @@ Use [task context](task-context.md). Supplied permitted context and an evaluatio

## Method

1. **Define the evaluated surface.** Name the model judgment, inputs, outputs, downstream actions, environment, and relevant failure impact. Separate model quality from deterministic tool authorization and application checks. Document the actual allowed action boundary and its source; missing authority remains unknown.
2. **Choose cases by risk and coverage.** Use permitted historical examples, expert-labeled cases, or clearly marked synthetic fixtures. Cover relevant segments, boundary conditions, known failure modes, and critical harms. Record input, expected outcome/rubric, provenance, and critical-failure rule per case. Keep evaluation cases separate from tuning where possible; no fixed case count proves safety.
3. **Agree the pass rule before the run.** Define quality thresholds, critical failures, coverage expectations, and acceptable uncertainty for this use. Use deterministic checks where possible and inspect subjective labels or judge reliability. Propose missing criteria for agreement; do not manufacture acceptance from the observed scores.
1. **Define the evaluated surface.** Name the model judgment, inputs, outputs, downstream actions, environment, and relevant failure impact. Separate model quality from deterministic tool authorization and application checks. Document the actual allowed action boundary and its source; missing authority remains unknown. When diagnosing quality failures, connect each observed failure to a permitted trace, expected behavior, and a code, human, or model check in the existing report; label anticipated risks separately.
2. **Choose cases by risk and coverage.** Use permitted historical examples, expert-labeled cases, or clearly marked synthetic fixtures. Cover relevant segments, boundary conditions, known failure modes, and critical harms. Record input, expected outcome/rubric, provenance, and critical-failure rule per case. Keep evaluation cases separate from tuning where possible; no fixed case count proves safety. State whether sampling seeks failure coverage or estimates usage frequency. Include random exploration where permitted; targeted or synthetic case results alone do not estimate production failure rates.
3. **Agree the pass rule before the run.** Define quality thresholds, critical failures, coverage expectations, and acceptable uncertainty for this use. Use deterministic checks where possible and inspect subjective labels or judge reliability. Validate new or changed code evaluators against independently expected passing, failing, and boundary cases. Check judgment proxies against expert labels: a citation ID found in retrieved sources does not prove that the answer is supported. Propose missing criteria for agreement; do not manufacture acceptance from the observed scores.
4. **Run the actual evaluated path.** Record model/provider version, prompts/configuration, retrieval corpus or tools, application revision, environment, fixtures, and run date. Repeat where variability matters. Report totals, per-segment results, critical failures, and limitations using [verification](verification.md). A model-only run does not prove the agent's tool boundary works.
5. **Verify action authority and controls.** Human approval is required where the user's policy or task requires it. Already agreed bounded automation may run within its documented actions, identities, environments, and limits; do not require fresh approval for every authorized action. Check enforcement outside the model, least privilege, input/output validation, cost/rate limits, stop conditions, observability, and recovery as applicable. Unknown or exceeded authority blocks those actions. Evaluation success never grants new authority.
6. **Make a scoped verdict.** Report **SHIP** only when agreed criteria pass, critical failures are zero, applicable authority/control checks pass, and material coverage gaps are resolved or the release is explicitly narrowed by the responsible decision-maker. Otherwise report **NO-SHIP** with the smallest corrective step: fix, gather evidence, descope, or reconsider the judgment surface. A SHIP verdict is technical evidence for the stated scope, not permission to deploy.
Expand Down
2 changes: 1 addition & 1 deletion skills/debug/.fde-generated.json
Original file line number Diff line number Diff line change
Expand Up @@ -6,7 +6,7 @@
"agents/openai.yaml": "f9178e65a1e2e9ee27f2f5917a36a50e3ebe64bd6e331b43379efcf1c7d5aac3",
"references/build.md": "daaa7be30ce6e4c42715dc708f877ab1d2270df0dffdcc33680085450b9f0b2f",
"references/debug.md": "bdc09abdd7a075574fc2daf91cecfd2858e67d75a964146a118bf782328192f6",
"references/eval-pack.md": "f883ea0be4777c9c3893db0a616b473f172140d5a3a1be4c449570c8608e4d5b",
"references/eval-pack.md": "40ca29fbee82ac1846d3a31ddf92ef20c6d7906ade53571decf81d2f79ee936c",
"references/integrate.md": "1cb7a60d7545b0bf224fce678a04ce6ccdf368c47877d9c8e4dc4272bb0d5b0c",
"references/qa.md": "41de4d70827c83291efa217e97d777f62ec2849827687fbba7e4b1d17484b87e",
"references/review.md": "36b77ed44989a125a36e9948ce19c2ced57d529bee06643e32dda96f896d9d5b",
Expand Down
6 changes: 3 additions & 3 deletions skills/debug/references/eval-pack.md
Original file line number Diff line number Diff line change
Expand Up @@ -6,9 +6,9 @@ Use [task context](task-context.md). Supplied permitted context and an evaluatio

## Method

1. **Define the evaluated surface.** Name the model judgment, inputs, outputs, downstream actions, environment, and relevant failure impact. Separate model quality from deterministic tool authorization and application checks. Document the actual allowed action boundary and its source; missing authority remains unknown.
2. **Choose cases by risk and coverage.** Use permitted historical examples, expert-labeled cases, or clearly marked synthetic fixtures. Cover relevant segments, boundary conditions, known failure modes, and critical harms. Record input, expected outcome/rubric, provenance, and critical-failure rule per case. Keep evaluation cases separate from tuning where possible; no fixed case count proves safety.
3. **Agree the pass rule before the run.** Define quality thresholds, critical failures, coverage expectations, and acceptable uncertainty for this use. Use deterministic checks where possible and inspect subjective labels or judge reliability. Propose missing criteria for agreement; do not manufacture acceptance from the observed scores.
1. **Define the evaluated surface.** Name the model judgment, inputs, outputs, downstream actions, environment, and relevant failure impact. Separate model quality from deterministic tool authorization and application checks. Document the actual allowed action boundary and its source; missing authority remains unknown. When diagnosing quality failures, connect each observed failure to a permitted trace, expected behavior, and a code, human, or model check in the existing report; label anticipated risks separately.
2. **Choose cases by risk and coverage.** Use permitted historical examples, expert-labeled cases, or clearly marked synthetic fixtures. Cover relevant segments, boundary conditions, known failure modes, and critical harms. Record input, expected outcome/rubric, provenance, and critical-failure rule per case. Keep evaluation cases separate from tuning where possible; no fixed case count proves safety. State whether sampling seeks failure coverage or estimates usage frequency. Include random exploration where permitted; targeted or synthetic case results alone do not estimate production failure rates.
3. **Agree the pass rule before the run.** Define quality thresholds, critical failures, coverage expectations, and acceptable uncertainty for this use. Use deterministic checks where possible and inspect subjective labels or judge reliability. Validate new or changed code evaluators against independently expected passing, failing, and boundary cases. Check judgment proxies against expert labels: a citation ID found in retrieved sources does not prove that the answer is supported. Propose missing criteria for agreement; do not manufacture acceptance from the observed scores.
4. **Run the actual evaluated path.** Record model/provider version, prompts/configuration, retrieval corpus or tools, application revision, environment, fixtures, and run date. Repeat where variability matters. Report totals, per-segment results, critical failures, and limitations using [verification](verification.md). A model-only run does not prove the agent's tool boundary works.
5. **Verify action authority and controls.** Human approval is required where the user's policy or task requires it. Already agreed bounded automation may run within its documented actions, identities, environments, and limits; do not require fresh approval for every authorized action. Check enforcement outside the model, least privilege, input/output validation, cost/rate limits, stop conditions, observability, and recovery as applicable. Unknown or exceeded authority blocks those actions. Evaluation success never grants new authority.
6. **Make a scoped verdict.** Report **SHIP** only when agreed criteria pass, critical failures are zero, applicable authority/control checks pass, and material coverage gaps are resolved or the release is explicitly narrowed by the responsible decision-maker. Otherwise report **NO-SHIP** with the smallest corrective step: fix, gather evidence, descope, or reconsider the judgment surface. A SHIP verdict is technical evidence for the stated scope, not permission to deploy.
Expand Down
2 changes: 1 addition & 1 deletion skills/evaluate/.fde-generated.json
Original file line number Diff line number Diff line change
Expand Up @@ -6,7 +6,7 @@
"agents/openai.yaml": "731f3c81d46444af542369f0aa27736d24530d00fc1aa717436ddefe3f3bdef7",
"references/build.md": "daaa7be30ce6e4c42715dc708f877ab1d2270df0dffdcc33680085450b9f0b2f",
"references/debug.md": "bdc09abdd7a075574fc2daf91cecfd2858e67d75a964146a118bf782328192f6",
"references/eval-pack.md": "f883ea0be4777c9c3893db0a616b473f172140d5a3a1be4c449570c8608e4d5b",
"references/eval-pack.md": "40ca29fbee82ac1846d3a31ddf92ef20c6d7906ade53571decf81d2f79ee936c",
"references/integrate.md": "1cb7a60d7545b0bf224fce678a04ce6ccdf368c47877d9c8e4dc4272bb0d5b0c",
"references/qa.md": "41de4d70827c83291efa217e97d777f62ec2849827687fbba7e4b1d17484b87e",
"references/review.md": "36b77ed44989a125a36e9948ce19c2ced57d529bee06643e32dda96f896d9d5b",
Expand Down
6 changes: 3 additions & 3 deletions skills/evaluate/references/eval-pack.md
Original file line number Diff line number Diff line change
Expand Up @@ -6,9 +6,9 @@ Use [task context](task-context.md). Supplied permitted context and an evaluatio

## Method

1. **Define the evaluated surface.** Name the model judgment, inputs, outputs, downstream actions, environment, and relevant failure impact. Separate model quality from deterministic tool authorization and application checks. Document the actual allowed action boundary and its source; missing authority remains unknown.
2. **Choose cases by risk and coverage.** Use permitted historical examples, expert-labeled cases, or clearly marked synthetic fixtures. Cover relevant segments, boundary conditions, known failure modes, and critical harms. Record input, expected outcome/rubric, provenance, and critical-failure rule per case. Keep evaluation cases separate from tuning where possible; no fixed case count proves safety.
3. **Agree the pass rule before the run.** Define quality thresholds, critical failures, coverage expectations, and acceptable uncertainty for this use. Use deterministic checks where possible and inspect subjective labels or judge reliability. Propose missing criteria for agreement; do not manufacture acceptance from the observed scores.
1. **Define the evaluated surface.** Name the model judgment, inputs, outputs, downstream actions, environment, and relevant failure impact. Separate model quality from deterministic tool authorization and application checks. Document the actual allowed action boundary and its source; missing authority remains unknown. When diagnosing quality failures, connect each observed failure to a permitted trace, expected behavior, and a code, human, or model check in the existing report; label anticipated risks separately.
2. **Choose cases by risk and coverage.** Use permitted historical examples, expert-labeled cases, or clearly marked synthetic fixtures. Cover relevant segments, boundary conditions, known failure modes, and critical harms. Record input, expected outcome/rubric, provenance, and critical-failure rule per case. Keep evaluation cases separate from tuning where possible; no fixed case count proves safety. State whether sampling seeks failure coverage or estimates usage frequency. Include random exploration where permitted; targeted or synthetic case results alone do not estimate production failure rates.
3. **Agree the pass rule before the run.** Define quality thresholds, critical failures, coverage expectations, and acceptable uncertainty for this use. Use deterministic checks where possible and inspect subjective labels or judge reliability. Validate new or changed code evaluators against independently expected passing, failing, and boundary cases. Check judgment proxies against expert labels: a citation ID found in retrieved sources does not prove that the answer is supported. Propose missing criteria for agreement; do not manufacture acceptance from the observed scores.
4. **Run the actual evaluated path.** Record model/provider version, prompts/configuration, retrieval corpus or tools, application revision, environment, fixtures, and run date. Repeat where variability matters. Report totals, per-segment results, critical failures, and limitations using [verification](verification.md). A model-only run does not prove the agent's tool boundary works.
5. **Verify action authority and controls.** Human approval is required where the user's policy or task requires it. Already agreed bounded automation may run within its documented actions, identities, environments, and limits; do not require fresh approval for every authorized action. Check enforcement outside the model, least privilege, input/output validation, cost/rate limits, stop conditions, observability, and recovery as applicable. Unknown or exceeded authority blocks those actions. Evaluation success never grants new authority.
6. **Make a scoped verdict.** Report **SHIP** only when agreed criteria pass, critical failures are zero, applicable authority/control checks pass, and material coverage gaps are resolved or the release is explicitly narrowed by the responsible decision-maker. Otherwise report **NO-SHIP** with the smallest corrective step: fix, gather evidence, descope, or reconsider the judgment surface. A SHIP verdict is technical evidence for the stated scope, not permission to deploy.
Expand Down
Loading
Loading