Skip to content

docs: role-based get-started, a cookbook of recipes, and an evaluate section - #2102

Merged
strawgate merged 4 commits into
mainfrom
docs/ia-entry-points
Jul 23, 2026
Merged

strawgate merged 4 commits into
mainfrom
docs/ia-entry-points

Conversation

@strawgate

@strawgate strawgate commented Jul 22, 2026 •

Copy link
Copy Markdown
Contributor

Adds three new documentation entry points, cross-linked and wired into the nav.

  • Get started by role — a "choose your path" hub plus five role pages (AI engineer, backend/SRE, platform/security, product/growth, solo builder), each a short ordered path to the pages that matter for that reader.
  • Cookbook — six runnable, end-to-end recipes: instrument and evaluate an agent, debug a slow endpoint, debug a slow tool call, roll out a prompt safely, and track and alert on LLM cost. Each ends with a troubleshooting section.
  • Evaluate — an overview of the datasets → scorers → scores → experiments model, plus a human-review page.

Every code example is verified to run end to end against real APIs, databases, and a local Logfire stack.

Companion routing (the Cookbook section and evaluate routes) is in pydantic/unified-docs#83, to merge after this.

Test plan

  • tests/test_docs.py (formatting + runnable examples) passes.
  • Internal links and heading anchors validated across all new pages.
  • Recipe flows verified live: query API, datasets, and slow-query / slow-tool-call traces.

Review in cubic

@macroscopeapp

macroscopeapp Bot commented Jul 22, 2026 •

Copy link
Copy Markdown

Macroscope skipped reviewing this pull request. Per-review cost limit exceeded (workspace setting).

This review would cost an estimated $5.10, which exceeds your per-review limit of $5.00.

The top 3 files driving up this estimate:

File Size Estimate
docs/cookbook/evaluate-an-agent.md 20.51KB $1.03
docs/cookbook/track-llm-cost.md 13.53KB $0.68
docs/cookbook/roll-out-a-prompt-safely.md 12.35KB $0.62

Tip

To get this pull request reviewed, you can:

  1. Comment @macroscope-app on this PR to request a manual review (monthly spend limits still apply).
  2. Exclude the file(s) above from review by adding a pattern to your .macroscope/ignore.md — note that creating this file replaces Macroscope's built-in default ignores rather than extending them.
  3. Raise your cost limit in your workspace billing settings.

Turn off this reminder going forward

@coderabbitai

coderabbitai Bot commented Jul 22, 2026 •

Copy link
Copy Markdown

Review Change Stack

Note

Reviews paused

It looks like this branch is under active development. To avoid overwhelming you with review comments due to an influx of new commits, CodeRabbit has automatically paused this review. You can configure this behavior by changing the reviews.auto_review.auto_pause_after_reviewed_commits setting.

Use the following commands to manage reviews:

  • @coderabbitai resume to resume automatic reviews.
  • @coderabbitai review to trigger a single review.

Use the checkboxes below for quick actions:

  • ▶️ Resume reviews
  • 🔍 Trigger review
📝 Walkthrough

Walkthrough

Adds role-based getting-started paths and navigation for Evaluate and Cookbook sections. Adds cookbook guides for debugging latency, evaluating agents, managing prompt rollouts, and tracking LLM cost. Adds evaluation overview and human-review documentation covering scorers, experiments, annotations, and feedback recording.

Possibly related PRs

🚥 Pre-merge checks | ✅ 4
✅ Passed checks (4 passed)
Check name Status Explanation
Title check ✅ Passed The title clearly summarizes the main additions: role-based getting started pages, cookbook recipes, and evaluate docs.
Description check ✅ Passed The description matches the changeset and accurately describes the new get-started, cookbook, and evaluate documentation.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch docs/ia-entry-points

Comment @coderabbitai help to get the list of available commands.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 4

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@docs/cookbook/evaluate-an-agent.md`:
- Around line 230-247: Make the no-tool evaluation deterministic by ensuring the
agent used in the `support-agent-no-tools` experiment cannot access or select
`order_status`, rather than only changing its system prompt. Update the agent
setup in `agent.py` or the evaluation configuration in `eval.py` to use a
separate tool-free agent or explicitly disable tool use, while preserving the
separate run name and expected failures for both order cases.

In `@docs/cookbook/track-llm-cost.md`:
- Around line 124-140: The daily cost alert example must use a calendar-day
window consistently. Update the SQL query in the “Alert when daily spend crosses
a threshold” section to add an explicit timezone-aware filter on
records.start_timestamp for today, and revise the surrounding window guidance so
it no longer recommends a rolling 24-hour window as the default.

In `@docs/evaluate/human-review.md`:
- Around line 48-65: Update the Python example around the logfire span and
record_feedback usage to use the repository’s configured indented Markdown
code-block style instead of fenced backticks, preserving the example content and
formatting. Do not change lint configuration unless this documentation
intentionally requires fenced blocks.

In `@docs/get-started/for-platform-and-security.md`:
- Around line 18-20: The documentation introduces non-standard acronyms without
expansion. In docs/get-started/for-platform-and-security.md lines 18-20, expand
SOC 2, HIPAA, GDPR, and OIDC at first use while retaining approved EU usage; in
docs/get-started/for-product-and-growth.md line 16, write OpenFeature Remote
Evaluation Protocol (OFREP) at first use.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro Plus

Run ID: fb721549-2d94-405f-ac03-bc515fa3e605

📥 Commits

Reviewing files that changed from the base of the PR and between 86f2514 and d15902a.

⛔ Files ignored due to path filters (2)
  • docs/images/guide/live-view-demo-shop.png is excluded by !**/*.png
  • docs/images/guide/llms-view.png is excluded by !**/*.png
📒 Files selected for processing (15)
  • docs/cookbook/debug-a-slow-endpoint.md
  • docs/cookbook/debug-a-slow-tool-call.md
  • docs/cookbook/evaluate-an-agent.md
  • docs/cookbook/index.md
  • docs/cookbook/roll-out-a-prompt-safely.md
  • docs/cookbook/track-llm-cost.md
  • docs/evaluate/human-review.md
  • docs/evaluate/overview.md
  • docs/get-started/choose-your-path.md
  • docs/get-started/for-ai-engineers.md
  • docs/get-started/for-backend-and-sre.md
  • docs/get-started/for-platform-and-security.md
  • docs/get-started/for-product-and-growth.md
  • docs/get-started/solo-builder.md
  • docs/nav.json
🔗 Linked repositories identified

CodeRabbit considers these linked repositories for cross-repo context during reviews:

  • pydantic/logfire (manual)
  • pydantic/pydantic-ai (manual)
  • pydantic/platform (auto-detected)
  • pydantic/pydantic (auto-detected)

Comment thread docs/cookbook/evaluate-an-agent.md Outdated
Comment thread docs/cookbook/track-llm-cost.md Outdated
Comment thread docs/evaluate/human-review.md
Comment thread docs/get-started/for-platform-and-security.md Outdated

@cubic-dev-ai cubic-dev-ai Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

1 issue found across 17 files

Confidence score: 5/5

  • In docs/evaluate/overview.md, the main risk is documentation usability: at ~1,360 words it exceeds the style target and may be harder for readers to scan or retain, but it does not appear to introduce product or runtime regression risk. This PR is safe to merge, with a follow-up to trim/split the page to the 300–800 word guideline for consistency.
Prompt for AI agents (unresolved issues)

Check if these issues are valid — if so, understand the root cause of each and fix them. If appropriate, use sub-agents to investigate and fix each issue separately.


<file name="docs/evaluate/overview.md">

<violation number="1" location="docs/evaluate/overview.md:16">
P3: Page is ~1,360 words; the style guide caps overview/concept pages at 300–800 words. The extra length here comes from covering multiple sub-topics (scaffold, offline/online, score shapes, three best-practice tips, consequences) that each add concept-level copy. An overview should orient and fan out within 800 words. Consider trimming or splitting: move the best-practice tips and score-shape details into dedicated how-to pages and keep only the mental model and fan-out in this overview.</violation>
</file>

Reply with feedback, questions, or to request a fix.

Re-trigger cubic

Comment thread docs/cookbook/evaluate-an-agent.md Outdated
Comment thread docs/cookbook/track-llm-cost.md
Comment thread docs/evaluate/overview.md
@strawgate
strawgate force-pushed the docs/ia-entry-points branch from d15902a to 35bf843 Compare July 22, 2026 06:14

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 5

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@docs/cookbook/debug-a-slow-endpoint.md`:
- Around line 37-42: Update the PostgreSQL docker run command to bind the
published port to loopback by replacing the current port mapping with a
127.0.0.1-bound mapping, while preserving the existing container and database
configuration.

In `@docs/cookbook/evaluate-an-agent.md`:
- Around line 234-238: Update the agent declaration in the evaluation example to
use the same support_agent variable referenced by eval.py, while preserving its
model, system prompt, and tool_choice settings so the tool remains registered
but cannot be called.
- Around line 136-145: Remove the duplicate logfire.configure() and
logfire.instrument_pydantic_ai() calls from the evaluation setup, along with the
now-unused logfire import if applicable. Keep Logfire initialization solely in
agent.py and preserve the remaining Dataset evaluation logic.

In `@docs/get-started/choose-your-path.md`:
- Line 16: Update the first-use sentence in choose-your-path to briefly define
“tool calls” and “tokens,” or link each term to an existing explanation, while
preserving the surrounding description and making the terms understandable to
unfamiliar readers.

In `@docs/get-started/for-product-and-growth.md`:
- Line 16: Update the feature-flags description to spell out OFREP on first use
as “OpenFeature Remote Evaluation Protocol (OFREP),” while preserving the
existing OpenFeature context and link.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro Plus

Run ID: 346e5e79-ef10-4f4f-bffe-9832e11e09e8

📥 Commits

Reviewing files that changed from the base of the PR and between d15902a and 35bf843.

⛔ Files ignored due to path filters (2)
  • docs/images/guide/live-view-demo-shop.png is excluded by !**/*.png
  • docs/images/guide/llms-view.png is excluded by !**/*.png
📒 Files selected for processing (15)
  • docs/cookbook/debug-a-slow-endpoint.md
  • docs/cookbook/debug-a-slow-tool-call.md
  • docs/cookbook/evaluate-an-agent.md
  • docs/cookbook/index.md
  • docs/cookbook/roll-out-a-prompt-safely.md
  • docs/cookbook/track-llm-cost.md
  • docs/evaluate/human-review.md
  • docs/evaluate/overview.md
  • docs/get-started/choose-your-path.md
  • docs/get-started/for-ai-engineers.md
  • docs/get-started/for-backend-and-sre.md
  • docs/get-started/for-platform-and-security.md
  • docs/get-started/for-product-and-growth.md
  • docs/get-started/solo-builder.md
  • docs/nav.json
🔗 Linked repositories identified

CodeRabbit considers these linked repositories for cross-repo context during reviews:

  • pydantic/logfire (manual)
  • pydantic/pydantic-ai (manual)
  • pydantic/platform (auto-detected)
  • pydantic/pydantic (auto-detected)
🚧 Files skipped from review as they are similar to previous changes (8)
  • docs/evaluate/overview.md
  • docs/get-started/solo-builder.md
  • docs/cookbook/debug-a-slow-tool-call.md
  • docs/cookbook/index.md
  • docs/nav.json
  • docs/get-started/for-platform-and-security.md
  • docs/cookbook/track-llm-cost.md
  • docs/cookbook/roll-out-a-prompt-safely.md

Comment thread docs/cookbook/debug-a-slow-endpoint.md
Comment thread docs/cookbook/evaluate-an-agent.md
Comment thread docs/cookbook/evaluate-an-agent.md Outdated
Comment thread docs/get-started/choose-your-path.md Outdated
Comment thread docs/get-started/for-product-and-growth.md Outdated
@strawgate
strawgate force-pushed the docs/ia-entry-points branch from 35bf843 to e0d12e5 Compare July 22, 2026 20:05
strawgate and others added 4 commits July 22, 2026 15:16
Six runnable recipes plus a landing page: instrument and evaluate an agent,
debug a slow endpoint, debug a slow tool call, roll out a prompt safely, and
track and alert on LLM cost. Every code example is verified to run end to end,
and each recipe closes with a short troubleshooting section.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
A "choose your path" hub plus five role pages (AI engineer, backend/SRE,
platform/security, product/growth, solo builder) that route each audience to
the pages that matter for them, each link with a one-line reason.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
An Evaluate section introducing the datasets/scorers/scores/experiments model
and how human review and user feedback land as the same kind of score.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
@strawgate
strawgate force-pushed the docs/ia-entry-points branch from e0d12e5 to 663cee3 Compare July 22, 2026 20:17
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant