Skip to content

【No.8】Agentic 多协议一致性测试 - #344

Closed
AllureCurtain wants to merge 2 commits into
redai-studio:mainfrom
AllureCurtain:task8/agentic-protocol-golden-tests
Closed

AllureCurtain wants to merge 2 commits into
redai-studio:mainfrom
AllureCurtain:task8/agentic-protocol-golden-tests

Conversation

@AllureCurtain

Copy link
Copy Markdown

Summary

  1. Add offline canonical request golden tests for Chat Completions, OpenAI Responses and Anthropic Messages: one semantic conversation is authored natively in each protocol and must project onto an identical messages / tools / chat_template_kwargs triple.
  2. Cover text, assistant tool call + tool result, URL and base64 image inputs, tools, thinking enable/disable, assistant reasoning, and system message roles, asserted in three layers: per-protocol golden triple, cross-protocol agreement, and one session state hash.
  3. Reject malformed requests — unknown/illegal role, missing or empty tool call id, empty string/list content, invalid image block, reserved template kwargs — asserting the exact message, param and cause type with full field paths such as messages[2].tool_calls[0].id.
  4. Tighten relax/agentic/session/state.py where malformed input used to pass validation silently: canonicalize every legal image spelling to one block form (dropping request-only keys such as OpenAI's detail), and require a non-empty tool_calls[].id.

Type of Change

Type Included Details
Tests Yes Cross-protocol golden tests plus error-path and tolerance pinning.
Bug fix Yes Tightens request normalization and validation shared by the three protocol adapters.
Runtime behavior Limited No scheduling, lifecycle, training or model execution changes.
Documentation No
Build or tooling No No dependency, CI or build configuration changes.

Test plan

  • pytest tests/agentic/test_protocol_golden.py -q — 114 items, CPU only; no network, model service or GPU involved.
  • pre-commit run --files on all touched files — every hook passes (ruff, docformatter, gitleaks, whitespace/EOF checks).
  • Branch is rebased onto the latest upstream main (0651812).

Conventions

  • Test layout follows the existing tests/ per-directory structure (tests/agentic/__init__.py), and golden values are authored inline, consistent with the existing canonicalization tests in tests/test_agentic_rollout.py.
  • Reuses check_messages / normalize_tools / normalize_template_kwargs as the task requires; no new dependencies.

Related to #321

# 🐛 Bug Fix

## Canonicalize image blocks in check_messages

- Add `_canonical_image_block`: every legal image spelling across the three
  ingress protocols normalizes to `{"type": "image_url", "image_url": {"url": ...}}`
- Drop request-only keys (e.g. OpenAI `detail`) from the canonical state so two
  spellings of one image cannot produce two state hashes
- Reject a missing, empty or non-string url, reporting the exact field path

## Require tool call ids

- `tool_calls[].id` must now be a non-empty string; a missing id used to pass
  validation and let a malformed tool call into the canonical session state
# ✅ Tests

## Add golden tests for the three agentic ingress protocols

- Add `tests/agentic/test_protocol_golden.py` (114 cases) pinning the canonical
  `messages` / `tools` / `chat_template_kwargs` triple and session hash for
  Chat Completions, OpenAI Responses and Anthropic Messages
- 10 semantic scenarios authored natively in all three protocols, asserted in
  three layers: per-protocol golden, cross-protocol agreement, one hash
- 9 protocol-specific rules; four tool-argument encodings collapse to one
  canonical JSON string
- 25 illegal inputs assert exact message, `param` and cause type
- 7 tolerated anomalies and 2 dataset-shaped turns pinned, so later
  tightening is deliberate
- Add `tests/agentic/__init__.py` matching the existing per-directory layout
@rai-studio-bot

rai-studio-bot commented Sep 20, 2026 •

Copy link
Copy Markdown
Contributor

Nyanpasu 审查看板

审查状态: ✅ 已通过

审查版本: 39bbf7463a8d975b48e3c0d8e704a56e1d3dae84

已完成完整变更审查,未发现需要修改的问题。真实规范化函数的隔离验证 114 项通过;标准 pytest 因缺少 Ray 无法收集,暂无 CI 检查,未运行模型服务或多节点 GPU 集成测试。

没有未解决的审查问题。

Powered by Nyanpasu with gpt-6-astra medium, please check the suggestions carefully.

@rai-studio-bot rai-studio-bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

已审查 39bbf746 的完整变更及共享校验的调用路径,未发现需要修改的问题,同意合入。

从当前源码提取真实规范化函数进行隔离验证,114 项测试通过;标准 pytest 因缺少 Ray 在收集阶段停止,因此不视为完整模块测试通过。当前未发现 CI 检查,未运行模型服务或多节点 GPU 集成测试(本环境缺少相应依赖与硬件)。

Powered by Nyanpasu with gpt-6-astra medium, please check the suggestions carefully.

@AllureCurtain

Copy link
Copy Markdown
Author

@SigureMo 任务已完成,请求 review。

@SigureMo

SigureMo commented Oct 8, 2026

Copy link
Copy Markdown
Member

本赛题已经由 #335 合入并锁定,感谢参与~

@SigureMo SigureMo closed this Oct 8, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants