📚 A curated list of papers & technical articles on AI Quality & Safety
-
Updated
Apr 14, 2025
📚 A curated list of papers & technical articles on AI Quality & Safety
Proven 2026 Multi-Agent AI Review System – Verdict-Driven Quality Control
Ship evals before you ship features.
Open-source AI model evaluation and benchmarking framework for LLMs (OpenAI, Ollama, Claude, Gemini)
Eval framework. Define correct, test against it, get results.
Find what your AI agent gets wrong — before you have a rubric. Qualitative eval for PMs.
Diagnose your AI agents in production. Extract policies from prompts, evaluate traces, generate diagnostic reports.
A framework-agnostic metric for measuring AI code generation quality. Sealed-envelope testing protocol + reference validators.
Open-source AI agent security testing framework. Test for prompt injection, data leakage, and privilege escalation before production.
Python SDK for IvyCheck
Provider-agnostic LLM evaluation harness: golden dataset, deterministic + LLM-as-judge scoring, RAG failure attribution, red-team suite, severity-weighted CI gate.
朱雀 Suzaku — AI 生成品質模組。諂媚抑制、建設性挑戰、輸出適配、上下文錨定、一致性守護。基於 LDRIT 設計。
Open Council: Claude Code skill for multi-agent AI quality control. Ten specialist employees draft the work. Ten industry-veteran board members review it. You see one verdict. Anti-sycophancy, anti-hallucination, multi-perspective review.
Universal skill enhancement layer for Claude Code. Sees what your skill was trying to do, grades the gap, drives the rewrite.
Evaluate your LLM apps with one function call. Hallucination detection, RAG scoring, and agent evals for OpenAI, Anthropic, and more. 14 evaluators, pytest plugin, composite trust scores.
AI Agent Ops framework for Claude Code — independent evaluator, adversarial review, and pre-commit quality gate for AI-generated code.
Self-improving AI quality system - auto-score, auto-regenerate, auto-learn.
A 5-layer adversarial quality gate for Claude Code. Catches factual errors, score inflation, and buried conclusions before your AI output ships.
AI evaluation platform for measuring reliability, evidence grounding, and benchmark performance of LLM systems.
Evidence-backed quality control for LLM outputs: rubrics, claim grounding, safety gates, human review, and reports.
Add a description, image, and links to the ai-quality topic page so that developers can more easily learn about it.
To associate your repository with the ai-quality topic, visit your repo's landing page and select "manage topics."