🐢 Open-Source Evaluation & Testing library for LLM Agents
-
Updated
Jul 21, 2026 - Python
🐢 Open-Source Evaluation & Testing library for LLM Agents
A comprehensive Python framework for evaluating LLM-extracted structured data against ground truth labels. Supports binary classification, scalar values, and list fields with detailed performance metrics, confidence-based evaluation, and statistical uncertainty quantification via non-parametric bootstrap confidence intervals.
Evaluation & testing framework for computer vision models
Build confidence in your AI with systematic slice-based testing
Lightweight Python library for automated data leakage detection in ML datasets. Detects duplicate rows, ID overlap, target leakage, and temporal leakage.
Per-layer ONNX export parity validator: PyTorch + ONNX Runtime cross-checked layer by layer in a C++ comparator with Python fallback
Add a description, image, and links to the ml-validation topic page so that developers can more easily learn about it.
To associate your repository with the ml-validation topic, visit your repo's landing page and select "manage topics."