The official evaluation suite and dynamic data release for MixEval.
-
Updated
Nov 10, 2024 - Python
The official evaluation suite and dynamic data release for MixEval.
MLOS is a project to enable autotuning for systems.
Python Multi-Process Execution Pool: concurrent asynchronous execution pool with custom resource constraints (memory, timeouts, affinity, CPU cores and caching), load balancing and profiling capabilities of the external apps on NUMA architecture
A toolkit for auto-generation of OpenAI Gym environments from RDDL description files.
NPBench - A Benchmarking Suite for High-Performance NumPy
Arline Benchmarks platform allows to benchmark various algorithms for quantum circuit mapping/compression against each other on a list of predefined hardware types and target circuit classes
Benchmarking machine learning inferencing on embedded hardware.
Telco pIPeline benchmarking SYstem
Benchmarking framework for Feature Selection and Feature Ranking algorithms 🚀
Deterministic runtime for agent evaluation
Framework for benchmarking deep learning operators for Apache MXNet
Codebase for Prompt Segmentation and Annotation Optimisation: Controlling LLM Behaviour via Optimised Segment-Level Annotations
STELLAR: A Search-Based Testing Framework for Large Language Model Applications" (SANER 2026) 🏆
Crossbar Parasitics Simulator – A tool for benchmarking parasitic resistance models in RRAM crossbars and evaluating neural networks under realistic hardware constraints.
A framework for benchmarking in python
PARROT (Performance Assessment of Reasoning and Responses On Trivia) is a novel benchmarking framework designed to evaluate Large Language Models (LLMs) on real-world, complex, and ambiguous QA tasks.
Public submission hub for PUMA local-LLM benchmark results
TraceOS standardizes AI experiments into reproducible, searchable, and comparable assets. One command runs experiments, generates reports, and produces structured analysis: capability vectors, failure taxonomy, and recommendations. Every run is tracked, traceable, and comparable. Built on ABC-130K (amazon-far/abc). Apache 2.0.
stat-my-agent ; benchmark consistency, tool-use, failure-recovery and goal-faithfulness — locally reproducible & shareable
scenario-driven load testing engine designed to be driven by AI coding assistants
Add a description, image, and links to the benchmarking-framework topic page so that developers can more easily learn about it.
To associate your repository with the benchmarking-framework topic, visit your repo's landing page and select "manage topics."