An eval harness and report on some of the latest LLMs that can be run locally
-
Updated
May 6, 2026 - Python
An eval harness and report on some of the latest LLMs that can be run locally
Automated RAG & LLM Evaluation Pipeline — LangChain · Ragas · Jenkins · GitHub Copilot Proxy ($0 cost)
This repository is a practical overview of inference engineering for Large Language Model deployment, with vLLM as the serving engine. It is meant to help you understand not only how to start a model server, but also how to reason about throughput, latency, batching, quantization, KV cache memory, and GPU VRAM requirements before deploying a model.
LLM prompt benchmark on 150 RAGBench samples × 5 strategies. RAG: 89.5% faithfulness, 10.5% ungrounded vs zero-shot: 31.3%, 68.7%. Qwen3 + Gemini judge.
PrefRank predicts which LLM response humans prefer using a 3-class classification framework. Combines linguistic, structural, and TF-IDF features with calibrated models (MLP, RF, XGBoost, LightGBM) — no embeddings required.
AI-powered email generation assistant built with Streamlit and GPT-4o-mini or gemini, featuring prompt engineering evaluation and custom metrics comparison.
Physician-reviewed evaluation of a local medical LLM on original USMLE-style clinical questions.
An agent proposes, a declarative policy authorizes, and the gate holds the credentials so the model never does. Gmail + Calendar + Notion.
Vision-language triage for forensic analysts: BLIP-2 + VQA to classify and sort images recovered from seized devices, replacing manual review of large media dumps.
A tool that scores an LLM's response to a prompt against a five-part rubric, using Claude as the judge
A production-oriented Retrieval-Augmented Generation system with hybrid retrieval, reranking, citations, evaluation, and regression testing.
Python agent that ingests funnel data, generates a written drop-off diagnosis using Claude, and includes an A/B eval of the design decision.
On-device LLM evaluation on Apple Silicon with MLX: Laya-mlx, Llama 3.2 3B at 4-bit, 8-bit and bf16 quantization, LoRA fine-tuning and adapter fusing, compared on accuracy, speed and memory for GitHub issue classification.
To associate your repository with the llmevaluation topic, visit your repo's landing page and select "manage topics."