Agentic RL 中文零基础教程(25 章):从概念到 GRPO 实战,含 TRL 最小可跑示例。第 25 章讲清 Jev / TypeSafe System One 判别模型与 RL 的能力边界 | Chinese Agentic RL tutorial, 25 chapters + Jev-vs-RL boundary analysis
-
Updated
Sep 24, 2026 - Python
Agentic RL 中文零基础教程(25 章):从概念到 GRPO 实战,含 TRL 最小可跑示例。第 25 章讲清 Jev / TypeSafe System One 判别模型与 RL 的能力边界 | Chinese Agentic RL tutorial, 25 chapters + Jev-vs-RL boundary analysis
A curated list of awesome resources about reward construction for AI agents. This repository covers cutting-edge research, and practical guides on defining and collecting rewards to build more intelligent and aligned AI agents.
Official Code for AutoLLMResearch: Training Research Agents for Automating LLM Experiment Configuration — Learning from Cheap, Optimizing Expensive
Self-Mutating Agent Gym
A curated list of agentic environment synthesis, evolution, quality & scaling — built on three surveys (AEE, Environment Scaling, ACE lens)
Synthetic environments for training robust tool-use agents via trajectory synthesis
End-to-end training pipeline for mobile game UA tool-calling agents, covering rule-based synthetic data generation across 7 workflows, OpenAI Messages conversion, Qwen3 LoRA SFT, GRPO/RLVR alignment, and benchmark evaluation.
Turn any real software into a replayable RL environment for training AI agents — deterministic replay, verifiable rewards, TRL & verifiers adapters.
AI 个人记忆训练库:让 AI 记住你的沟通习惯与环境事实,越用越懂你(两级记忆架构,Claude Code 适配)
14 Principles. 23 Real Corrections. A systematic methodology for training AI agents through correction loops.
AI schooling: experienced agents train newly hatched ones, then examine them cold to graduate them. The coop is the local RAPP neighborhood - several twins (human and AI), one world, no collisions. Pattern dedicated to the public domain.
The Flight Simulator for Production AI Agents — Generate high-quality synthetic trajectories for training reliable SRE, DevOps, and infra agents.
The Vercel for Agent Training - Train production-ready AI agents with 95% tool reliability
Deep reinforcement learning project comparing DQN variants including baseline DQN, Double DQN, curriculum learning, and reward shaping.
Neurochemical behavior training for AI agents — PentaDrive model with 5 drives, 3 phases, MCP server, and structured training modules
Local macOS recorder for voluntary computer-use demonstrations and agent-training datasets. Timestamped inputs, cursor trails, and read-only MCP. Experimental.
Open-source benchmark for measuring whether AI agents improve across unseen missions, with validity audits, rotated mission packs, adapter tests, and traceable reports.
Correctness-by-construction distillation for agentic tool-calling models (LatticeAG Forge series)
A bilingual, buildless, single-file landing page for an agent training-environment platform. No bundler, no framework, no runtime CDN.
To associate your repository with the agent-training topic, visit your repo's landing page and select "manage topics."