ChinaTravel: A Real-World Benchmark for Language Agents in Chinese Travel Planning
-
Updated
Aug 24, 2026 - Python
ChinaTravel: A Real-World Benchmark for Language Agents in Chinese Travel Planning
[ACL 2025] A curated list of papers and resources based on "PlanGenLLMs: A Modern Survey of LLM Planning Capabilities"
Adaptive Iterative Feedback Prompting for Obstacle-Aware Path Planning via LLMs - LM4Planning - AAAI2025
Awesome-LLM-Planning
PolicyEval checks whether an LLM's output follows the rules you set
Benchmark measuring how LLM planners fail at robot tasks, not just how often: trap-labelled instructions with machine-checked ground truth, differential-tested against pyperplan, no human or LLM judging anywhere. Every label proof re-run in CI.
Repo for exploring the (in)effectiveness of chain of thought in planning
Heterogeneous GNN policies for constrained zero-shot morphology transfer in legged locomotion with ROS2/Gazebo deployment, YOLO-based vision, and LLM-guided navigation.
The project is intended to evolve into a study prototype for LLM-based systems (e.g., agents, planners, tool calling)
Benchmarks a 1.5B reasoning model (DeepSeek-R1) vs. SOTA non-reasoning LLMs for text-to-SQL using a generator-discriminator planning framework. Proposes a novel method to extract soft scores from chain-of-thought outputs for fine-grained candidate ranking.
Provenance-locked benchmark and evaluation pipeline for verifiable LLM planning across five PDDL domains.
Add a description, image, and links to the llm-planning topic page so that developers can more easily learn about it.
To associate your repository with the llm-planning topic, visit your repo's landing page and select "manage topics."