Skip to content
xiaojunlanPublic

About

CARE: corrective data generation and FSR-Bench evaluation for robotic manipulation

Resources

Stars

49 stars

Watchers

0 watching

Forks

Repository files navigation

CARE: Experience-Guided Atomic Corrective Execution

CARE learns short corrective behaviors from failures encountered during robot execution. The code is organized around failure-distribution estimation → corrective-data collection → policy evaluation. It also includes the FSR-Bench recovery evaluation scenarios. This repository is a code-only snapshot built on RoboTwin 2.0; large datasets, assets, model checkpoints, videos, and virtual environments are not included.

Installation

Install the simulation environment by following the official RoboTwin 2.0 installation guide. Use the RoboTwin 2.0 Python environment and its required SAPIEN, CuRobo, and rendering dependencies. The commands below must be run from this repository's root:

cd care
conda activate care

Download/configure the RoboTwin 2.0 assets separately under assets/. The copied task_config/_embodiment_config.yml expects paths such as ./assets/embodiments/aloha-agilex/. Provide the policy checkpoint and launch a compatible HTTP π₀ inference service before running evaluation. The evaluation clients use GET /health and POST /set_instruction, /predict, and /reset; check the service first:

curl http://127.0.0.1:5000/health

policy/pi0/scripts/serve_policy.py is a WebSocket server and does not implement this HTTP interface. Use your existing compatible HTTP service. The local script/_install.sh belongs to an older RoboTwin layout; follow the official guide for environment setup.

Quick start

The examples use GPU 0, seed 0, an HTTP model service at http://127.0.0.1:5000, and 100 evaluation episodes. Change these values for your setup. Evaluation outputs go to eval_result/; collected trajectories go to the save_path in the selected task_config/*.yml (normally data/).

1. Run a nominal policy and model its failures

Use the no-correction evaluation entry point to obtain failed rollouts for failure analysis:

bash policy/pi0/eval_put_object_cabinet_no_atom.sh \
  put_object_cabinet demo_randomized 0 http://127.0.0.1:5000 100

The script runs policy/pi0/eval_put_object_cabinet_no_atom.py; its five arguments are task, task config, seed, server URL, and number of trials. It writes the aggregate success rate to eval_result/<task>/pi0/<config>/<label>/<timestamp>/_result.txt and can save videos when eval_video_log: true.

Distribution-modeling input: this evaluator currently records success/failure, but it does not export the stage-critical dx, dy, dz, yaw values required to fit the paper's failure distribution. Record those offsets at grasp closure or object release, grouped by task/stage/arm. Fit the failure distribution from nominal-policy failures before using its parameters for corrective-data generation; this repository does not currently include a distribution-fitting command. Do not mix corrective demonstrations into the nominal failure statistics.

2. Generate corrective demonstrations

The task environment samples a perturbation, advances the simulation to a failure state, then uses its motion-planning oracle to collect a short correction. For example, envs/put_object_cabinet_regrasp.py contains the grasp-offset sampler and re-grasp behavior:

# Set episode_num, collect_data, and save_path in task_config/demo_randomized.yml first.
bash collect_data.sh put_object_cabinet_regrasp demo_randomized 0

collect_data.sh calls script/collect_data.py and writes trajectories to data/put_object_cabinet_regrasp/demo_randomized/ with the default config. Other task environments can be collected in the same form: bash collect_data.sh <envs task name> <task config name> <gpu id>. The bundled collect_data.sh tries script/.update_path.sh, which is absent in this snapshot; that call is suppressed and collection continues. Ensure the embodiment paths in task_config/_embodiment_config.yml already resolve before running.

Fitted-parameter handoff: the task environments keep distribution parameters in their Python sampling functions. Update the appropriate sampler with the checked failure distribution before collecting experience-guided data.

Process the collected HDF5 episodes, assign stage labels, attach per-frame instructions, then convert the processed data to LeRobot. One entry point handles task-specific automatic rules for handover_block, lift_pot, stamp_seal, place_bread_skillet, place_dual_shoes, and put_object_cabinet; other tasks can keep existing joint_action/phase labels or pass --phase for a single-stage correction dataset. Handover and shoe transitions use gripper-based approximations; inspect their labels before training. For example:

python dealdata/process_task_data.py place_bread_skillet --config demo_randomized
bash policy/pi0/process_data_pi0.sh place_bread_skillet demo_randomized 50
bash policy/pi0/generate.sh \
  policy/pi0/processed_data/place_bread_skillet-demo_randomized-50 \
  lerobot_data/place_bread_skillet

Set 50 to the number of collected episodes. The last argument is an existing local LeRobot dataset to append to. policy/pi0/scripts/process_data.py defines the instruction for each task and phase; add a mapping there when introducing a new phase. The conversion keeps frames from every phase. Dataset paths and a short command reference are in dealdata/readme.md.

3. Train CARE policy

Train the processed LeRobot data with Future Robots. Configure its dataset and task-specific training config for the CARE trajectories. For this release, disable MA-VLA agent-order shuffling and image masking in the DroidMultiInputs_atom data transform:

shuffle_prob=0
image_mask_prob=0

Then compute normalization statistics and launch training from the Future Robots repository using your CARE config:

uv run scripts/compute_norm_stats.py --config-name YOUR_CARE_CONFIG
XLA_PYTHON_CLIENT_MEM_FRACTION=1 uv run scripts/train.py YOUR_CARE_CONFIG --exp-name=care_run

Point the evaluation model service to the resulting checkpoint. These two settings disable the named augmentations; check other augmentation options separately in the selected training config.

4. Evaluate CARE on RoboTwin 2.0

This handover-block evaluator uses point-cloud monitoring and stage-level correction:

python policy/pi0/eval_vla_handover_block_pointcloud.py \
  --config policy/pi0/deploy_policy.yml \
  --server_url http://127.0.0.1:5000 \
  --overrides \
  --task_name handover_block \
  --task_config demo_randomized \
  --ckpt_setting handover_block \
  --seed 0 --policy_name pi0 --test_num 100

--ckpt_setting is the result-directory label; the actual checkpoint is loaded by the HTTP model service. Results are saved under eval_result/handover_block/pi0/demo_randomized/. This evaluator may also write point-cloud and debug outputs, so allocate storage outside the code-only snapshot for large runs.

5. Evaluate FSR-Bench

FSR-Bench has five tasks, each with Easy and Hard evaluators. All ten evaluator scripts below are present. Run commands from the care/ directory.

Task Difficulty Evaluator (policy/pi0/) TASK (--task_name, --ckpt_setting)
Handover Block Easy eval_vla_handover_block_benchmark_easy.py handover_block_benchmark_easy
Handover Block Hard eval_vla_handover_block_benchmark_hard.py handover_block_benchmark_hard
Lift Pot Easy eval_vla_lift_pot_benchmark_easy.py lift_pot_benchmark_easy
Lift Pot Hard eval_vla_lift_pot_benchmark_hard.py lift_pot_benchmark_hard
Stamp Seal Easy eval_stamp_seal_benchmark_easy.py stamp_seal_benchmark_easy
Stamp Seal Hard eval_stamp_seal_benchmark_hard.py stamp_seal_benchmark_hard
Place Bread Skillet Easy eval_vla_place_bread_skillet_benchmark_easy.py place_bread_skillet_benchmark_easy
Place Bread Skillet Hard eval_vla_place_bread_skillet_benchmark_hard.py place_bread_skillet_benchmark_hard
Place Dual Shoes Easy eval_vla_place_dual_shoes_benchmark_easy.py place_dual_shoes_benchmark_easy
Place Dual Shoes Hard eval_vla_place_dual_shoes_benchmark_hard.py place_dual_shoes_benchmark_hard

Set EVAL to the evaluator filename without .py, and TASK to the value in the same row. For example, Handover Block Easy runs as follows:

EVAL=eval_vla_handover_block_benchmark_easy
TASK=handover_block_benchmark_easy
CONFIG=demo_randomized
python "policy/pi0/${EVAL}.py" \
  --config policy/pi0/deploy_policy.yml \
  --server_url http://127.0.0.1:5000 \
  --overrides \
  --task_name "$TASK" \
  --task_config "$CONFIG" \
  --ckpt_setting "$TASK" \
  --seed 0 --policy_name pi0 --test_num 100

The example uses the included task_config/demo_randomized.yml; choose a suitable config for each benchmark scenario. policy/pi0/eval_vla_place_dual_shoes_benchmark_hard2.py is an additional Hard evaluator variant with a fixed starting seed of 100000. This code snapshot does not contain the full fixed test-initialization manifest for all 36 scenarios, so these evaluator commands alone do not establish the paper's exact FSR-Bench protocol.

Repository layout

  • envs/: nominal tasks, corrective-data oracles, and FSR-Bench recovery environments.
  • policy/pi0/: π₀ evaluation clients and policy code.
  • script/, collect_data.sh: RoboTwin data collection and evaluation entry points.
  • task_config/: simulation and collection settings.
  • dealdata/: trajectory processing and phase extraction.

Acknowledgement

This work builds on RoboTwin 2.0. We thank the RoboTwin 2.0 authors for the simulation platform, task environments, and data-collection infrastructure. Installation and asset preparation follow the official RoboTwin 2.0 documentation.

License

See LICENSE and the licenses of bundled third-party components.

About

CARE: corrective data generation and FSR-Bench evaluation for robotic manipulation

Resources

Stars

49 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages