Skip to content

Add EmbodiedMemory-Bench - #6

Open
zwq2018 wants to merge 1 commit into
OpenEnvision:mainfrom
zwq2018:add/embodiedmemory-bench
Open

zwq2018 wants to merge 1 commit into
OpenEnvision:mainfrom
zwq2018:add/embodiedmemory-bench

Conversation

@zwq2018

@zwq2018 zwq2018 commented Sep 29, 2026

Copy link
Copy Markdown

Adds EmbodiedMemory-Bench: Benchmarking Embodied Memory for Long-Horizon Embodied Tasks (arXiv, September 2026) to 6.6 Physical Robotics and Multi-Robot Systems.

Uses one canonical benchmark entry and attaches the accompanying memory system, code, and data to it. The entry explicitly identifies simulation and interactive execution; it does not claim real-robot validation.

The paper introduces 2,554 interactive episodes for visual recall, dynamic state tracking, interaction outcomes, and experience generalization, together with the Embodied-Memorizer external memory system.

Resources: code, project, dataset.

Validation: checked the target listing and open PRs for duplicates, preserved the existing entry format, and inspected the changed-file diff.

Disclosure: submitted on behalf of the paper authors.

Protocol and placement

The primary artifact is the interactive benchmark, so it has one entry in Section 6.6, with the bundled EMem baseline and resources attached. Inputs include the task instruction, accumulated visual observations and interaction history; the evaluated agent outputs navigation/interaction actions in AI2-THOR. Returned observations and execution outcomes update the memory used for later choices. This is simulation, and the reported metrics concern task success and memory-augmented interaction efficiency rather than real-robot validation. Task families and rollout limits are defined by the released episode manifests and evaluator; reproduction should retain those settings.

Checklist

  • Searched for duplicate titles, arXiv identifiers, and repository URLs.
  • Selected one canonical placement based on the primary evaluated contribution.
  • Verified the first public month and preprint status from primary evidence.
  • Added direct primary links; all four resource destinations returned HTTP 200.
  • Proposed system passes all four system tests — N/A: benchmark entry, with the bundled baseline attached.
  • Used the benchmark template and reverse-chronological placement after the section introduction.
  • Preserved existing entries and anchors; verified the website parser includes the new record once.
  • Identified environment, observation interface and metrics without unverified release/venue claims.
  • Disclosed author affiliation above.
  • Ran local checks and documented their scope below.

Validation: all 8 existing catalog tests pass with node --test tests/catalog.test.mjs. npx --yes awesome-lint README.md reports the same single missing-Git-metadata diagnostic on both source and edited downloaded snapshots, with no content diagnostics. An equivalent HTTP checker covered all four added resource destinations (paper/code/project/dataset), each HTTP 200; it did not recheck the entire historical link collection. Markdown was rendered and the added resource links and metadata inspected.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant