Repository navigation
Conversation
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Adds EmbodiedMemory-Bench: Benchmarking Embodied Memory for Long-Horizon Embodied Tasks (arXiv, September 2026) to 6.6 Physical Robotics and Multi-Robot Systems.
Uses one canonical benchmark entry and attaches the accompanying memory system, code, and data to it. The entry explicitly identifies simulation and interactive execution; it does not claim real-robot validation.
The paper introduces 2,554 interactive episodes for visual recall, dynamic state tracking, interaction outcomes, and experience generalization, together with the Embodied-Memorizer external memory system.
Resources: code, project, dataset.
Validation: checked the target listing and open PRs for duplicates, preserved the existing entry format, and inspected the changed-file diff.
Disclosure: submitted on behalf of the paper authors.
Protocol and placement
The primary artifact is the interactive benchmark, so it has one entry in Section 6.6, with the bundled EMem baseline and resources attached. Inputs include the task instruction, accumulated visual observations and interaction history; the evaluated agent outputs navigation/interaction actions in AI2-THOR. Returned observations and execution outcomes update the memory used for later choices. This is simulation, and the reported metrics concern task success and memory-augmented interaction efficiency rather than real-robot validation. Task families and rollout limits are defined by the released episode manifests and evaluator; reproduction should retain those settings.
Checklist
Validation: all 8 existing catalog tests pass with
node --test tests/catalog.test.mjs.npx --yes awesome-lint README.mdreports the same single missing-Git-metadata diagnostic on both source and edited downloaded snapshots, with no content diagnostics. An equivalent HTTP checker covered all four added resource destinations (paper/code/project/dataset), each HTTP 200; it did not recheck the entire historical link collection. Markdown was rendered and the added resource links and metadata inspected.