Windowed Context Memory
A sliding window of recent interactions helps distinguish completed steps from the next action.
Understanding hidden states can help robots make better decisions.
Remember the repetition count.
Keep the earlier task context.
Track the execution progress.
Expert demonstrations in RLBench Three complementary demands on memory
Frame and keyframe counts average the 15 task means; variations are summed. Success-rate gain is over the strongest baseline.
Seeing is not always knowing.
The same goal and the same observation can require different actions when the histories differ.
“Repeat twice” specifies what to do, but does not reveal how many repetitions are already complete. The manipulation policy must infer this hidden execution state from its own interaction history.
HIDE evaluates the memory needed for these decisions through 15 RLBench tasks. Counts, earlier observations, and execution progress can remain unresolved by current visual and proprioceptive input. Retaining the relevant history helps a policy decide what to do next, and when to stop.
Recent advances in robot learning have enabled manipulation policies to perform increasingly diverse tasks and generalize across environments. However, reliable execution often depends on hidden task states that cannot be determined from current observations alone, making interaction history essential. We introduce HIDE, a benchmark for evaluating manipulation memory under partial observability. HIDE comprises 15 tasks covering repetition counting, historical-state recall, and execution-progress tracking, with randomized initial configurations and decision points where similar observations require different actions depending on prior events. We further propose SEEK, a framework combining three complementary memory mechanisms to retain historical evidence and track execution state. Evaluations reveal substantial limitations in existing policies on HIDE, while memory augmentation improves task success in both simulation and real-world experiments. Individual mechanisms benefit some tasks but can degrade others; their combination achieves the highest average success rate on HIDE among the evaluated configurations. These findings highlight the importance of maintaining internal representations of hidden task states and matching memory design to task-specific information requirements.
Past views & actions
Task-relevant hidden state
History-informed decisions
Testing memory, revealing blind spots.
Fifteen tasks test memory for counts, past evidence, and execution progress—states the current view may leave unresolved.
Choose demo 1, 2, or 3 to explore a different episode of every task.
Scripted expert demonstrations, not evaluated policy rollouts. Each clip holds its final frame for 2 seconds before replaying.
Language patterns, trajectory lengths, and variations across the benchmark.
| Task | Language template | Avg. frames | Avg. keyframes | Variations | Variation type |
|---|---|---|---|---|---|
| Task means / total | 15 tasks | 305.3 | 12.1 | 375 | — |
Bracketed terms are variable placeholders; [N] is the requested count. The footer reports equal-weight task means for frames and keyframes, and the total number of task variations.
Giving shape to hidden states.
Recent context, persistent evidence, and explicit progress.
Three complementary designs within a shared manipulation policy.
A sliding window of recent interactions helps distinguish completed steps from the next action.
A protected initial memory preserves earlier evidence alongside the most recent interaction context.
An explicit counter updates at predicted stage boundaries to tell visually similar repetitions apart.
↗ 11.7 percentage points
over SAM2Act+, the strongest
baseline on average.
HIDE evaluates whether a policy can use the past when the present is not enough.
| Method | Average | Repetition | Historical recall | Progress |
|---|---|---|---|---|
| SEEK | 62.9 | 61.6 | 59.2 | 68.0 |
| SAM2Act+ | 51.2 | 46.4 | 46.4 | 60.8 |
| RVT2 | 42.9 | 43.2 | 42.4 | 43.2 |
| SAM2Act | 42.7 | 41.6 | 39.2 | 47.2 |
| RVT | 35.7 | 37.6 | 38.4 | 31.2 |
| GR00T-N1.7 | 26.4 | 6.4 | 49.6 | 23.2 |
| π₀ | 22.1 | 10.4 | 36.0 | 20.0 |
| MME | 18.4 | 2.4 | 35.2 | 17.6 |
| μVLA | 18.1 | 12.8 | 28.8 | 12.8 |
| π₀.₅ | 14.4 | 5.6 | 24.8 | 12.8 |
| OpenVLA-OFT | 13.9 | 8.0 | 27.2 | 6.4 |
| OpenVLA | 1.6 | 3.2 | 1.6 | 0.0 |
RC: repetition counting; HSR: historical-state recall; EPT: execution-progress tracking. Each HIDE task is evaluated on 25 held-out episodes.
Memory, put to work.
Across counting, stacking, cleaning, and searching, SEEK achieves 89% average success, compared with 47% for SAM2Act+ and 13% for π₀.₅.
| Method | Avg. | (a) Button | (b) Cups | (c) Desk | (d) Chip |
|---|---|---|---|---|---|
| π₀.₅ | 13 | 0 | 20 | 0 | 32 |
| SAM2Act+ | 47 | 32 | 48 | 44 | 64 |
| SEEK | 89 | 76 | 100 | 80 | 100 |
Avg. is the unweighted mean across four tasks. N is the requested number of button presses or cups to stack; M is the total number of cups. The videos illustrate the tasks and do not constitute the full evaluation set. Long still intervals are shortened for viewing.
Loading real-world footage…
Choose a numbered clip under each method. Empty slots await footage. Playback speed applies to all real-world videos; each clip stays on its final frame when it ends.
There is more to explore.
Read the research. Explore the benchmark.
Code and data links will be added as they become available.
Benchmark, memory design,
and experimental findings.
Training and evaluation code
for memory-aware manipulation.
Tasks and demonstrations
for hidden-state-dependent skills.
@misc{shi2026hide,
title = {Benchmarking and Enhancing Skill-Level Memory for Partially Observable Robotic Manipulation},
author = {Yansong Shi and Jiange Yang and Xijie Yang and Shaowei Zhang and Yuhan Zhu and Tao Lu and Limin Wang},
year = {2026},
eprint = {2609.38886},
archivePrefix = {arXiv},
primaryClass = {cs.RO},
url = {https://arxiv.org/abs/2609.38886}
}