ROBOTIC MANIPULATION · MEMORY · BENCHMARK

Benchmarking and Enhancing Skill-Level Memory
for Partially Observable Robotic Manipulation

Understanding hidden states can help robots make better decisions.

Yansong Shi · Jiange Yang · Xijie Yang · Shaowei Zhang · Yuhan Zhu · Tao Lu · Limin Wang

15memory-dependent tasks
305.3average frames
12.1average keyframes
375total task variations
62.9%+11.7 pp
average success on HIDE

Frame and keyframe counts average the 15 task means; variations are summed. Success-rate gain is over the strongest baseline.

Seeing is not always knowing.

The same goal and the same observation can require different actions when the histories differ.

“Repeat twice” specifies what to do, but does not reveal how many repetitions are already complete. The manipulation policy must infer this hidden execution state from its own interaction history.

HIDE evaluates the memory needed for these decisions through 15 RLBench tasks. Counts, earlier observations, and execution progress can remain unresolved by current visual and proprioceptive input. Retaining the relevant history helps a policy decide what to do next, and when to stop.

Abstract

Recent advances in robot learning have enabled manipulation policies to perform increasingly diverse tasks and generalize across environments. However, reliable execution often depends on hidden task states that cannot be determined from current observations alone, making interaction history essential. We introduce HIDE, a benchmark for evaluating manipulation memory under partial observability. HIDE comprises 15 tasks covering repetition counting, historical-state recall, and execution-progress tracking, with randomized initial configurations and decision points where similar observations require different actions depending on prior events. We further propose SEEK, a framework combining three complementary memory mechanisms to retain historical evidence and track execution state. Evaluations reveal substantial limitations in existing policies on HIDE, while memory augmentation improves task success in both simulation and real-world experiments. Individual mechanisms benefit some tasks but can degrade others; their combination achieves the highest average success rate on HIDE among the evaluated configurations. These findings highlight the importance of maintaining internal representations of hidden task states and matching memory design to task-specific information requirements.

01Observe

Past views & actions

02Think & Remember

Task-relevant hidden state

03Act

History-informed decisions

Testing memory, revealing blind spots.

Fifteen tasks test memory for counts, past evidence, and execution progress—states the current view may leave unresolved.
Choose demo 1, 2, or 3 to explore a different episode of every task.

BUILT ON RLBENCH
3 DEMOS PER TASK
DEMO VARIATION

Loading demonstrations…

Scripted expert demonstrations, not evaluated policy rollouts. Each clip holds its final frame for 2 seconds before replaying.

Task statistics

Language patterns, trajectory lengths, and variations across the benchmark.

From the appendix · Frames and keyframes are per-task averages.
TaskLanguage templateAvg. framesAvg. keyframesVariationsVariation type
Task means / total15 tasks305.312.1375—

Bracketed terms are variable placeholders; [N] is the requested count. The footer reports equal-weight task means for frames and keyframes, and the total number of task variations.

Giving shape to hidden states.

Recent context, persistent evidence, and explicit progress.
Three complementary designs within a shared manipulation policy.

WCM 01

Windowed Context Memory

A sliding window of recent interactions helps distinguish completed steps from the next action.

RECENT CONTEXT
PAM 02

Persistent Anchor Memory

A protected initial memory preserves earlier evidence alongside the most recent interaction context.

HISTORICAL EVIDENCE
SCM 03

Stage-Counter Memory

An explicit counter updates at predicted stage boundaries to tell visually similar repetitions apart.

EXECUTION PROGRESS
FIG. 01 SEEK: memory-augmented, coarse-to-fine manipulation. Historical visual information and stage progress condition action prediction.
AVERAGE TASK SUCCESS
62.9%

↗ 11.7 percentage points
over SAM2Act+, the strongest
baseline on average.

HIDE evaluates whether a policy can use the past when the present is not enough.

Success rate by method

View the full comparison table
Success rates (%) on HIDE. Overall and category averages.
MethodAverageRepetitionHistorical recallProgress
SEEK62.961.659.268.0
SAM2Act+51.246.446.460.8
RVT242.943.242.443.2
SAM2Act42.741.639.247.2
RVT35.737.638.431.2
GR00T-N1.726.46.449.623.2
π₀22.110.436.020.0
MME18.42.435.217.6
μVLA18.112.828.812.8
π₀.₅14.45.624.812.8
OpenVLA-OFT13.98.027.26.4
OpenVLA1.63.21.60.0

RC: repetition counting; HSR: historical-state recall; EPT: execution-progress tracking. Each HIDE task is evaluated on 25 held-out episodes.

REAL-WORLD EXPERIMENTS

Memory, put to work.

Across counting, stacking, cleaning, and searching, SEEK achieves 89% average success, compared with 47% for SAM2Act+ and 13% for π₀.₅.

Real-world results · Success rates (%)
MethodAvg.(a) Button(b) Cups(c) Desk(d) Chip
π₀.₅13020032
SAM2Act+4732484464
SEEK897610080100

Avg. is the unweighted mean across four tasks. N is the requested number of button presses or cups to stack; M is the total number of cups. The videos illustrate the tasks and do not constitute the full evaluation set. Long still intervals are shortened for viewing.

SELECTED TRIALS

SEEK

Loading real-world footage…

There is more to explore.

Read the research. Explore the benchmark.
Code and data links will be added as they become available.

01 / PAPER ↗

The research

Benchmark, memory design,
and experimental findings.

Read the paper ↗
02 / CODE

The implementation

Training and evaluation code
for memory-aware manipulation.

Coming soon
03 / DATA

The benchmark

Tasks and demonstrations
for hidden-state-dependent skills.

Coming soon

Citation

@misc{shi2026hide,
  title = {Benchmarking and Enhancing Skill-Level Memory for Partially Observable Robotic Manipulation},
  author = {Yansong Shi and Jiange Yang and Xijie Yang and Shaowei Zhang and Yuhan Zhu and Tao Lu and Limin Wang},
  year = {2026},
  eprint = {2609.38886},
  archivePrefix = {arXiv},
  primaryClass = {cs.RO},
  url = {https://arxiv.org/abs/2609.38886}
}

0:00 / —
01 Front
02 Left shoulder
03 Right shoulder

Three synchronized views · 20 FPS · 2 s final-frame hold

About this task

FROM THE APPENDIX

Instruction pattern

REAL-WORLD TRIAL

Enlarged SEEK architecture.