Generalizes across simulation and real-world settings.
Evaluation spans seen tasks, randomized scenes, and novel real-world environments.
| Benchmark | Split | Success (%) |
|---|---|---|
| LIBERO | seen | 98.8 |
| LIBERO-Plus | unseen | 85.8 |
| RoboTwin — clean | seen | 85.8 |
| RoboTwin — randomized | unseen | 75.7 |
| RoboCasa365 | seen | 48.1 |
| Real robot — novel environments | unseen | 70.0 |
Full comparisons against prior work are in the paper.
A single goal says where. A chain says how.
Dense video rollouts provide detailed guidance but are expensive; action-only policies lack an explicit long-horizon visual plan. ProWAM instead predicts a sparse chain of visual sub-goals indexed by relative progress r and conditions each action chunk on that plan.










pick the kettle, place it on a burner






t frame index · r relative progress, re-anchored to the latest observation every replan
attention mask · query → key
Every new observation re-anchors the plan.
At each replanning step, the latest observation becomes r = 0, and ProWAM predicts a fresh chain of sub-goals. The visualization advances automatically — drag the slider to inspect individual sub-goals, or select a replanning step.
“turn on the stove and put the moka pot on it”
“use both arms to place the burger and the fries on the green tray”
“grab the hamburger and the fries carton, then set them on the green tray”
“move the red and green blocks to the center and place green atop red”
“move the red block to the center, then put the green block on top”
“pick up the bowls on the counter and stack them on top of one another in the open cabinet. Place the smaller bowl on top of the larger bowl”
“pick up the knife from the drawer and place it on the cutting board. Then place the meat from the plate to the cutting board”
“grab a lemon wedge from the fridge and one ice cube from the ice bowl, and put them in the glass of lemonade”
Zero-shot on a real robot in novel environments.
Rollouts on a physical arm, sped up to fit. Exterior and wrist cameras play in sync.
No action labels. No target-domain data. Still a visual plan.
Stage 1 trains on instruction-conditioned video only. Given one frame and an instruction from LIBERO or RoboTwin, it imagines the scene at each progress point r.
r = 0
r = 0
r = 0
r = 0
r = 0
r = 0
r = 0
r = 0
r = 0
r = 0
r = 0
r = 0
r = 0
r = 0
r = 0
r = 0Every tile is out of distribution · first frame at r = 0, then imagined milestones at r = 0.1 … 0.9 · hover to pause
@article{zhang2026prowam,
title = {World Action Modeling with Progressive Visual Planning},
author = {Zhang, Fei and An, Zhaochong and Frost, Duncan and Wang, Yikai
and Liu, Pengfei and Zhang, Ya and Drozdzal, Michal and Bar, Amir},
journal = {arXiv preprint arXiv:2610.02508},
year = {2026}
}