ProWAM

World Action Modeling with Progressive Visual Planning

1Shanghai Jiao Tong University   2SII   3Meta   4Imperial College London
ProWAM first imagines progress-indexed sub-goals, then executes actions grounded in them

Imagine the milestones, then act toward them.

0
LIBERO-Plus failuresvs best prior · 85.8% SOTA
0
RoboTwin randomizedover best prior · 75.7% absolute
0
Real robotunseen environments · 70.0%
0
RoboCasa365of 14 on the leaderboard · 48.1% overall
Results

Generalizes across simulation and real-world settings.

Evaluation spans seen tasks, randomized scenes, and novel real-world environments.

BenchmarkSplitSuccess (%)
LIBEROseen98.8
LIBERO-Plusunseen85.8
RoboTwin — cleanseen85.8
RoboTwin — randomizedunseen75.7
RoboCasa365seen48.1
Real robot — novel environmentsunseen70.0

Full comparisons against prior work are in the paper.

Method

A single goal says where. A chain says how.

Dense video rollouts provide detailed guidance but are expensive; action-only policies lack an explicit long-horizon visual plan. ProWAM instead predicts a sparse chain of visual sub-goals indexed by relative progress r and conditions each action chunk on that plan.

∅
Zero imaginationfast, no foresight
Final goalwhere, not how
ProWAMsparse, progress-indexed
Full imaginationdense, slow
X1
X2
X3
G1
G2
Video Expert ℒvideo + ℒsub 1 forward pass KV cache
T
Iref
X1
X2
X3
G1
G2
promptfirst frame video latent sub-goal latent
pick the kettle, place it on a burner
current observation
t=0 · r=0
t=1
t=2
t=3
sub-goal at r=0.4
r=0.4
sub-goal at r=0.8
r=0.8
task timeline

t frame index  ·  r relative progress, re-anchored to the latest observation every replan

A0
A1
A2
Action Expert ℒact ×T steps
A0
A1
A2
action chunk
X0XGA X0 X G A

attention mask · query → key

Closed-loop replanning

Every new observation re-anchors the plan.

At each replanning step, the latest observation becomes r = 0, and ProWAM predicts a fresh chain of sub-goals. The visualization advances automatically — drag the slider to inspect individual sub-goals, or select a replanning step.

LIBEROunseenK = 1

“turn on the stove and put the moka pot on it”

observation
observe
imagined sub-goal
imagine
execution
execute
replan
RoboTwinseenK = 4

“use both arms to place the burger and the fries on the green tray”

observation
observe
imagined sub-goal
imagine
execution
execute
replan
RoboTwinrandomizedK = 4

“grab the hamburger and the fries carton, then set them on the green tray”

observation
observe
imagined sub-goal
imagine
execution
execute
replan
RoboTwinseenK = 4

“move the red and green blocks to the center and place green atop red”

observation
observe
imagined sub-goal
imagine
execution
execute
replan
RoboTwinrandomizedK = 4

“move the red block to the center, then put the green block on top”

observation
observe
imagined sub-goal
imagine
execution
execute
replan
RoboCasa365seenK = 6

“pick up the bowls on the counter and stack them on top of one another in the open cabinet. Place the smaller bowl on top of the larger bowl”

observation
observe
imagined sub-goal
imagine
execution
execute
replan
RoboCasa365seenK = 6

“pick up the knife from the drawer and place it on the cutting board. Then place the meat from the plate to the cutting board”

observation
observe
imagined sub-goal
imagine
execution
execute
replan
RoboCasa365unseenK = 6

“grab a lemon wedge from the fridge and one ice cube from the ice bowl, and put them in the glass of lemonade”

observation
observe
imagined sub-goal
imagine
execution
execute
replan
observationimagined sub-goalexecution
Real world

Zero-shot on a real robot in novel environments.

Rollouts on a physical arm, sped up to fit. Exterior and wrist cameras play in sync.

exterior
wrist
“stack the bowls”18×
exterior
wrist
“stack the blue cube on the red cube”11×
exterior
wrist
“put the banana in the box”16×
exterior
wrist
“put the orange on the right side of the apple”6×
Stage-1 imagination

No action labels. No target-domain data. Still a visual plan.

Stage 1 trains on instruction-conditioned video only. Given one frame and an instruction from LIBERO or RoboTwin, it imagines the scene at each progress point r.

RoboTwin: imagined progress for “grab the synthetic leather white shoe from the table with the right arm and place it on the mat” r = 0
RoboTwingrab the synthetic leather white shoe from the table with the right arm and place it on the mat
LIBERO: imagined progress for “put the bowl on the stove” r = 0
LIBEROput the bowl on the stove
RoboTwin: imagined progress for “take both arms to grasp the light wood roller” r = 0
RoboTwintake both arms to grasp the light wood roller
LIBERO: imagined progress for “put the wine bottle on the rack” r = 0
LIBEROput the wine bottle on the rack
LIBERO: imagined progress for “open the top drawer and put the bowl inside” r = 0
LIBEROopen the top drawer and put the bowl inside
RoboTwin: imagined progress for “place red block and green block at the center and then put green block on red block” r = 0
RoboTwinplace red block and green block at the center and then put green block on red block
LIBERO: imagined progress for “put the bowl on the plate” r = 0
LIBEROput the bowl on the plate
RoboTwin: imagined progress for “grasp the mesh-patterned microphone head and pass it across” r = 0
RoboTwingrasp the mesh-patterned microphone head and pass it across
RoboTwin: imagined progress for “make sure the cylindrical light blue cup ends up on the flat round wooden coaster” r = 0
RoboTwinmake sure the cylindrical light blue cup ends up on the flat round wooden coaster
LIBERO: imagined progress for “put the wine bottle on top of the cabinet” r = 0
LIBEROput the wine bottle on top of the cabinet
RoboTwin: imagined progress for “hold the ceramic bowl with earthy brown trim steady and slide the medium-sized bowl for holding food on” r = 0
RoboTwinhold the ceramic bowl with earthy brown trim steady and slide the medium-sized bowl for holding food on
LIBERO: imagined progress for “put the cream cheese in the bowl” r = 0
LIBEROput the cream cheese in the bowl
LIBERO: imagined progress for “push the plate to the front of the stove” r = 0
LIBEROpush the plate to the front of the stove
RoboTwin: imagined progress for “use the right arm, the left arm, and the left arm to center large block, medium block, and small block in order” r = 0
RoboTwinuse the right arm, the left arm, and the left arm to center large block, medium block, and small block in order
LIBERO: imagined progress for “turn on the stove” r = 0
LIBEROturn on the stove
RoboTwin: imagined progress for “grab the smooth red can, drop it into the yellow basket for carrying stuff, then use another arm for the yellow basket for carrying stuff” r = 0
RoboTwingrab the smooth red can, drop it into the yellow basket for carrying stuff, then use another arm for the yellow basket for carrying stuff

Every tile is out of distribution · first frame at r = 0, then imagined milestones at r = 0.1 … 0.9 · hover to pause

Citation
@article{zhang2026prowam,
  title   = {World Action Modeling with Progressive Visual Planning},
  author  = {Zhang, Fei and An, Zhaochong and Frost, Duncan and Wang, Yikai
             and Liu, Pengfei and Zhang, Ya and Drozdzal, Michal and Bar, Amir},
  journal = {arXiv preprint arXiv:2610.02508},
  year    = {2026}
}