World Action Model

InternW0-Δ

An Embodied World Model Bridging Predictive Dynamics and Actions

Jointly learning visual dynamics and robot actions,
with Causal Imprint capturing action-relevant future changes.

Read the paper and explore code and models.

Robot manipulationProject video

Demonstrations

Real-world
robot manipulation.

Explore demonstrations with grippers
and dexterous hands.

Qualitative demonstrations. Playback-speed annotations, where shown, are part of the supplied footage.

Data curation

Filtering the
training data.

Examples of filtered or masked segments,
grouped by the reason for exclusion.

Robot and human demonstrations are curated and aligned for pretraining. The examples below illustrate four types of data filtering, including segment removal and action-target masking.

Method

World Action Model
with directed attention.

A pretrained video expert and an action expert,
with separate parameters and masked interaction.

Architecture overview. Wan VAE encodes visual observations. T5 and proprioception condition the video expert; task-conditioned VLM features and proprioception condition the action expert. The two experts are coupled through directed attention.
View PDF
Sparse Memory Context (SMC)
Anchor, recent, and current frames preserve episode context and recent execution history without a dense video history.
Task-conditioned scene semantics
A frozen VLM encodes the instruction together with current observations, providing scene-grounded context for the action expert.
Directed attention

Video and action experts exchange information through masked joint attention while retaining their own cross-attention and feed-forward layers. Causal Imprint tokens (Δ) attend to R, C, and Δ; action tokens attend to A, R, C, Δ, and Act. Neither can attend to future video tokens F.

Causal Imprint (CI)

Learnable queries capture action-relevant future changes from recent and current observations. Latent-difference supervision and future-feature alignment shape these representations, which the action expert uses alongside observed visual features.

4D-aware representation distillation

A frozen Track4World teacher provides offline-cached geometry and motion descriptors. A query decoder aggregates clean A/R/C features, and an auxiliary loss distills these priors into the video expert during training.

Experiments

Results across
benchmark settings.

LIBERO-Plus, RoboTwin,
RoboDojo, and EBench.

All axes start at zero. RoboDojo uses a 0–30% range; the other panels use 0–100%.

Open research

Paper, code,
and models.

Explore the paper, code repository, and model collection.

Authors & affiliations

See the paper for the full author list and affiliations.