InternW0
A Foundational Physical World Model
for Efficient Real-World Interactions



98.6%
LIBERO average
Mean success rate across all four task suites.93.12%
RoboTwin 2.0-Full
Average success across clean and randomized settings.75.60%
Clean2Random average
Clean-only post-training, evaluated on both scene settings.16.47 Hz
Model-side policy updates
60.73 ms per action update on RTX 5090D; video prediction runs asynchronously.Introduction
A Quick Overview of InternW0
See InternW0 carry out a multi-stage laboratory workflow before exploring the architecture, training data, and evaluation results.
Model
Asynchronous video–action modeling with contact-aware control
InternW0 uses a mixture-of-transformers backbone with separate video and action experts. Observation-conditioned context routing keeps longer-horizon predictions useful as new visual and physical feedback arrives.

Joint video–action flow matching learns future visual dynamics and continuous control. A frozen Wan VAE encodes video, while a frozen DINOv3 encoder supplies current visual observations. Robot trajectories supervise both experts; egocentric videos train the video pathway without action labels. Domain-specific interfaces and soft prompts support heterogeneous embodiments, with force and tactile channels added during contact-aware post-training.
Data
7,233.5 hours across 25 training domains
Seven datasets combine real-robot trajectories, simulated manipulation, and egocentric laboratory video. A unified 37-dimensional state–action interface and validity masks accommodate different embodiments, while square-root domain sampling balances the training mixture.

| Source | Type | Domains | Hours | Episodes |
|---|---|---|---|---|
| InternData-A1 | Simulated | 6 | 3,494.3 | 568,194 |
| AgibotWorld | Real | 2 | 2,620.0 | 158,380 |
| RoboCOIN | Real | 13 | 438.1 | 59,952 |
| Galaxea | Real | 1 | 314.9 | 15,362 |
| MolmoAct | Real | 1 | 70.3 | 3,424 |
| RoboDojo | Simulated | 1 | 20.5 | 3,465 |
| Source | Hours | Episodes | Supervision |
|---|---|---|---|
| EgoLab | 275.4 | 3,192 | Future-video prediction without robot action labels |
Benchmarks
Strong single-arm manipulation and bimanual generalization
LIBERO evaluates 40 single-arm tasks across spatial, object, goal, and long-horizon suites. RoboTwin 2.0 covers 50 bimanual tasks, with Full and Clean2Random protocols separating in-domain adaptation from generalization to unseen randomized scenes.
| Benchmark | Metric | InternW0 |
|---|---|---|
| Average success rate | 98.6% | |
| Average success rate | 93.12% | |
| Average success rate | 75.60% |
Egocentric video improves transfer

Efficient action updates

Experiments
From everyday manipulation to scientific workflows
Five real-world tasks test instruction following, object discrimination, long-horizon coordination, and contact-aware tool use. Four use a dual-arm platform; quantitative pipetting uses a 20-DoF dexterous hand on a 7-DoF arm with hybrid force–position control.

The real-world suite spans Make Sandwich, Pick Industrial Parts, Sort Tubes, MoF Experiments, and Quantitative Pipetting.
- Make Sandwich
- Pick up two bread slices and a piece of meat, then stack them in the prescribed order.
- Pick Industrial Parts
- Identify industrial parts and place each one into its matching box.
- Sort Tubes
- Follow an instruction specifying the arm, tube color, and destination box.
- MoF Experiments
- Complete 15 ordered subtasks for solution preparation, including pouring, funnel removal, flask transfer, stopper insertion, and stirring.
- Quantitative Pipetting
- Coordinate five stages of dexterous tool use: pickup and reorientation, tip attachment, aspiration, dispensing, and tip ejection and return.
Real-world results
| Method | Make Sandwich · SR | Pick Industrial Parts · SR | Sort Tubes · SR | MoF · Progress | Pipetting · Progress |
|---|---|---|---|---|---|
| π0.5 | 73.3 | 53.1 | 86.7 | 50.2 | 46.7 |
| Fast-WAM | 40.0 | 50.6 | 66.7 | 10.7 | 18.7 |
| InternW0 | 73.3 | 82.7 | 88.9 | 68.4 | 65.3 |
Everyday and industrial manipulation
Dual-arm demonstrations of ordered assembly, part sorting, and language-conditioned tube selection.
Make Sandwich
Pick Industrial Parts
Sort Tubes
MoF Experiments
Close-up views of flask placement, stopper insertion, and stirrer operation within the 15-stage solution-preparation workflow.
Flask placement
Stopper insertion & stirring
Quantitative Pipetting
A dexterous hand coordinates pipette pickup, tip attachment, aspiration, dispensing, and tip ejection with force-aware control.
Quantitative Pipetting
MoF Experiments — subtask results
| Subtask | π0.5 | Fast-WAM | InternW0 |
|---|---|---|---|
| 1 · Pick up funnel from rack | 100.0 | 100.0 | 100.0 |
| 2 · Insert funnel into flask | 100.0 | 46.7 | 100.0 |
| 3 · Pick up graduated cylinder | 100.0 | 13.3 | 93.3 |
| 4 · Pour liquid into flask via funnel | 73.3 | 0.0 | 80.0 |
| 5 · Place graduated cylinder back | 73.3 | 0.0 | 80.0 |
| 6 · Hold flask steady on platform | 73.3 | 0.0 | 80.0 |
| 7 · Pick up funnel from flask | 73.3 | 0.0 | 80.0 |
| 8 · Insert funnel into rack hole | 46.7 | 0.0 | 66.7 |
| 9 · Return to initial position | 46.7 | 0.0 | 66.7 |
| 10 · Place flask onto stirrer | 33.3 | 0.0 | 66.7 |
| 11 · Pick up stopper from rack | 6.7 | 0.0 | 66.7 |
| 12 · Insert stopper into flask | 6.7 | 0.0 | 66.7 |
| 13 · Press left button of stirrer | 6.7 | 0.0 | 26.7 |
| 14 · Press right button of stirrer | 6.7 | 0.0 | 26.7 |
| 15 · Return to initial position | 6.7 | 0.0 | 26.7 |
| Overall progress rate | 50.2 | 10.7 | 68.4 |
Subtask success rates and overall progress rate (%) from Table 6. Overall progress is the fraction of subtasks completed correctly and in order.
Quantitative Pipetting — subtask results
| Subtask | π0.5 | Fast-WAM | InternW0 |
|---|---|---|---|
| 1 · Pickup & reorientation | 93.3 | 46.7 | 86.7 |
| 2 · Tip attachment | 40.0 | 13.3 | 66.7 |
| 3 · Liquid aspiration | 33.3 | 13.3 | 66.7 |
| 4 · Liquid dispensing | 33.3 | 13.3 | 60.0 |
| 5 · Tip ejection & return | 33.3 | 6.7 | 46.7 |
| Overall progress rate | 46.7 | 18.7 | 65.3 |
Subtask success rates and overall progress rate (%) from Table 7. Contact-aware post-training combines force history with visual and proprioceptive observations.
Contact-aware manipulation

Citation
BibTeX
@techreport{internw0,
title = {InternW0: A Foundational Physical World Model for Efficient Real-World Interactions},
author = {{Physical Intelligence Team, Shanghai AI Laboratory}},
institution = {Shanghai AI Laboratory},
year = {2026},
url = {https://internrobotics.github.io/internw0/}
}