“Welcome to the real world.”
01 Benchmark Results
We report success rate (SR) alongside Score, which credits partial completion at the end of an episode. Both metrics are averaged equally across tasks.
GPT-6-Astra + Single-shot In-Context Learning (ICL)
Task-averaged success rate
Comparison across models
| Model | ||
|---|---|---|
| 1OpenWAM-α | 55.3 | 0.701 |
| 2AMapbot1 | 48.9 | 0.639 |
| 3GPT-6-Astra | 46.7 | 0.654 |
| 4Qwen-RobotManip | 45.6 | 0.608 |
| 5π₀.₅ | 41.4 | 0.544 |
| 6InternVLA-A1.5 | 34.2 | 0.468 |
| 7π₀ | 33.7 | 0.475 |
| 8GigaBrain-0.7 | 33.3 | 0.460 |
| 9Fast-WAM | 25.6 | 0.371 |
02 Experiment Setup: Benchmark composition and evaluation protocol
Evaluation Tasks
EBench: generalization under controlled variations
Our main evaluation covers all 26 tasks in the EBench test split, spanning nine scene categories across household, retail, and industrial settings. These include 19 mobile and 7 tabletop tasks, with task-level annotations of mobility, manipulation precision, and task horizon. We evaluate generalization under controlled variations in objects, backgrounds, instructions, and their combinations. Across the 510 evaluated episodes, we measure task success rate (SR) and Score, which captures normalized partial completion at the end of an episode. Both metrics are first averaged within each task and then averaged equally across the 26 tasks.
Beyond EBench: generalization to task compositions
We evaluate generalization to an unseen task and the ability to compose it with an already-mastered task in sequence. The compositional task asks to place a bookmark and then a pen on the same open book, combining the familiar "Bookmark on Book" task with a new pairing, "Pen on Book". The compositional evaluation setting details the training-task coverage, prescribed sequence, and scoring protocol.
Model Adaptation and Evaluation Protocol
We compare GPT-6-Astra with seven post-trained Vision-Language-Action (VLA) and World-Action Model (WAM) policies using their reported EBench leaderboard results. These policies are trained on the public EBench training dataset, which contains 6,600 episodes across 26 tasks. GPT-6-Astra receives one annotated training demonstration per task as in-context input, without parameter updates.
To construct this context, a separate GPT-6-Astra instance, configured with high reasoning effort, reviews one video trajectory from the training set and autonomously selects and annotates 10–18 keyframes. The annotations are paired with the corresponding robot states and recorded actions to form the demonstration context. All evaluation episodes of the same task reuse this context across seeds. During execution, GPT-6-Astra uses the demonstration to inform decisions based on current observations, rather than replaying the recorded coordinates.
Physical execution in the main evaluation follows the same task-specific physics-step limits as the baseline policies. We impose no additional aggregate cap on GPT-6-Astra's tokens, tool calls, or inference time. GPT-6-Astra may issue a termination signal before the execution limit; the server determines the final task success and Score. If the policy stops early, the runner holds the robot and advances the simulator to obtain that terminal result.
Observations and Direct Robot Control
Inputs: camera views and robot state
GPT-6-Astra serves directly as the control agent with high reasoning effort. At each decision step, it receives RGB observations and the current robot state, including end-effector (EEF) poses, gripper state, base pose, and the simulator timestep. The initial observation includes images from the overview, left wrist, and right wrist cameras; subsequent tool calls return the camera views requested by the model.
The model receives no privileged information like depth data, ground-truth object poses, camera calibration, segmentation, intermediate task scores, or subgoal-completion signals. Current elapsed simulation steps are reported, but the maximum physics-step budget remains hidden.
Outputs: Actions through a fixed interface
Given the current EEF pose, gripper state, and base position, GPT-6-Astra adaptively outputs one or more EEF waypoints, together with gripper and base targets and execution durations. The waypoints are interpolated into an action chunk and executed open-loop before the next observation is returned. No post-trained VLA or WAM generates or refines GPT-6-Astra's actions.
The interaction follows an observation–decision–execution–feedback loop. After the trajectory finishes executing, the control interface returns the updated proprioceptive observations and the requested camera views for the next decision. The model does not receive intermediate observations or make new decisions between the interpolated segments within that chunk.
Execution Demo
Loading episode interactions…
03 Score and Ranking
04 General Analysis: Investigating the performance gap
Failure Patterns: Where execution breaks down
Under the EBench protocol, scoring is not binary; instead, an episode receives partial credit based on the intermediate stages successfully completed, with a full Score of 1.0 assigned only upon total task completion. Across 510 evaluated episodes, 237 achieve complete success, while 188 of the 273 unsuccessful runs still register positive partial credit (with only 85 scoring zero). For episodes that gain partial credit but fall short of full success, this score breakdown points to two common failure modes: either the execution lacks the precision required to satisfy strict final termination criteria, or the agent stalls at an intermediate step and times out before completing the task.
This behavioral difference is evident when comparing tasks with similarly low success rates: failed (zero-score) episodes make up 75% of gear installation, 65% of utensils placement and 50% of glasses packing, whereas they are completely absent in nut tightening, peg insertion, and shop (and account for only 30% in bottle). In these latter tasks, the agent advances through early stages despite missing full completion, a distinction that would be lost when looking at final success alone.
The breakdown below splits each task's episodes into complete successes, incomplete runs and failures, with each row tagged by the precision and horizon the task demands.
These terminal scores motivate examining where precise contact fails and how local retries influence overall task progress. The recordings below connect these execution challenges directly to the reported task results.
Qualitative analysis from the recordings
The gap in tabletop dexterous-and-precise tasks
All seven tabletop tasks in EBench demand dexterous or precise manipulation, which we refer to as tabletop dexterous-and-precise (D&P) tasks: frame placement, cup flipping and glasses packing through two-arm coordination, flipping and folding; the remaining four through tight contact. Even on the three D&P tasks labeled low or medium precision, GPT-6-Astra averages 31.7% success, compared with Qwen-RobotManip’s 91.7%. Cup flipping reaches 10% versus 90%; glasses packing reaches 20% versus 90%.
Precision labels alone do not explain the capability gap. The recordings point to two failures that the labels do not capture. Glasses packing is labeled medium precision, yet closing the case depends on folding the temples precisely: GPT-6-Astra completes the transfer into the case, but its corrective contacts leave the temples protruding, so the lid cannot close.
The second is bimanual coordination. Without a demonstration, the agent does not compose a two-arm plan for frame placement and attempts single-arm variants that leave the frame unsupported; with one, it divides the work between the arms (see the frame ICL case) and reaches 65% in the main evaluation, above OpenWAM-α’s 50%. Cup flipping is a purely bimanual problem that the demonstration does not resolve: GPT-6-Astra reaches 10% against Qwen-RobotManip’s 90%. Across the five tasks we identify as involving a handover (our annotation, not an EBench label), GPT-6-Astra averages 34.0% against OpenWAM-α’s 48.0%, placing fifth among the eight systems evaluated.
Neither failure shows up in the success rate on its own, and neither is named by a single task label. The next step is therefore to take the annotations EBench does supply — mobility, horizon and precision — and ask how much of this deficit they account for.
Capability Differences: Ahead on short-horizon tasks, behind on tabletop D&P and long tasks
Taking those annotations in turn gives a first account of where the deficit sits. Averaged within the benchmark's own task groups, GPT-6-Astra has the highest success rate of the eight models on the 11 short-horizon tasks (73.2%, against OpenWAM-α’s 65.9%), but ranks sixth on the 15 long-horizon tasks (27.3%, against OpenWAM-α’s 47.6%) and sixth on the seven tabletop D&P tasks (20.0%, against Qwen-RobotManip’s 50.0%). The lead therefore comes from short procedures, eight of the eleven of which demand only low precision; once a task runs long or requires dexterous, precise contact, the post-trained policies stay ahead.
Task-level strengths and weaknesses
Adaptive Capabilities: Exploration, recovery and self-correction
The apple interaction is consistent with experience use within an episode: the agent states a possible failure cause, changes its transport action, and eventually completes the task. Recovery can also fail while extending the robot’s movement through the scene, as the safety cases show.
Safety in Execution: Risk behaviors in the recorded trajectories
Exploration and recovery also expose safety weaknesses. Unsuccessful dishwasher and apple-to-fruit-bowl episodes show objects slipping from the gripper, followed by repeated arm and base repositioning to search for them. During exploration, the agent keeps requesting EEF targets after an object has fallen outside its workspace. The coffee-bean case shows awkward EEF poses and interference between held objects and the scene. Together, these failures highlight the need to preserve grasp stability, enforce workspace constraints and account for contact throughout manipulation.
05 Case Studies
The selected comparisons illustrate two observations: in-context learning (ICL) can guide task-specific manipulation, and task-level adaptation does not guarantee precise execution. The teacup and glasses recordings reveal complementary capabilities: GPT-6-Astra revises grasps and recovers disrupted goals, while post-trained policies execute fine manipulation more accurately in these examples.
In-Context Learning
Adaptation and Precision
More recorded casesSelected GPT-6-Astra rollouts across all 26 tasks
06 Compositional Generalization
Evaluation Setting
Beyond EBench, we further evaluate a compositional unseen task that requires placing a bookmark and a pen on the same open book sequentially. Bookmark on book and pen-placement tasks with other targets (in a pen holder or on a ruler) appear in the robot-policy training set, whereas pen on book does not. This setting therefore tests more than the sequencing of two familiar tasks: the combined objective includes a sub-task absent from the robot-policy training data. Each method is evaluated over 10 rollouts and GPT-6-Astra receives no in-context demonstrations. The prescribed sequence places the bookmark first and then the pen, and a partial credit of 0.5 is awarded for completing the first placement in that order.
Compositional Task Execution
07 Conclusions and Future Works
The evaluation places GPT-6-Astra’s capabilities between flexible task interpretation and reliable physical execution. With a single annotated demonstration per task, it ranks third on the leaderboard, and second among the eight evaluated models. The agent leads on average across short-horizon tasks, while trailing on average in the tabletop dexterous-and-precise (D&P) and long-horizon groups. Qualitatively, the execution recordings illustrate the model attempting alternative approaches, recovering after a disrupted subgoal, and using a failed attempt to inform its subsequent action.
Meanwhile, precise alignment and physical contact remain clear execution bottlenecks, where leading post-trained policies still hold a considerable edge. The logs also point to a fundamental control trade-off: while the agent selects corrections at the boundaries of action chunks, physical contact and grasp stability often evolve within the chunk itself. In practice, large end-effector adjustments and awkward object poses raise legitimate safety concerns for real-world deployment. These observations point toward several practical directions.
Combine post-trained policies with agents
The comparisons suggest complementary strengths between frontier reasoning agents and low-level controllers. GPT-6-Astra’s semantic planning and recovery provide ways to navigate unfamiliar states that often challenge demonstration-trained policies. Conversely, post-trained controllers deliver the responsive, precise execution required for contact-sensitive operations. A natural next step is exploring hybrid setups where the agent handles exploration and goal revision, while a dedicated local controller maintains high-frequency feedback for delicate contacts.
Making this coordination work in practice depends on clear handoff protocols. The agent might propose candidate goals and contact assumptions, while the local controller monitors for slip or unexpected resistance, interrupting execution to return control when needed. Frameworks like RPent [1], which treat learned VLA skills as callable tools for a planner, illustrate one promising architecture; assessing how well such designs hold up under unexpected scene changes and contact constraints remains an open question. Yet, how to effectively combine a strong agent with post-trained models hasn’t been fully explored.
Teach smaller policies to plan, explore, and recover
A second direction involves transferring the agent’s corrective strategies to smaller on-device models. The observed rollouts capture meaningful exploratory behaviors, such as re-orienting for a better view or adjusting an approach angle after a misstep. Training setups could explicitly pair initial goals and failure states with revised strategies and outcomes—incorporating both failed corrections and successful recoveries. This could help a policy learn when it makes sense to retry versus when an approach should be abandoned.
This approach could essentially broaden behavior cloning toward intention imitation, focusing on matching corrective strategies to observed states. The main practical hurdle is preserving this reasoning across different scenes and embodiments without simply overfitting to specific demonstration trajectories or compounding execution errors. If issues around transfer fidelity, safety, and collection overhead can be managed, agent-generated exploration could offer a helpful complement to purely human demonstrations.
Ground exploratory experience into reusable tools
The apple recovery episode illustrates how interaction feedback can inform subsequent actions, but scaling this into systematic recursive self-improvement (RSI) requires deliberate design. Rather than relying on ephemeral context histories or rigid hardcoded assertions, the core challenge lies in distilling trial-and-error rollouts into efficient, structured knowledge. Frameworks like Show-Harness [2] and RPent [1] demonstrate the viability of exposing learned behaviors as modular interfaces; building on this, an RSI loop must actively curate what to keep. It needs to distinguish robust, generalizable recovery strategies from accidental successes or brittle environment hacks that happen to clear a single trial. By leveraging the agent’s emergent reasoning to synthesize, verify, and catalog these validated strategies into reusable tools, the system can continuously expand its behavioral library without human-in-the-loop engineering. The critical test for such an architecture is whether these accumulated capabilities genuinely transfer across unseen tasks—consistently reducing retries and improving execution efficiency rather than overfitting to specific historical failures.
Treat safety as an active constraint in execution
Task success alone cannot fully satisfy physical deployment. Successfully retrieving an object after a drop is encouraging, but in real-world environments, the drop itself may have already damaged the hardware or workspace. Large trajectory corrections, blind gripper closures, and uncertain contact heights all demand active runtime constraints. Rather than relying solely on post-hoc error recovery, the system needs low-level guardrails—such as dynamic force and velocity limits, clearance verification, and reflexive interrupts—integrated directly into the action pipeline. At the same time, the agent must be sensitive to perceptual ambiguity, actively pausing to adjust camera viewpoints or request human intervention when geometry is uncertain.
Ultimately, these findings suggest that high-level task planning, exploration, and error recovery cannot be developed in isolation from physical control. Moving forward, the key challenge is balancing the agent’s behavioral flexibility with the precision, responsiveness, and safety demanded by physical reality—while finding practical ways to transfer the experience gained in one task to the next.
09 References
- RLinf. RPent: Architecture. Online documentation, 2026. Available: rpent.readthedocs.io/en/latest/rst_source/development/architecture.html.
- Yanzhe Chen, Zechen Bai, Zhijun Cao, et al. Show-Harness: Just a VLM Agent Can Play Robots. arXiv preprint arXiv:2609.10522, 2026. Available: arxiv.org/abs/2609.10522.
- Tianxing Chen, Yue Chen, Zixuan Li, et al. RoboDojo: A Unified Sim-and-Real Benchmark for Comprehensive Evaluation of Generalist Robot Manipulation Policies. arXiv preprint arXiv:2607.04434, 2026. Available: arxiv.org/abs/2607.04434.
- Wenbo Zhang, Kaixuan Wang, Yutao Ouyang, et al. An Unexpected Robot Policy: Early Evaluations of GPT-6 Astra on RoboDojo and Beyond. Technical report, 2026. Available: robodojo-benchmark.com/report/gpt-6-astra-eval.
- Jiayi Su, Yixin Zheng, Mi Yan, Li Yi, Zhizheng Zhang, and He Wang. GPT 6 Astra as an Embodied Policy. Technical report and code, 2026. Available: anonymous-report-421.github.io/public-website.
- Anonymous. PhysEvo: Astra can act, let it. Research preview, 2026. Available: anonymous-report-777.github.io.
- Fangcheng Liu, Yeqing Shen, Anda Cheng, et al. Embodied In-Context Learning for GPT-6 Astra. Research preview, 2026. Available: mosi-ai.github.io/RoboICL-GPT6-Astra.github.io.
- Pantograph. Benchmarking Frontier LLMs on Robots. Online report, 2026. Available: pantograph.com/journal/vlm-harness.
- Zihan Jack Zhang, Sravanthi Machcha, Sabrina Zou, Tzu Kit Chan, and Jay Chooi. GPT-6 Astra vs MolmoAct2 on bimanual robotic manipulation. Online report, 2026. Available: openai.robocurve.org/stationerybench.
- Zeyu Shen, Haoxiang You, Yilang Liu, et al. EmbodiedSWE: Coding Agents for Long-Horizon Dexterous Robotics. Project page, 2026. Available: embodiedswe.github.io.

