Skip to main content
LLMgram · AI News · 2026-09-13

GPT-6 Astra completes 7 of 100 StationeryBench dual-arm tasks as MolmoAct2 scores zero

GPT-6 Astra completes 7 of 100 StationeryBench dual-arm tasks as MolmoAct2 scores zero

Early results on StationeryBench, a dual-arm robotics benchmark, suggest GPT-6 Astra may represent a meaningful advance in embodied spatial reasoning. The Decoder reports that Astra completed seven of one hundred tasks while MolmoAct2 failed to finish any, and Astra's median progress score reached forty-six compared with twelve for the competitor. A researcher quoted in the coverage described the gap as a step change in spatial reasoning, a domain where general multimodal models have often lagged specialized robotics stacks. For teams weighing foundation models against purpose-built spatial systems, the headline numbers are directionally striking even though absolute success remains low. Treat the signal as provisional: only seven full completions out of one hundred trials leaves room for variance, replication, and broader competitor comparisons before production conclusions.

Sources

GPT-6 Astra completes 7 of 100 StationeryBench dual-arm tasks as MolmoAct2 scores zero

GPT-6 Astra completes 7 of 100 StationeryBench dual-arm tasks as MolmoAct2 scores zero

On StationeryBench, the model completed 7 out of 100 tasks with dual-arm robots, while competitor MolmoAct2 couldn't finish a single one. Astra's median progress score hit 46 out of 100, MolmoAct2 managed 12.

Key takeaway

On StationeryBench, GPT-6 Astra's seven full dual-arm task completions against MolmoAct2's zero mark an early sign that general models may be closing spatial gaps with robotics specialists.

What happened

The Decoder reports early StationeryBench results for GPT-6 Astra on dual-arm robotics tasks, where the model finished seven of one hundred trials while MolmoAct2 completed none.

Coverage cites a median progress score of forty-six for Astra versus twelve for MolmoAct2, and a researcher characterizes the performance gap as a step change in spatial reasoning.

Evidence

  • GPT-6 Astra completed seven of one hundred StationeryBench dual-arm tasks while MolmoAct2 completed none.

    The Decoder · attributed

    On StationeryBench, the model completed 7 out of 100 tasks with dual-arm robots, while competitor MolmoAct2 couldn't finish a single one.

  • Astra's median progress score on StationeryBench was forty-six compared with twelve for MolmoAct2.

    The Decoder · attributed

    Astra's median progress score hit 46 out of 100, MolmoAct2 managed 12.

  • A researcher described the result as a step change in spatial reasoning.

    The Decoder · attributed

    A researcher calls it a "step change in spatial reasoning."

Why it matters

Robotics and embodied AI teams may need to re-evaluate investments in specialized spatial stacks if general foundation models keep improving on manipulation benchmarks.

Limits and uncertainties

Only seven of one hundred tasks were fully completed, leaving a very small success sample.

The comparison featured MolmoAct2 alone and may not reflect the full state of the art.

Early benchmark results often lack independent replication across diverse physical environments.

Practical implications

Builders should pilot general multimodal models on embodied manipulation workflows before committing to specialized spatial-only stacks.

Teams should demand reproducible StationeryBench runs and broader competitor baselines before changing production robotics architecture.

What to watch

Independent replication of StationeryBench scores for GPT-6 Astra and additional robotics model baselines beyond MolmoAct2.

Whether Astra's median progress gains translate into higher full-task completion rates on larger trial sets.

Sources

LLMgram editorial selection and synthesis · @llmgram. LLMgram is not the original publisher of this information.
Continue on LLMgram: Open in AI Signal →
Original reporting: GPT-6 Astra appears to show a "step change" in spatial reasoning based on early benchmarks