Skip to content
OpenAIAINewsRobotics

GPT-6 Astra Scores 95% in Robocurve Robot Control Test

GPT-6 Astra hit 95% on Robocurve's robot control task against Claude Fable 5.1's 40%, using 6.2x fewer output tokens at 2.3x lower cost.

Table of Contents

GPT-6 Astra scored 95% on a Robocurve robot control task, outperforming Claude Fable 5.1's 40% result while using fewer output tokens and costing less.

The test moves the frontier-model competition into embodied control, where a model's reasoning has to translate into physical action rather than text, code, or on-screen work. Astra completed the task with 6.2x fewer output tokens and at 2.3x lower cost than Fable 5.1.

On harder precision tasks, Astra and Fable 5.1 reportedly reached the same success rate, completing 2 of 20 trials. Astra still used 3.9x fewer output tokens and was 1.6x cheaper on those attempts, which suggests the larger shift may be efficiency as much as raw task completion.

Robocurve Tests Frontier Models on Physical AI

Robocurve is a public benefit corporation focused on measuring and reporting frontier robotics capabilities. The company describes its work as an attempt to create open, independent evaluations for physical AI at a time when much of the robotics field is still shaped by selective demo videos and narrow lab results.

Its open-source framework, Inspect Robots, runs different policies, including large language model agents and vision-language-action models, against compatible robot embodiments or simulators. It records auditable logs, grader scores, model transcripts, configuration details, and visualizations so runs can be reviewed instead of treated as one-off demonstrations.

Test-time scaling in robotics capabilities

Robotics evaluation is harder than standard software benchmarking. A model can appear capable in a video while still failing under small variations in object placement, lighting, gripper pose, timing, or task wording. Robocurve standardizes how models are tested across real or simulated robots so those differences become measurable.

Inspect Robots also supports multiple embodiments, including bimanual I2RT YAM arms, Franka arms, AgiBot A2 arms, Unitree G1 arms, SO-ARM followers, WidowX hardware, ROS-connected arms, Isaac Lab simulation, and mock environments. That range keeps the benchmark infrastructure from being tied to a single robot or private lab setup.

Astra Shows a Robotics Efficiency Gap

Astra's 95% success rate on the robot control task is the headline number, but the efficiency gap may be the more important technical signal.

Robotics agents have to plan, observe, issue action chunks, respond to failures, and adjust to feedback. A model that reaches the same or better result with fewer output tokens can reduce latency and cost, which makes real-time or near-real-time control more practical. In physical AI, delays can decide whether a robot responds smoothly enough to finish a task, so they matter more than in a typical software interface.

Astra's reported advantage over Fable 5.1 points to two improvements at once: higher task success in the tested setting and more compact action reasoning. The cost difference also matters because robotics evaluation and deployment can require repeated trials, camera inputs, logs, and long interaction loops. A cheaper model per completed task is easier to test, iterate, and deploy at scale.

OpenAI has also framed GPT-6 Astra as a major step forward in broader computer-use and reasoning tasks, saying the model surpassed its human action-efficiency baseline on 96% of ARC-AGI-3 levels and improved task completion speed in computer-use evaluations. The Robocurve result is separate from those internal claims, but it fits the same theme. Frontier models are increasingly judged by how well they operate tools, software, and physical systems, in addition to how they answer questions.

Precision Manipulation Remains the Bottleneck

The harder precision-task result keeps the benchmark in perspective. Astra matching Fable 5.1 at 2 successful trials out of 20 suggests that fine-grained manipulation remains a major challenge even when a model improves on broader control tasks.

Picking up a large object, moving a block, or completing a structured gross-motor action can differ meaningfully from tasks that require millimeter-level alignment, delicate grasping, force control, or recovery after small mistakes. A small planning error can accumulate into a failed grasp, a collision, or an object that moves out of the expected frame.

The result suggests frontier LLMs may become more useful to developers as high-level robot-control policies, especially when paired with structured tools, action validators, and robotics-specific frameworks. The same result also shows why robotics progress will still depend on better embodiment interfaces, perception systems, low-level control, simulation transfer, and benchmark design.

Robocurve's test adds another data point to a growing question for the AI market: whether the next wave of model competition will be measured by agency in real environments rather than benchmark scores alone. If models can control robots more reliably and cheaply, the practical surface area of AI expands from digital work into physical labor. If they remain brittle on precision tasks, deployment will likely stay concentrated in constrained environments where tooling and task design can absorb model weaknesses.

Frontier models are beginning to show measurable gains in physical AI evaluation, and open frameworks are increasingly judging those gains by exposing both success rates and failure modes.

Comments

Latest