Skip to content

OpenRoboto and Axis Robotics Launch Open Axis Benchmark for Robot Model Evaluation

The living benchmark rotates robot-manipulation tasks from Axis Robotics' growing library to test whether models can generalize beyond fixed evaluation sets.

Table of Contents

OpenRoboto and Axis Robotics have launched Open Axis Benchmark, a living evaluation engine for robot-manipulation models that makes benchmark performance harder to game through overfitting.

The benchmark draws from Axis Robotics' task library and replaces static test suites with versioned rounds of evaluation. Each round locks in a fresh task set, and tasks rotate out when they stop separating models, including when every model clears them, or no model can complete them.

The benchmark is live on the simulation track, with model submissions open through the project's OpenRoboto site. The launch gives Bittensor's SN80 robotics subnet a more formal evaluation layer for comparing physical AI models as teams compete to improve manipulation performance.

0:00
/0:00

How Open Axis Benchmark Works

Open Axis Benchmark addresses a recurring problem in robotics evaluation: fixed benchmarks lose usefulness once model builders learn the test.

Physical AI models need to do more than perform well in isolated demonstrations. A useful robot-control model must adapt across objects, positions, contact conditions, and multi-step tasks that vary from one environment to another.

Static benchmarks make results easier to compare, but they carry a weakness. Once the same task set is used repeatedly, teams can tune against the benchmark itself. This risk runs deeper in robotics, where simulated success does not always translate cleanly into real-world deployment.

Open Axis Benchmark balances those tradeoffs by freezing and versioning each evaluation round while keeping the task source dynamic. Participants can compare results within a round, and the benchmark still changes over time.

That structure fits OpenRoboto's role as a Bittensor robotics subnet. Subnet competition depends on evaluation systems that reward useful model improvements rather than familiarity with a narrow test set. A living benchmark gives SN80 a clearer way to test whether submitted robot policies can handle new manipulation challenges as the field advances.

Robot models are improving faster while many evaluation sets stay frozen, so a score can end up reflecting how well a model has adapted to a known benchmark instead ofhow well it generalizes to unfamiliar manipulation tasks.

Open Axis uses a rotating evaluation pool instead. Each round locks a fresh task set from the Axis Library so participants know the version being evaluated, while the broader library keeps expanding behind it. A single round stays reproducible without becoming a permanently fixed target.

The Axis library includes more than 6,000 tasks and 5.5 million trajectories across grasping, placement, articulated objects, multi-step manipulation, and related robot-control problems.

Axis Robotics built the benchmark with OpenRoboto on top of its data engine.

Tasks leave the active set when they stop separating models, keeping evaluations on tasks that produce real differences between systems, not stale ones that no longer distinguish anything.

OpenRoboto Launches Shift to Collect First-Person Robotics Data on Bittensor
The SN80 project is adding a data collection track that rewards accepted recording hours and turns workplace video into training data for robotics models.

Simulation First, With Real-Robot Validation Ahead

Season 1 submissions are open and describe the path as moving "from a sim leaderboard" toward validation on a physical xArm 6 robot.

Simulation lets more participants test and submit models quickly. Still, physical robot validation more strongly checks whether policies survive real-world constraints such as hardware tolerances, perception noise, friction, object variation, and timing differences.

This enables Benchmark v1 to add expanded tasks, reproducible evaluation, and stronger robot-transfer gates, with the final gate running policies on a physical xArm 6.

Comments

Latest