> ## Content Index
> Fetch the complete content index at: https://www.tao.media/llms.txt
> Use this file to discover other available public pages before exploring further.

# ORO Launches AssistantEval Benchmark for Replayable AI Assistant Evaluation
- URL: https://www.tao.media/oro-launches-assistanteval-benchmark-for-replayable-ai-assistant-evaluation/
- Published: 2026-10-08T23:04:44.000Z
- Updated: 2026-10-08T23:22:48.000Z
- Description: The Bittensor Subnet 15 team published a public leaderboard that grades assistants from service logs in resettable synthetic environments.
- Author: Tristan Hillerich
- Tags: Oro, Subnets, News

[ORO](https://x.com/oroagents?ref=tao.media), the team behind Bittensor Subnet 15, has launched [AssistantEval](https://assistanteval.com/?ref=tao.media), a public benchmark for evaluating AI assistants in controlled, replayable environments.

The launch gives developers and users a way to compare assistants on practical tasks such as calendar changes, inbox handling, marketplace actions, and payments-style workflows without relying only on what an assistant says it did.

AssistantEval tests assistants inside synthetic services that record every action, and the grade comes from those service logs. That makes the benchmark less about whether an assistant produces a convincing response and more about whether it actually completed the task, asked for the right permission, used the right information, and avoided unsafe actions.

0:00 

/0:30 

1× 

 Assistant Eval trailer

## How AssistantEval Tests AI Assistants

AssistantEval addresses a problem that is hard to solve in live productivity apps. Real email, calendar, and payment services cannot be reset to the same state for every run. If one assistant changes a meeting, sends an email, or approves a transaction, the next assistant no longer faces the same environment.

ORO's benchmark instead places each assistant in a small synthetic world with services such as a calendar, mailbox, and marketplace. Each task starts from a known state, gives the assistant access through run-specific keys, and records what happens through service logs. Every run saves both the conversation and the underlying service actions, including inputs, outputs, and final state.

The design allows tasks to be repeated and compared more cleanly. A simulated user sends the opening request, answers only what the assistant asks, and grants approval only under the task's rules. ORO also uses twin tasks, where one factual detail changes and the correct assistant behavior should change with it.

That structure is meant to test more than task completion. An assistant may need to ask before acting, seek missing information, recover from service failures, report back accurately, or avoid taking an action it was not authorized to take. The score depends on what the log shows, not merely on whether the assistant claims success.

## What the First Leaderboard Shows

At launch, AssistantEval's public leaderboard placed Claude Opus 5.5 first with 80% of tasks completed. GPT-6 Astra followed at 75%, Claude Fable 5.1 at 74%, and Pally at 70%. Muse scored 57%, while Grok Bot scored 51%.

ORO's published columns also track more specific behaviors, including truthfulness, median task time, unsafe runs, asking first, asking for information, using real information, and reporting back. Runs can pass, fail, or unsafe-fail, with unsafe failures covering unauthorized actions or violations of task restrictions.

The early results show why ORO is emphasizing repeated trials rather than one-shot benchmark scores. AssistantEval reported that 31% of assistant-task results that passed at least once failed on another attempt. It also reported that 27% of repeated twin-pair results split, meaning an assistant consistently passed one version of a paired task and consistently failed the other.

Those gaps point to reliability problems that may not appear in a single demo. A model can look competent when it completes one request, but AssistantEval is designed to show whether that behavior survives retries, near-duplicate scenarios, and small changes in user context.

## Why It Matters for Bittensor

ORO previously launched ORO Bench for synthetic shopping environments on Bittensor Subnet 15\. AssistantEval expands that evaluation work into broader personal-assistant behavior, including permission handling, clarification, recovery, and proactivity across common digital services.

[ORO Launches ORO Bench With Daily Synthetic Shopping Environments on SN15Bittensor Subnet 15 replaces static ShoppingBench evaluation with versioned EnvPacks, seven task families, and family-specific verifiers for agentic commerce.![](https://storage.ghost.io/c/78/0b/780ba906-b1a7-4bf0-873c-bdd5c32e5331/content/images/icon/Group-1321319358-76ecdf07-3bf0-41fa-8fb1-2a289b054e54.png)Intelligence![](https://storage.ghost.io/c/78/0b/780ba906-b1a7-4bf0-873c-bdd5c32e5331/content/images/2026/09/ChatGPT-Image-Sep-23--2026--04_48_12-PM.png)](https://www.tao.media/oro-launches-oro-bench-with-daily-synthetic-shopping-environments-on-sn15/)

The benchmark could reach beyond a leaderboard. If labs or assistant developers use ORO's environments for regression testing, they can become reusable test cases for agentic products. A failed run can be turned into a repeatable task, then used to check whether a future model or workflow actually improves.

ORO is building the benchmark as evaluation infrastructure for assistants that act on behalf of users, measuring whether they can complete user-facing work accurately, safely, and consistently when every action is recorded.