Skip to content

What Is XDOF? The Robot Data Startup Building Infrastructure For Physical AI

XDOF is building data infrastructure for physical AI as robotics companies race to solve one of the field's hardest bottlenecks: real-world training data.

XDOF Robotics Data Startup

Table of Contents

XDOF, a robot data startup founded by UC Berkeley researchers, is reportedly in late-stage talks to raise a Series B at a valuation of about $1.2 billion less than three months after emerging from stealth.

The round would be led by 8VC, which reported that the deal terms are not final and could still change. XDOF had already raised a $70 million Series A in June from investors including Thrive Capital, Andreessen Horowitz, Lux, Spark Capital and WndrCo.

As AI labs and robotics companies push toward general-purpose robots, many are running into a problem that language models did not face in the same way: the internet does not contain enough structured physical interaction data to teach robots how to reliably manipulate the real world.

XDOF is trying to build that missing layer. The company collects and organizes real-world teleoperation data, develops data pipelines and annotation systems, and supports evaluation workflows for robotics companies and frontier AI labs working on physical AI.

XDOF's Role In The Robot Data Stack

XDOF was co-founded in 2024 by Philipp Wu, Fred Shentu, and Nemo Jin, and their pitch is that robotics companies need large volumes of high-quality data, but collecting it is operationally difficult.

It can require robots, warehouses, calibrated hardware, trained operators, labeling systems, simulation support, and real-world evaluation. Many AI labs may prefer to buy that infrastructure from a specialist rather than build the entire data operation internally. XDOF’s goal is to support robotics data across many hardware forms rather than one narrow robot setup as the need for different kinds of robots expands.

The company emerged from stealth in June with about 60 employees and 20 customers, including several unnamed frontier AI labs. Its reported annualized revenue is now approaching $50 million, a figure that helped prompt investors to approach the startup about another financing round.

Why Robot Training Data Is Different

Large language models benefited from vast stores of existing digital data: websites, books, code repositories, forums, documents and other text-based material. Robotics does not have an equivalent open web of usable physical experience.

A robot learning to fold clothing, place objects into containers, open drawers or assemble parts needs more than video of the task. It needs data that connects observations to actions. Useful robot training data can include camera views, joint positions, timing, trajectories, force interactions, object states, corrective interventions, and failure cases.

0:00
/0:16

Folding a T-shirt — a policy trained with SARM’s reward signal.

The difference makes robotics data harder to collect and harder to standardize. A demonstration that looks clear to a person may not contain the signals a robot policy needs. A dataset can also be large without being useful if it lacks task diversity, precise control data, consistent labeling, or evaluation feedback.

Beyond collecting demonstrations, XDOF is trying to create a feedback loop where data collection, cleaning, annotation, model evaluation, and targeted improvement feed into one another.

That loop matters because physical AI systems have to perform in environments where small failures can break the task. A text model can produce a weak answer and try again. A robot that misses a grasp, misaligns an object, or fails to recover from a small disturbance can fail the entire job.

Teleoperation As A Bridge To Autonomy

XDOF's work builds on teleoperation, a process in which a human operator controls a robot while the system records the demonstration.

The approach lets humans perform tasks through robots today while producing the examples needed to train more autonomous systems later. A teleoperated demonstration can show not only the desired outcome, but the sequence of movements, corrections, and timing needed to complete the task.

Wu and Shentu previously worked on GELLO, a low-cost teleoperation system designed to make it easier to collect robot manipulation data. That research became part of the foundation for XDOF.

XDOF plans to work across three tiers of data collection. The most valuable tier is teleoperation data collected on the actual robot a customer wants to deploy. A second tier uses lower-cost teleoperated robot systems to collect more general data. A third tier involves humans wearing sensors while performing everyday tasks, creating egocentric data that can help models learn from human movement and physical context.

The structure gives XDOF a way to serve both customer-specific deployment needs and broader robotics model development. Data collected on a customer's target robot may be most directly useful, while more general demonstrations can help build reusable physical skills.

ABC-130K Shows The Scale Of The Data Push

XDOF also helped release ABC-130K, an open-source bimanual robot teleoperation dataset developed as part of the broader ABC project.

ABC-130K is the largest bimanual teleoperation dataset to date. It includes 134,806 episodes across 195 tasks and 3,553 hours of real-world manipulation data. The dataset spans task categories such as pick-and-place, folding, sorting, tool use, handover, insertion, and assembly.

The release also includes training code, models, simulation infrastructure, and evaluation resources. The broader stack includes 400 hours of simulation data across 20 scenes and more than 100 hours of real-world evaluation. The project introduced ABC-DiT and ABC-VLA, two model families used to study behavior cloning for robot manipulation.

ABC-130K does two things for XDOF. It is a research contribution that gives the robotics community access to a larger open dataset, and it backs the company's core thesis that robot learning depends on organized, high-volume, high-quality physical interaction data.

The dataset also shows why the infrastructure layer is becoming valuable. Scaling robot data takes more than recording videos. It requires hardware choices, collection protocols, data formats, evaluation rubrics, and training workflows that researchers and companies can reuse.

Why Investors Are Paying Attention

XDOF's reported valuation talks suggest investors see robot data as one of the next infrastructure markets in AI.

During the first AI infrastructure wave, the most valuable bottlenecks were compute, models, chips, cloud capacity, and data pipelines for digital tasks. Robotics adds another bottleneck: physical experience. If companies want robots that can generalize across homes, factories, warehouses, restaurants and other real environments, they need a steady supply of reliable interaction data.

That demand could make companies such as XDOF strategically important even if they do not own the final robot product. In that sense, the company resembles earlier AI data infrastructure firms, but with a more operationally complex mandate. Text and image labeling can happen through software workflows. Robot data often requires physical equipment, trained operators, robot maintenance, sensor calibration, and real-world evaluation.

Complexity can also be seen as a risk, as XDOF's labor and hardware model may be expensive to scale. Large AI labs may eventually decide that robot data is too strategic to outsource. Dataset quality can vary widely across tasks and hardware. The company's opportunity depends on whether it can turn a difficult service business into a durable infrastructure layer.

Robotics is moving from isolated demonstrations toward commercial deployment, and the market is beginning to reward companies that solve unglamorous bottlenecks. We have seen the same pattern when reporting on robotics, from Figure's Index data engine to Bittensor-linked robotics efforts such as OpenRoboto and Axis Robotics.

The Future For XDOF

XDOF's reported Series B talks signal more than a robotics funding round. They point to physical AI entering its infrastructure phase. If robots are going to become reliable general-purpose systems, the industry needs ways to produce, organize, and evaluate the data that teaches machines how to act.

XDOF's opportunity is to become that data layer. The wider lesson for AI is that the next bottleneck may be the physical work required to help robots learn the world, not another chatbot or model API.

Comments

Latest