Skip to content

Reward AI Introduces OM-1 Robot Policy for Human-Learned Manipulation

Reward AI’s Omnibody stack uses a wearable hand interface to collect multimodal human demonstration data and run one robot policy across arms, mobile robots, and humanoids.

Table of Contents

Reward AI has introduced OM-1, an in-house general-purpose robot policy designed to learn manipulation directly from human behavior and transfer that capability across different robot bodies.

The model is part of the company's broader Omnibody stack, described as "One Model, One Data Interface, Any Body." The system combines a wearable manipulation-capture device, a unified data collection pipeline, a multimodal robot policy, and a control layer intended to make the same learned behaviors work on machines ranging from industrial robot arms to humanoids.

Reward AI is working on a problem central to robotics, building systems that manipulate real-world objects with more human-like speed and adaptability. Rather than relying on teleoperation or collecting task data on each robot body, OM-1 learns from robot-free human demonstrations gathered while people wear its Omnibody Hand device.

OM-1 Learns From Human Manipulation Instead of Robot Teleoperation

Reward AI's core argument is that human manipulation contains information that is difficult to capture when people are forced to operate through a robot-specific control interface. Teleoperation can be useful, but it often slows the operator down, constrains natural motion, and ties the resulting data to the kinematics of a particular machine.

OM-1’s policy learns from human demonstrations collected through the company's One Data Interface, then generates robot actions directly from those demonstrations. No teleoperation data and no on-robot experience are used in OM-1 training.

Robot learning systems often face a bottleneck between human skill and machine execution. If a person has to teach a robot through a remote interface, the demonstration may reflect the limitations of the interface as much as the underlying task. Reward AI is trying to preserve more of the original human behavior, including fast reaches, subtle contact changes, force application, and in-hand adjustment.

The model can pick up a new task, including long-horizon tasks with challenging dynamics, from less than 30 minutes of data. Data efficiency remains a major constraint for general-purpose robotics, so that would matter if it holds across broader task sets.

The Omnibody Hand Captures Dexterous Behavior

The Omnibody stack begins with Omnibody Hand, a wearable device built to capture manipulation at the speed and fluency of ordinary human movement. The work builds on Reward AI's prior research around DexCap, a portable motion-capture system for dexterous manipulation.

Instead of replicating every human joint, Omnibody Hand uses a compact seven-degree-of-freedom design focused on the grasp functions Reward AI sees as most important. The device captures thumb-index pinching, thumb and index finger flexion, and coordinated movement of the middle, ring, and little fingers for power grasps.

That design is meant to preserve practical manipulation behaviors such as choosing contact points, shifting between precision and power grasps, and reorienting objects in the hand. These functions matter more than copying human anatomy joint by joint.

The company also emphasizes ergonomics. A wearable that slips, fits poorly, or forces the wearer to compensate can change the grasp itself, lowering the quality of the demonstration data. Omnibody Hand is designed to accommodate differences in hand size and finger proportion, while an integrated distal flexion mechanism reduces the need for per-user link adjustment.

A Single Data Interface Combines Vision, Touch, Motion, and Force

Reward AI's One Data Interface turns human movement into multimodal training data. The system combines high-frequency tactile feedback, proximity sensing, global-shutter in-hand cameras, hand-pose tracking, and force measurement.

Each signal captures a different part of manipulation. Vision provides scene context, proximity sensing helps capture the approach before contact, tactile feedback records the contact itself, and pose and force data describe how the hand moved and how much effort was required. For tasks such as conveyor-belt sorting, missing even a short interval of interaction can make the demonstration far less useful, because grasps and corrections happen quickly.

Reward AI also uses electromagnetic sensing to improve hand-pose tracking during rapid motion. Visual-inertial tracking can struggle with fast reversals because it depends on visual update rates, while its augmented approach provides a higher-fidelity positional signal and compensates for environmental electromagnetic disturbance.

The company tested this by rigidly mounting both tracking systems to a single structure, moving them between fixed mechanical stops at eight speeds, and measuring overshoot. Averaged over ten runs per speed, Reward AI said its approach reduced mean overshoot error by 60% at high speed, from 24.9 millimeters to 9.5 millimeters.

How OM-1 Turns Human Data Into Robot Actions

OM-1 consumes the multimodal streams collected by Omnibody Hand, including images, tactile signals, inter-finger proximity, and hand-pose trajectories. Reward AI says the model processes each modality at the native sampling rate of its sensor instead of forcing all inputs into a single frequency.

High-frequency tactile and motion signals can carry information that slower visual streams miss, such as the moment contact begins, whether an object is secure, or how quickly force changes during a manipulation. Preserving those signals helps the policy reason about physical interaction over time.

The model outputs actions that include motion direction, speed, force, and the timing of key events such as grasping and moving. Those inputs and outputs create a common policy interface for expressing manipulation behavior across robot bodies.

The final layer is control. OM-1's control system runs underneath the policy at high frequency, converting predicted actions into actuation while accounting for the specific dynamics of each robot. The controller is trained with reinforcement learning in simulation to handle velocity- and acceleration-dependent dynamics, external disturbances, system delays, and smooth transitions between successive policy outputs.

That control layer is designed to let the policy run across different machines without requiring each robot to become a separate learning problem. It also helps maintain continuous motion when inference latency varies, a practical issue for robots expected to move quickly while handling contact-rich tasks.

Why OM-1 Matters for General-Purpose Robotics

OM-1 is a systems-level robotics release centered on data collection, model training, and control architecture, rather than a consumer product announcement or a factory deployment update.

Its relevance comes from the problem Reward AI is trying to solve: converting natural human manipulation into reusable robot intelligence without collecting new data for every body, task, or deployment. If the company's approach scales, data gathered from people wearing Omnibody Hand today could train robots that have not yet been designed.

That idea fits a broader shift in robotics toward generalist policies, multimodal learning, and cross-embodiment control. Robots need more than better perception or larger models. They also need training data that captures physical interaction at useful speed, and control systems that can execute learned actions on real hardware.

Reward AI's OM-1 announcement connects those pieces into one pipeline. The company is betting that human-first manipulation data, a shared policy interface, and body-specific control can shorten the path from human demonstration to useful robot behavior.

Comments

Latest