Table of Contents
Tavus has introduced Griffin, a full-duplex video-to-video AI model designed to hold face-to-face conversations in real time.
Griffin is Tavus’s first Human Interaction Model, or HIM, a class of system built to see, hear, speak, move, and react continuously during a conversation rather than waiting for a user to finish speaking before responding.
48% of participants in a live one-minute study believed they were talking to a real person when paired with Griffin-Lite, the research-preview version of the model. Previous Tavus systems using its Conversational Video Interface had a pass rate of 2.4% or lower under the same type of evaluation.
Introducing Griffin, the first model to pass the video Turing test.
— Tavus (@tavus) October 1, 2026
48% of people who talked to it live thought it was a real human. Previous systems have had a pass rate <3%. It is #1 on NVIDIA's benchmark for full-duplex AI video.
It’s the first Human Interaction Model (HIM). pic.twitter.com/eb11XbC8Xr
Griffin Is Built for Full-Duplex Video Conversation
Most real-time AI video systems are built as cascades. One model transcribes speech, another model decides what to say, and additional systems generate the voice, face, and video output. That architecture can work, but each handoff can add latency and discard parts of the interaction, such as tone, facial expression, gaze, or what is happening on camera.
Griffin is designed around a different structure. The model continuously processes audio and video while also generating speech and video output. That lets it backchannel while a user is talking, stop when interrupted, respond to visual context, adjust its tone, and use nonverbal behavior such as nods, facial expressions, gaze changes, and gestures.
The system has two main components. A Continuous Conversational Modeling engine reads incoming audio and video, decides when and how to respond, and emits expressive controls. An Audio-Visual Generation engine then turns those controls into streaming speech and video. Because those processes run concurrently, Tavus says Griffin can update its behavior during a sentence instead of waiting for a new turn.
Conversation is not only a sequence of words. People use pauses, glances, interruptions, expressions, tone, and timing to decide whether to continue, clarify, wait, or respond. Tavus built Griffin to treat those signals as part of the core interface rather than as optional extras layered onto a chatbot.
Griffin Led on NVIDIA’s VideoFDB Benchmark
Tavus also reported that Griffin-Lite ranked first on both tracks of NVIDIA’s Video Full-Duplex Benchmark, a benchmark for evaluating audio-visual conversation quality and timing.
On VideoFDB’s generation track, which evaluates the model’s own speech and video response, Griffin-Lite scored 3.83 out of 5. Tavus said that placed it 1.03 points ahead of the next-highest system, Gemini 2.5 + Anam at 2.80, and 0.09 below the human-reference score of 3.92.
On the perception track, which evaluates whether the model understands the conversational moment using audio and video input, Griffin-Lite scored 3.73. Tavus said that was 0.29 points ahead of the strongest reported baseline, MiniCPM-o 4.5 in its audio-only configuration, and 0.47 below the human-reference score of 4.20.
The benchmark also measures takeover-rate alignment, a timing metric for how closely a model’s decisions about when to speak match reference conversations. Tavus said Griffin-Lite scored 62.8% on generation and 73.8% on perception, the highest of any system on both tracks.
Tavus separately tested Griffin-Lite in a live face-to-face study with 54 participants recruited through an independent research platform in the U.S. and Europe. Participants were told they would be matched with another person for a one-minute video call and were only informed afterward that the partner was an AI model. Of those participants, 26 said they believed the partner was a real person.
Participants also rated Griffin-Lite favorably on naturalness, trustworthiness, whether they would enjoy speaking with it again, whether the partner seemed to be listening, and how well the conversation flowed.
How Griffin Generates Speech and Video in Real Time
Griffin’s technical design is aimed at reducing the delay between what a person says or does and what appears in the model’s response.
For speech, Tavus says Griffin-Lite uses Tavec, a convolutional audio codec that maps 48 kHz audio into a compact continuous latent representation of 40 values per frame at 100 frames per second. A causal decoder can emit audio packets as small as 10 milliseconds, allowing speech generation and playback to proceed while an utterance is still being formed.
For video, Tavus says Griffin-Lite uses a fast diffusion-based generator that produces 720p video in 320-millisecond chunks. The video model takes a reference image, streaming audio, and streaming controls from the conversational model, then generates one video latent at a time. Tavus said the approach lets Griffin control the whole frame in real time, including face, gaze, upper-body motion, arms, hands, and background movement.
In audio-to-video latency tests against four published streaming diffusion models, Tavus said Griffin-Lite averaged 0.43 seconds on H100s, about half the latency of the next-fastest method. The model also ranked first against those baselines on DOVER, FID, and THEval visual-quality measures, while placing second on LSE-C lip-sync confidence.
The product claim depends on timing as much as visual quality. A lifelike avatar that responds too late, talks over users at the wrong moment, or fails to react to visual context can still feel artificial. Tavus argues that Griffin improves the conversational layer and the generated audio-video layer at the same time.
Safety Is Part of the Release Plan
Tavus is not releasing Griffin broadly yet. The same capabilities that make Human Interaction Models useful for natural human-machine communication can also make them deceptive if users believe they are interacting with a real person.
The risk applies directly to Griffin, which is designed to look, sound, and respond like a live conversational partner. Tavus said it is working on safe disclosure features and additional alignment procedures before wider availability.
For now, the research preview gives Tavus a way to test Griffin with a controlled group while continuing to operate its existing PAL products and developer tools. The company’s current platform already lets developers build real-time face-to-face AI agents using Tavus models for rendering, perception, turn-taking, and speech.
The Griffin announcement moves Tavus deeper into the race to make AI interfaces less dependent on text boxes, voice commands, and menu-like prompting. Its central claim goes beyond generating a realistic talking person. Griffin is meant to participate in the timing and nonverbal flow of a live conversation.
If Tavus releases Griffin with clear disclosure and safety controls, users could interact with AI agents through face-to-face communication rather than through prompts alone.