US artificial intelligence startup Tavus released a new conversational model called Griffin on October 1, claiming it is the first AI system to pass a real-time video Turing test. In a test organized by the company, 54 participants held one-minute video calls with an AI powered by Griffin-Lite, and 26 of them — 48% — believed they were conversing with a real person.
Tavus founder and CEO Hassaan Raza said the original motivation behind building Tavus was simple: “Talking to a machine should feel as natural as chatting with a friend or colleague.”
The result stands in stark contrast to the company’s previous system. Tavus’s last-generation setup, stitched together from three models — Phoenix 4.5, Sparrow-2, and Raven-1 — faced 41 participants in the same test, and only one believed they were talking to a real person, a pass rate of just 2.4%.
From pipeline to full-duplex architecture
The core difference between Griffin and traditional voice conversation systems lies in architecture. Conventional systems use a pipeline approach: first perform speech recognition, then pass the text to a language model to generate a response, and finally output via speech synthesis to drive a digital avatar. This means the system typically has to wait for the user to finish speaking before it can begin responding.
Griffin instead uses what Tavus calls a “full-duplex video-to-video” architecture, splitting the system into two parts running in parallel: a continuous conversation modeling engine that persistently senses the user’s audio and video signals and decides in real time when to respond, what to say, and how to express it; and an audiovisual generation engine that synchronously converts those control signals into sound and video.
The model re-evaluates the conversation state at sub-second intervals. While the user is still speaking, Griffin can nod, interject, adjust its expressions, and even begin responding; if the user suddenly pauses, it doesn’t simply interpret the pause as “the other person has finished speaking.”
Tavus defines Griffin as a new category — a “Human Interaction Model” (HIM) — emphasizing that the goal is not to let people “operate computers” but to let people “work with computers.”
On technical metrics, Griffin-Lite’s streaming video generator can produce 720p video in 320-millisecond chunks and accepts real-time control signals, allowing the conversation model to directly control gestures, gaze, and emotional shifts. Speech generation is likewise designed for real-time interaction, capable of cloning a speaker’s voice from roughly 10 seconds of audio and generating speech incrementally rather than waiting for a full sentence to finish before outputting.
NVIDIA benchmark performance
Tavus submitted Griffin-Lite to NVIDIA’s VideoFDB full-duplex audiovisual conversation benchmark for evaluation. On the generation side, Griffin-Lite scored 3.83, against a human reference value of 3.92 and a second-best system score of 2.80. On the perception side, Griffin-Lite scored 3.73, against a human reference value of 4.20 and a second-best system score of 3.44. Tavus said that among the 15 models evaluated in the benchmark, Griffin-Lite posted the highest perception score.
In video generation, Griffin-Lite ranked first on three metrics: DOVER, FID, and THEval. Its audio-to-video latency averaged 0.43 seconds on NVIDIA H100 chips, which Tavus says is half that of the next-fastest method.
Notably, test subjects didn’t merely think Griffin “looked human.” More than half said they never once considered during the call that the other party might not be a real person; those who did grow suspicious typically began doubting within the first 20 seconds.
Safety considerations and limited release
Despite the striking test results, Griffin is not currently available to general customers. Tavus explicitly stated that the ability of a human interaction model to communicate naturally with people also means it could cause humans to mistakenly believe they are not facing an AI. The company believes further safety and alignment measures are needed before an official release, and it is developing safety disclosure features.
For now, Griffin-Lite is available only as a research preview to a small group of screened testers; interested parties can submit an application form through the Tavus website.
This cautious approach has real-world grounding. In January 2025, hackers linked to North Korea used deepfake technology to impersonate trusted contacts in Zoom or Teams video calls, with security researchers attributing the intrusions to BlueNoroff, a subgroup of the Lazarus Group. Victims were induced during the calls to install malware disguised as an audio repair tool.
Cryptocurrency exchange Kraken also flagged a suspected North Korean job applicant in 2025, whose security team asked spontaneous questions during the interview — such as requesting a government ID or asking about local restaurant names — and the candidate struggled.
Industry context and funding
AI’s ability to mimic humans in conversation is not a new topic. A study from the University of California, San Diego found that OpenAI’s GPT-4.5, when prompted to play an introverted, web-savvy young person, convinced judges it was human in 73% of text conversations. Griffin’s breakthrough lies in extending that capability into the real-time audiovisual domain.
Tavus completed a $40 million Series B round led by CRV in November 2025. The company’s earlier system, which scored 2.4%, was assembled from three separate models handling vision, conversation, and perception tasks respectively.
Regarding the test methodology itself, some observers have pointed out differences from the classic Turing test proposed by British mathematician Alan Turing in 1950. In the classic test, the judge knows they are facing a hidden human and a hidden machine and must determine which is the machine. In Tavus’s test, participants were told they would be chatting with another participant and did not know the other party could be an AI. A community note on social platform X also pointed out that the results had not been independently verified and did not follow standard protocols. Tavus said participants were recruited through what it called an “independent research platform” and that NVIDIA independently executed the benchmark evaluation.
Tavus’s vision for Griffin is ambitious: in the future, a student could face an AI directly and explain what they didn’t understand, with the AI adjusting its explanation based on facial expressions and reactions; a worker could rehearse a difficult conversation with an AI; a customer could hold a broken part up to the camera and let the AI understand the problem through what it sees and the ongoing exchange. No need to remember the right commands or find the right menu — just talk it out, as you would with another person.

