Tavus Says Griffin-Lite Fooled 48% of People in One-Minute Live Video Calls
Tavus says Griffin-Lite fooled 48% (26 of 54) in one-minute live video calls versus 2.4% for its prior stack. Research preview only; VideoFDB scores are vendor-reported.

Tavus says its new Griffin-Lite model fooled nearly half of people on a one-minute live video call. In the company’s own study, 26 of 54 participants (48%) believed their partner was a real person. Its prior stack, built around Phoenix-4.5, convinced only 1 of 41 (2.4%).
The San Francisco video-AI company announced Griffin on October 1, 2026, calling it the first “Human Interaction Model” and the first system to pass a real-time video Turing test. Those labels are Tavus’s. Griffin-Lite is a research preview for select testers only. Wider customer access is on hold while the company works on disclosure and safety features.
The claim matters because face-to-face video is where trust decisions get made in hiring, support, sales, and scams. A model that can nod, interrupt, and react to what it sees is a different risk profile from a talking head that waits for you to finish.
What Tavus Actually Announced
Griffin is Tavus’s name for a full-duplex, video-to-video system. Perception, conversational timing, speech, and video generation run together instead of through a classic cascade of speech-to-text, language model, then text-to-speech and a face renderer.
In Tavus’s framing, cascade systems throw away tone and camera cues at each handoff and add delay. Griffin-Lite, the company says, reassesses the conversation at sub-second intervals and can backchannel, yield, or interrupt while it keeps listening and watching.
Griffin-Lite is available today to a select group of early testers as a research preview. A wider release of a more powerful model is promised after safety work. Tavus is explicit that Griffin-Lite will not be available for general customer use right now.
The 48% Study, and What It Does Not Prove
Tavus recruited participants through an independent research platform. People were told they would be matched with another participant for a one-minute video call about what they were looking forward to that year. Their partner was a PAL running on Griffin-Lite, generating face, voice, and responses in real time.
Only at the end of the survey were they asked whether it had crossed their mind that their partner might not be a real person. Every participant was then told it was AI.
Key Tavus-reported numbers:
- Griffin-Lite: 26 of 54 (48%) said their partner was a real person.
- Phoenix-4.5 + Sparrow-2 + Raven-1: 1 of 41 (2.4%) under the same protocol.
- People who said “real” averaged 79% confidence; people who said “AI” averaged 81%.
- Doubters who suspected AI often did so within the first 20 seconds.
On a 7-point scale, Griffin-Lite averaged 5.4 for seeming natural, 5.6 for seeming trustworthy, and 5.8 for whether people would enjoy talking again. Conversation flow scored 4.9, the lowest of the five axes Tavus listed.
Treat this as a vendor study, not an independent replication. The call was one minute, the topic was light, and participants were not warned they might meet an AI. Longer calls or adversarial users could look different. Tavus still calls the result a milestone: to its knowledge, the first model to pass a video Turing test.
NVIDIA VideoFDB Scores Tavus Reports
Tavus also cites NVIDIA’s Video Full-Duplex Benchmark (VideoFDB), which scores generation and perception for face-to-face audio-visual conversation. Tavus says NVIDIA scored the results in September 2026 with a language-model judge. The numbers below are as published on Tavus’s Griffin page.
| Track | Griffin-Lite | Next-best cited | Human reference |
|---|---|---|---|
| Generation (0-5) | 3.83 | Gemini 2.5 + Anam, 2.80 | 3.92 |
| Perception (0-5) | 3.73 | MiniCPM-o 4.5 (audio-only), 3.44 | 4.20 |
| Takeover-rate alignment | 62.8% gen / 73.8% perc. | Not listed as higher | Reference timing |
Tavus says Griffin-Lite is the only case where the same model was evaluated on both VideoFDB tracks, and that it led published baselines on both. Generation sits 0.09 below the human reference. Perception sits 0.47 below. Those gaps are vendor-reported and should be checked against NVIDIA’s own leaderboard before anyone treats them as settled science.
Separately, Tavus reports that Griffin-Lite’s video generator averaged 0.43 seconds of true audio-to-video latency on H100 GPUs, about half the next-fastest published streaming diffusion baseline it compared. It also claims first place on DOVER, FID, and THEval among those baselines, and second on LSE-C lip-sync confidence at 7.27.
Confirmed vs Unconfirmed
| Claim | Status |
|---|---|
| Griffin announced Oct 1, 2026 by Tavus | Confirmed on Tavus’s Griffin page |
| 48% (26/54) thought Griffin-Lite was human after a 1-minute call | Vendor-reported Tavus study |
| Prior Tavus stack fooled 2.4% (1/41) | Vendor-reported same protocol |
| VideoFDB generation 3.83 vs human 3.92 | Vendor-reported; Tavus says NVIDIA scored it |
| “First” real-time video Turing test pass | Vendor claim; not an external standards body finding |
| Public customer API or general release date | Not available; research preview only |
| Pricing in USD | Not published for Griffin-Lite |
Why Tavus Is Holding the Wider Release
Tavus’s safety section is blunt. The same traits that make a Human Interaction Model useful for tutoring or support also make it able to deceive someone into thinking they are not talking to AI.
The company says further alignment and safety procedures are required for safe release. It is working on safe disclosure features and says it is talking with organizations that tackle AI safety. Griffin-Lite stays limited to trusted testers until that work lands.
That gating pattern is familiar. OpenAI held back GPT-6.1 Astra after alignment tests missed its own bar, and Google limited Gemini 4 Argon to trusted cyber defenders first. Holding a convincing video model behind a research preview is the least surprising part of this launch.
For builders, the practical question is disclosure: will the product watermark itself, announce that it is AI at call start, or rely on user interfaces to label the agent? Tavus has not published those details yet. It only says disclosure features are in progress.
What It Means for Developers in the US, Canada, Australia, and India
There is no public Griffin-Lite price list in USD, INR, CAD, or AUD. There is also no general API access to evaluate. Teams that already use Tavus conversational video interfaces should treat Griffin as a research track, not a drop-in upgrade.
If and when a wider release arrives, the product surface will stress identity, consent, and fraud controls more than raw model scores. Banks, hospitals, schools, and marketplaces that already fight deepfake account recovery and AI phishing will need stronger liveness and call-labeling policies. That is the same problem space as Google’s selfie video sign-in work and other deepfake detection tools.
Indian product and support teams that experiment with AI video agents for onboarding or tutoring should assume platforms will push for clear AI disclosure. A one-minute demo that fools half a panel is a marketing hook. It is also a preview of how quickly voice-plus-face agents can be misused for social engineering if disclosure is optional.
How Griffin Claims to Work
Tavus describes a two-part system. A continuous conversational modeling engine reads incoming audio and video, decides when and how to respond, and emits control signals for words, emotional stance, expression, and gesture. An audio-visual generation engine turns those signals into streaming speech and 720p video.
Speech generation uses a fast autoregressive diffusion transformer conditioned on about 10 seconds of reference audio, with a continuous codec Tavus calls Tavec. Video generation is distilled from a large diffusion teacher into a few-step autoregressive generator that produces 320 ms chunks at 25 fps, according to the company.
Demo clips on the Griffin page show coaching a Rubik’s Cube solve, playing Simon Says, co-writing a story with interruptions, and guiding soldering. Those clips are curated demos.
What to Watch Next
Three signals would move this from vendor claim to verifiable product news: an independent replication of the face-to-face study with longer calls, NVIDIA VideoFDB leaderboard entries that match Tavus’s posted scores without Tavus as the only messenger, and a public disclosure design plus pricing for a general release.
Until then, log Griffin as a notable research preview with strong self-reported numbers and an explicit safety hold. The 48% figure is dramatic. It is also Tavus measuring Tavus on a one-minute call.
Frequently Asked Questions
Did Tavus’s Griffin pass a video Turing test?
Tavus says yes. In its study, 48% of 54 participants believed Griffin-Lite was a real person after a one-minute live video call. No external standards body certified that result.
Can anyone use Griffin-Lite today?
No general customer access. Tavus says Griffin-Lite is a research preview for select trusted testers while it works on disclosure and safety features.
Are the NVIDIA VideoFDB scores independent?
Tavus says NVIDIA scored VideoFDB using published metrics and its own judge in September 2026. The figures 3.83 generation and 3.73 perception are reported on Tavus’s page. Cross-check NVIDIA’s leaderboard before treating them as final.
How did Griffin-Lite compare with Tavus’s older system?
Under the same one-minute protocol, Tavus reports 48% (26 of 54) for Griffin-Lite versus 2.4% (1 of 41) for Phoenix-4.5 with Sparrow-2 and Raven-1.
Is there public pricing for Griffin?
Not yet. Tavus has not published USD list prices for Griffin-Lite because it is not generally available to customers.