AI has learned to pretend to be itself during video calls – Griffin model passes the “Video Turing Test”

AI has learned to pretend to be itself during video calls – Griffin model passes the “Video Turing Test”

American start-up company Tavus has launched the Griffin model for instant video communication. In the company’s experiment, 26 out of 54 participants (48%) mistook Griffin-Lite’s virtual interlocutor for a person after one minute of conversation. Developers called the result the first pass of the Turing test for a video communication format.

    Image source: tavus.io

Image source: tavus.io

However, participants were informed beforehand that they would be talking to another person. The experiment thus demonstrates the avatar’s persuasiveness in brief conversations, but does not prove that the system can survive targeted testing. The previous Tavus system, under similar conditions, was mistaken for a single person in only 1 out of 41 participants (2.4%).

Model Griffin Has a completely new architecture. Previously, a sequential approach was used: speech was first recognized, text was then passed to a language model to generate a response, and the result was output via a speech synthesis system that controlled the avatar. In other words, the system must wait for the user to finish speaking before it can start responding.

Tavus’ Griffin architecture is described as “Full Duplex Video to Video System”. It consists of just two components: a continuous conversation simulation engine that constantly monitors the user’s audio and video signals and decides in real time when and how to respond; and an audiovisual generation engine that synchronously converts these signals into audio and video. Conversation state is evaluated at intervals of less than one second. As the user speaks, Griffin can nod, interject, change facial expressions, and even begin to respond—pauses in speech are not necessarily taken as a sign that the person has finished speaking.

In Tavus, Griffin neural networks are classified as “A model of human interaction” (Human Interaction Model-HIM), that is, people do not control the computer, but work with it. Griffin-Lite generates 720p video in 320-ms chunks and can clone sound from 10-second recordings. Compared with the four video generators, the model leads in terms of DOVER, FID and THEval indicators. The average delay on converting incoming audio to video on the Nvidia H100 is 0.43 seconds, not the overall system response time.

Nvidia performed separate evaluations on the VideoFDB benchmark. Griffin-Lite scored 3.83 for audiovisual action generation, compared to the human baseline score of 3.92. In the perception category, the result was 3.73 points, compared to 4.20 points for humans and 3.44 points for the nearest competitor. Language models give scores on a five-point scale based on specified criteria; they describe aspects of communication rather than general equivalences about a person.

Remarkably, test participants didn’t just think Griffin just looked like a human being—more than half said that during the conversation they didn’t even think the interlocutor might be anything other than a worthless person. Those who are skeptical often start to doubt within the first 20 seconds of a conversation.

Griffin-Lite is available to a limited number of testers. Ahead of a wider launch, Tavus intends to refine security measures, including notifying interlocutors of interactions with artificial intelligence. The company recommends using the system to train users, rehearse difficult conversations and provide remote assistance.

If you find an error, select it with your mouse and press CTRL+ENTER.

Exit mobile version