Atmee.ai
ResearchOctober 2, 20266 min read

Your AI Avatar Is a Terrible Listener. GLARE Wants to Fix That.

AI can make faces talk, but it still can't make them listen. GLARE teaches avatars to nod, smile, and laugh at the right moment by reading how a speaker's voice changes.

Atmee Research
Atmee Research
Research Team
Your AI Avatar Is a Terrible Listener. GLARE Wants to Fix That.
GLARE turns speaker audio and a single listener photo into a video of a listener who reacts at the right moments.

Picture a video call with someone who stares at you, perfectly still, while you tell your best story. No nod, no smile, nothing. Then they burst out laughing three seconds before the punchline.

That's roughly where AI-generated "listeners" are today, and our new NeurIPS'26 paper called GLARE (Generating Listening Heads with Appropriate REactions)[1] sets out to change it.

Talking is solved-ish. Listening isn't.

AI has become very good at making faces talk, but a conversation has two people in it. While one talks, the other is constantly giving small signals: a nod that says "go on," a smile at the funny part, a frown that says "wait, what?" These reactions are how we tell each other we're paying attention.

What makes them hard is timing. A reaction at the wrong moment feels worse than no reaction at all. Laugh too early and you look like you weren't listening. Nod too late and you look distracted.

The researchers found three reasons AI listeners struggle:

  • 📼 No good training data. Plenty of conversation videos exist, but nobody had labeled when the listener actually reacts.
  • 🤖 Models learn to wiggle, not respond. Existing systems learn general head movement, so the listener looks alive without really engaging with what's being said.
  • 📏 Nobody measures what matters. The usual metrics ask "does this look realistic?" rather than "did it react at the right moment?"

How GLARE tackles it

Step 1: Build a proper dataset. The Atmee AI team went through about 147 hours of real two-person conversations and marked more than 64,000 listener reactions in six flavors: nodding, head shaking, smiling, laughing, frowning, and looking surprised. Each one records exactly when it starts and ends. The result is a large, labeled record of how people react while listening.

Step 2: Teach the model to hear how things are said. What someone says matters, but how they say it matters a lot too. We tend to react when a speaker's voice rises with excitement, stresses a word, or shifts in emotion. GLARE uses an audio AI model to flag these "something just happened in their voice" moments and passes them along as hints, a bit like a friend nudging you: "this is the good part, react now."

Step 3: Grade it on reactions. GLARE is rewarded for making the right reaction at the right moment, not just for producing a nice-looking face. It's like judging an actor not just on looking natural, but on hitting their cues at the right moment.

Here's how the pieces fit together:

GLARE framework overview
Speaker audio and a photo of the listener go in, and a reacting listener video comes out.

Prosody Conditioning: listening to the voice, not just the words

Prosody is the music of speech: pitch, loudness, emphasis. It's the difference between a flat "that's great" and an excited "that's great!", and it's often what prompts a listener to react. GLARE asks an audio AI model to flag the moments where the speaker's voice gets expressive, then passes that timeline to the generator as a hint about where reactions belong.

Temporal Reaction Loss: checking the reactions as it learns

As it learns, GLARE also has to say, moment by moment, which reaction it's making (nodding, smiling, laughing, frowning, head shaking, surprised, or none). Those answers are compared with the labeled dataset, and every missed or mistimed reaction counts against it. That way it can't get by on generic movement. It has to learn what each reaction is and when it should happen.

Together: prosody tells the model where to pay attention, and the reaction loss makes sure it responds correctly.

A new judge for listeners

The team also built a better way to score AI listeners. The metrics ask practical questions:

  • R-F1: Did it make the right reaction?
  • R-tIoU: Did the timing line up?
  • R-ATD Did it jump the gun? Reacting early is penalized more than reacting a beat late, since that's what feels weird to real people.
  • R-FID Does the reaction itself look convincing?

So... does it work?

Yes, pretty convincingly. The Atmee AI team compared GLARE with five earlier methods (best in bold):

RealTalk dataset

MethodR-F1 ↑R-tIoU ↑R-ATD ↓R-FID ↓
L2L0.450.5178.324.7
DIM0.580.55133.123.0
ViCo0.430.5792.526.1
ListenFormer0.330.6573.423.1
DyStream0.540.6963.215.3
GLARE0.590.7057.415.2

Seamless dataset

MethodR-F1 ↑R-tIoU ↑R-ATD ↓R-FID ↓
L2L0.410.55112.935.7
DIM0.300.50157.334.4
ViCo0.390.52140.337.9
ListenFormer0.370.51128.735.3
DyStream0.410.55110.830.8
GLARE0.460.58102.128.9
  • 🎯 GLARE reacts at the right time more often. It had more correct reactions and tighter timing across the board.
  • ✨ It also looks better. Improving the reactions didn't cost visual quality. GLARE came out ahead there too.
  • 🗿 No more statues. In side-by-side comparisons, several older methods mostly produced frozen faces. GLARE's nods and smiles matched what the real listener did.
  • 🧩 Both additions pull their weight. Rewarding correct reactions gave the biggest boost. The voice hints added extra polish on top.

Why you should care

As AI avatars show up in virtual assistants, customer service, games, and video calls, they'll spend a lot of their time listening. An avatar that nods in the right places and smiles at your joke will feel far more human than one that stares blankly.

GLARE isn't the final word. People react differently to the same moment, and word meaning and personality aren't modeled yet. But it's a clear step toward AI that seems to actually be listening to you.

References

  1. [1]GLARE: Generating Listening Heads with Appropriate REactions

Try it

Talk to an avatar that listens

Sign up in minutes. No download. Browser-based.

Explore Avatars →
#research#avatars#listening#reactions#prosody
Your AI Avatar Is a Terrible Listener. GLARE Wants to Fix That. — Atmee.ai