StreamingTalker: Breaking the Latency Barrier in 3D Facial Diffusion Models
StreamingTalker: Audio-driven 3D Facial Animation with Autoregressive Diffusion Model
StreamingTalker is a novel speech-driven 3D facial animation framework that utilizes an Autoregressive (AR) Diffusion Model to generate realistic, synchronized facial motions in a streaming manner. It achieves SOTA performance on BIWI and VOCASET datasets, significantly reducing inference latency for long-form audio.
TL;DR
StreamingTalker introduces an Autoregressive (AR) Diffusion Model for speech-driven 3D facial animation. By shifting from global sequence denoising to a frame-by-frame streaming approach, the model achieves SOTA lip-sync accuracy, maintains constant 25ms latency, and handles audio of arbitrary length without performance degradation.
The "Long Audio" Bottleneck
In the world of 3D digital humans, diffusion models have become the gold standard for generating "natural" motion. However, they suffer from a "fixed-horizon" problem. Most models are trained on short clips (e.g., 4 seconds). When presented with a 60-second audio track, they either fail to generalize or require processing the entire 60 seconds before showing a single frame—a dealbreaker for real-time virtual assistants.
The technical challenge lies in Temporal Context. How do you maintain the expressive diversity of diffusion while ensuring the causality required for streaming?
Methodology: Autoregressive Condition Predictor
The core innovation is the Condition Predictor that bridges historical motion with future generation.
1. Latent Space via VQ-VAE
Instead of operating on raw 3D mesh vertices (which are high-dimensional and noisy), StreamingTalker uses a VQ-VAE to compress facial motions into a compact, discrete latent space. This stabilizes the learning process.
2. The AR Diffusion Loop
The architecture (shown below) consists of two main stages:
- AR Condition Predictor: A Transformer decoder that takes a fixed window of past motion latents (e.g., the last 60–120 frames), current HuBERT audio embeddings, and a speaker identity. It outputs a "dynamic condition."
- MLP Diffusion Head: A lightweight network that uses the dynamic condition to steer a standard denoising process, predicting the next latent frame from Gaussian noise.

By utilizing ALiBi-based causal attention, the model is naturally biased toward recent history, allowing it to extrapolate to much longer sequences than seen during training.
Experiments and Benchmarks
StreamingTalker was evaluated on two industry-standard datasets: BIWI (rich expressions) and VOCASET (standard speech).
Quantitative Edge
The model secured a first-place finish in Lip Vertex Error (LVE) and Face Dynamics Distance (FDD). On long sequences (2000+ frames), it showed a massive improvement over FaceFormer and DiffSpeaker, which tend to "drift" or become unstable as time progresses.
| Method | VOCASET LVE ↓ | BIWI FDD ↓ |
|---|---|---|
| DiffSpeaker (Prior SOTA) | 3.1478 | 3.8535 |
| Ours (StreamingTalker) | 2.7206 | 3.6690 |
The Speed Revolution
The most striking result is the latency graph. Traditional diffusion models (FaceDiffuser, DiffSpeaker) have latents that scale linearly with audio duration. StreamingTalker stays flat at 25ms, providing a true "streaming" experience.

Visual Evidence: Why it looks better
Qualitative results show that the model captures difficult phonemes (bilabials like /m/, /p/, /b/) much better than previous methods. The mouth fully closes during these sounds, and vowel shapes (like /o/ or /u/) are significantly rounder and more natural.

Deep Insight & Conclusion
The genius of StreamingTalker is the fixed-window history strategy. The ablation study revealed that using "all history" actually hurts performance because it introduces a distribution shift at test time. By enforcing a consistent window of past data, the model perceives every segment of a 10-minute speech as if it were a fresh, local task.
Limitations: The model is still tied to specific training identities. Generalizing to a completely unseen face template without fine-tuning remains the "final boss" of this field.
Final Takeaway: For anyone building real-time interactive avatars or VR social platforms, this paper provides a robust blueprint for low-latency, high-fidelity 3D motion synthesis.
