StreamingTalker: Breaking the Latency Barrier in 3D Facial Diffusion Models

StreamingTalker: Audio-driven 3D Facial Animation with Autoregressive Diffusion Model

2025-01-01
Yifan Yang, Zhi Cen, Sida Peng, Xiangwei Chen, Yifu Deng, Xinyu Zhu, Fan Jia, Xiaowei Zhou, Hujun Bao
Summary
Problem
Method
Results
Takeaways
Abstract

StreamingTalker is a novel speech-driven 3D facial animation framework that utilizes an Autoregressive (AR) Diffusion Model to generate realistic, synchronized facial motions in a streaming manner. It achieves SOTA performance on BIWI and VOCASET datasets, significantly reducing inference latency for long-form audio.

TL;DR

StreamingTalker introduces an Autoregressive (AR) Diffusion Model for speech-driven 3D facial animation. By shifting from global sequence denoising to a frame-by-frame streaming approach, the model achieves SOTA lip-sync accuracy, maintains constant 25ms latency, and handles audio of arbitrary length without performance degradation.

The "Long Audio" Bottleneck

In the world of 3D digital humans, diffusion models have become the gold standard for generating "natural" motion. However, they suffer from a "fixed-horizon" problem. Most models are trained on short clips (e.g., 4 seconds). When presented with a 60-second audio track, they either fail to generalize or require processing the entire 60 seconds before showing a single frame—a dealbreaker for real-time virtual assistants.

The technical challenge lies in Temporal Context. How do you maintain the expressive diversity of diffusion while ensuring the causality required for streaming?

Methodology: Autoregressive Condition Predictor

The core innovation is the Condition Predictor that bridges historical motion with future generation.

1. Latent Space via VQ-VAE

Instead of operating on raw 3D mesh vertices (which are high-dimensional and noisy), StreamingTalker uses a VQ-VAE to compress facial motions into a compact, discrete latent space. This stabilizes the learning process.

2. The AR Diffusion Loop

The architecture (shown below) consists of two main stages:

  • AR Condition Predictor: A Transformer decoder that takes a fixed window of past motion latents (e.g., the last 60–120 frames), current HuBERT audio embeddings, and a speaker identity. It outputs a "dynamic condition."
  • MLP Diffusion Head: A lightweight network that uses the dynamic condition to steer a standard denoising process, predicting the next latent frame from Gaussian noise.

Model Architecture

By utilizing ALiBi-based causal attention, the model is naturally biased toward recent history, allowing it to extrapolate to much longer sequences than seen during training.

Experiments and Benchmarks

StreamingTalker was evaluated on two industry-standard datasets: BIWI (rich expressions) and VOCASET (standard speech).

Quantitative Edge

The model secured a first-place finish in Lip Vertex Error (LVE) and Face Dynamics Distance (FDD). On long sequences (2000+ frames), it showed a massive improvement over FaceFormer and DiffSpeaker, which tend to "drift" or become unstable as time progresses.

MethodVOCASET LVE ↓BIWI FDD ↓
DiffSpeaker (Prior SOTA)3.14783.8535
Ours (StreamingTalker)2.72063.6690

The Speed Revolution

The most striking result is the latency graph. Traditional diffusion models (FaceDiffuser, DiffSpeaker) have latents that scale linearly with audio duration. StreamingTalker stays flat at 25ms, providing a true "streaming" experience.

Inference Latency Comparison

Visual Evidence: Why it looks better

Qualitative results show that the model captures difficult phonemes (bilabials like /m/, /p/, /b/) much better than previous methods. The mouth fully closes during these sounds, and vowel shapes (like /o/ or /u/) are significantly rounder and more natural.

Qualitative Comparison

Deep Insight & Conclusion

The genius of StreamingTalker is the fixed-window history strategy. The ablation study revealed that using "all history" actually hurts performance because it introduces a distribution shift at test time. By enforcing a consistent window of past data, the model perceives every segment of a 10-minute speech as if it were a fresh, local task.

Limitations: The model is still tied to specific training identities. Generalizing to a completely unseen face template without fine-tuning remains the "final boss" of this field.

Final Takeaway: For anyone building real-time interactive avatars or VR social platforms, this paper provides a robust blueprint for low-latency, high-fidelity 3D motion synthesis.

Find Similar Papers

Try Our Examples

  • Search for recent papers on audio-driven 3D facial animation that utilize autoregressive transformers or state-space models for real-time inference.
  • What are the original papers that proposed ALiBi (Attention with Linear Biases) and how has it been applied to handle long-sequence extrapolation in generative tasks?
  • Explore research that integrates emotional latent spaces or style-controllable encoders into diffusion-based 3D digital human animation.
Contents
StreamingTalker: Breaking the Latency Barrier in 3D Facial Diffusion Models
1. TL;DR
2. The "Long Audio" Bottleneck
3. Methodology: Autoregressive Condition Predictor
3.1. 1. Latent Space via VQ-VAE
3.2. 2. The AR Diffusion Loop
4. Experiments and Benchmarks
4.1. Quantitative Edge
4.2. The Speed Revolution
5. Visual Evidence: Why it looks better
6. Deep Insight & Conclusion