CRR: Breaking the "Independence" Barrier in Video Random Access

Crowdsourcing Based Cross Random Access Point Referencing for Video Coding

2020-01-01
Hualong Yu, Xiaoding Gao, Lu Yu
Summary
Problem
Method
Results
Takeaways
Abstract

This paper introduces Cross Random-access-point Referencing (CRR), a novel video coding structure that allows Random Access Point (RAP) pictures to utilize External Reference Pictures (ERPs) from non-adjacent segments. By integrating a crowdsourcing-based optimization for ERP selection, CRR achieves state-of-the-art gains of 12.00% on VVC common test sequences and up to 25.48% on long-duration drama content.

TL;DR

To support jumping to different timestamps (Random Access), modern video encoders like VVC insert "Intra" frames that don't look back at previous data. This wastes bits because the same background often reappears later. Cross Random-access-point Referencing (CRR) solves this by allowing these entry points to look "sideways" at shared External Reference Pictures (ERPs). Using Crowdsourcing Theory, the authors optimize which pictures to store as "references" for the whole video, slashing bitrates by up to 25% for long-form drama series.

Problem & Motivation: The "Amnesia" of Modern Encoders

In current standards (HEVC, VVC), a video is divided into Random Access Segments (RASs). To ensure you can start playing from the middle of a file, each RAS begins with an IDR or Intra-RAP picture.

  • The Pain Point: These segments are independent. If a character is talking in Scene A, then we cut to Scene B, then back to Scene A, the encoder "forgets" Scene A and must re-encode it from scratch at the next Random Access Point.
  • The Gap in Prior Work: Methods like DRAP (Dependent RAP) allowed some dependency but only on the immediately preceding intra frame. This is useless for episodic content or news where Scene A and Scene C are similar, but Scene B is totally different.

Methodology: Crowdsourcing the Best References

The authors treat the selection of reference frames as an optimization problem. Instead of just picking the first frame of a segment, they look for External Reference Pictures (ERPs) that provide the best "deal" for the entire video.

1. The CRR Structure

Unlike conventional structures where the Decoded Picture Buffer (DPB) is cleared at every RAP, CRR maintains an "External DPB." CRR Structure In Fig 1(c), you can see multiple ERPs being used across non-consecutive segments, breaking the linear chain of traditional coding.

2. Crowdsourcing ERP Selection

How do you pick which frames should be ERPs? If you pick too many, the "cost" (bitrate of the ERPs themselves) outweighs the "contribution" (savings in the RASs). The authors map this to Crowdsourcing Theory:

  • Tasks: Compressing the video segments.
  • Users: Candidate ERP frames.
  • Profit Function: The reduction in RD-cost minus the cost to transmit the ERP.

By proving the function is submodular (the law of diminishing returns applies), they use a Local-Search-Based (LSB) algorithm to iteratively add or remove ERPs until the total bitrate is minimized.

Experiments & Results: Massive Wins for Long-Form Content

The authors tested CRR against VVC (VTM 3.0) and various DRAP flavors.

SOTA Comparison

On standard "Common Test Condition" (CTC) sequences, CRR gained 12.00%. However, the real power shows in 2-minute drama clips (like Sherlock or The Big Bang Theory):

  • Average Gain: 25.48% BD-rate saving.
  • Efficiency: In the movie "The Man from Earth," CRR effectively identified scene clusters, allowing the encoder to reuse background data across 37 different fragments using only a handful of ERPs.

Results Table

System Integration

One of the paper's highlights is the practical design for DASH (Streaming over HTTP). They propose a separate "External Track" for ERPs. When a user seeks to a random point, the player fetches the necessary ERP just seconds before the main segment, ensuring "Random Access" functionality remains intact without needing a massive local buffer.

Critical Analysis & Conclusion

Takeaway

CRR transitions video coding from a "forgetful" stream to a "context-aware" library. By treating ERP selection as a crowdsourcing problem, it provides a mathematically sound way to balance storage overhead vs. compression gain.

Limitations

  • Encoding Complexity: CRR adds about 12% to the encoding time. While small for offline VOD, it might be too heavy for live broadcast.
  • Content Sensitivity: If a video has zero repeating scenes (e.g., constant fast motion in a forest), CRR might actually introduce a slight overhead.

Future Prospect

This work lays the groundwork for "Cloud-based Reference Libraries," where a streaming service could maintain a shared set of background reference frames for an entire series (e.g., the "Central Perk" set in Friends), used by all episodes to achieve unprecedented compression levels across the whole library.

Find Similar Papers

Try Our Examples

  • Search for recent papers on library-based video coding or long-term reference frame management in VVC and ECM (Enhanced Compression Model).
  • What are the original theoretical foundations of submodular function maximization in the context of Rate-Distortion Optimization (RDO)?
  • Explore how the Cross Random-access-point Referencing concept can be applied to multi-view video coding or 360-degree video streaming where spatial redundancy is high.
Contents
CRR: Breaking the "Independence" Barrier in Video Random Access
1. TL;DR
2. Problem & Motivation: The "Amnesia" of Modern Encoders
3. Methodology: Crowdsourcing the Best References
3.1. 1. The CRR Structure
3.2. 2. Crowdsourcing ERP Selection
4. Experiments & Results: Massive Wins for Long-Form Content
4.1. SOTA Comparison
4.2. System Integration
5. Critical Analysis & Conclusion
5.1. Takeaway
5.2. Limitations
5.3. Future Prospect