Forcing-KV Releases Hybrid KV Cache Compression for Autoregressive Video Diffusion

Forcing-KV studies the heterogeneous roles of attention heads in autoregressive video diffusion models and proposes a hybrid KV-cache compression strategy. Static heads use structured pruning, while dynamic heads use segment-wise similarity pruning.

The public preprint reports over 29 FPS on an NVIDIA H200, about 30% cache-memory reduction, and up to 2.82× speedup at 1080P. Code and demo videos are released with the paper.

Sources