Lightweight VAE
Efficient encoding and LQ-conditioned decoding reduce autoencoding overhead, with cached features for incremental video processing.
FastVRFastVR is a one-step diffusion framework for video restoration, supporting arbitrary-scale super-resolution, and streaming inference for long videos.
1 Alibaba Group2 Xidian University
Reported model runtime on a single H20 GPU; excludes file I/O and model loading.
01 / LONG-VIDEO RESULTS
AIGC and Real-World videos.
Drag to compare input and FastVR enhancement.
Select a video, then press play to load its input and FastVR result.
If playback is unavailable, access the video files on GitHub.
Diffusion-based video restoration recovers realistic details, but its large-scale use is limited by the cost of VAE encoding and decoding and the quadratic complexity of full attention.
FastVR combines a lightweight VAE with chunk-wise causal attention to enable efficient, one-step streaming restoration. Continuous trajectory learning and velocity consistency regularization improve restoration quality during training.
Experiments on synthetic and real-world benchmarks show strong perceptual quality and temporal consistency, with 11 FPS at 1080p on a single NVIDIA H20 GPU.
02 / UNDER THE HOOD
Efficient inference. Restoration-oriented training.
Efficient encoding and LQ-conditioned decoding reduce autoencoding overhead, with cached features for incremental video processing.
Bidirectional interaction within each chunk; bounded preceding context across chunks. KV Cache reuses history during streaming.
Latent trajectory learning and velocity consistency are followed by pixel-space L1 and DISTS supervision.
03 / LESS WAITING. MORE VIDEO.
1080p video · single NVIDIA H20 GPU
Single H20 GPU · 1080p. Runtime excludes file I/O and model loading.
04 / MEASURED, NOT JUST SEEN
FastVR achieves the highest MUSIQ and CLIP-IQA scores across all five reported datasets using one-step diffusion. Full-reference and no-reference metrics capture different aspects of restoration quality.
↑ Higher is better. ↓ Lower is better. Ties share the same rank at the reported precision. Values are reproduced from the paper.
A CLOSER LOOK
Original PNG comparisons. Click to inspect fine details.
@misc{chen2026fastvrefficientstreamingvideo,
title = {FastVR: Efficient Streaming Video Restoration with One-Step Diffusion},
author = {Xiaoxu Chen and Qin Yang and Haoran Bai and Sibin Deng and Ying Chen},
year = {2026},
eprint = {2609.36757},
archivePrefix = {arXiv},
primaryClass = {cs.CV},
url = {https://arxiv.org/abs/2609.36757}
}