FastVR

Efficient Streaming Video Restoration
with One-Step Diffusion

FastVR is a one-step diffusion framework for video restoration, supporting arbitrary-scale super-resolution, and streaming inference for long videos.

Xiaoxu Chen1,∗ Qin Yang1,2,∗ Haoran Bai1 Sibin Deng1 Ying Chen1,†

1 Alibaba Group2 Xidian University

∗ Equal contribution · † Corresponding author

11 FPS1080p on NVIDIA H20
21.96 GBPeak GPU memory
1 stepFrom LQ to restored video
StreamingBounded temporal memory

Reported model runtime on a single H20 GPU; excludes file I/O and model loading.

01 / LONG-VIDEO RESULTS

Long-video restoration.

AIGC and Real-World videos.
Drag to compare input and FastVR enhancement.

Input
FastVR
0:00 0:00 SYNCED PLAYBACK

AIGC

Real-World

Select a video, then press play to load its input and FastVR result.

If playback is unavailable, access the video files on GitHub.

THE IDEA

High-quality restoration.
One-step diffusion.

Explore the method

Diffusion-based video restoration recovers realistic details, but its large-scale use is limited by the cost of VAE encoding and decoding and the quadratic complexity of full attention.

FastVR combines a lightweight VAE with chunk-wise causal attention to enable efficient, one-step streaming restoration. Continuous trajectory learning and velocity consistency regularization improve restoration quality during training.

Experiments on synthetic and real-world benchmarks show strong perceptual quality and temporal consistency, with 11 FPS at 1080p on a single NVIDIA H20 GPU.

02 / UNDER THE HOOD

Designed for one-step diffusion.

Efficient inference. Restoration-oriented training.

LightweightVAE is used for inference. Both training stages use the frozen official Wan VAE.
01

Lightweight VAE

Efficient encoding and LQ-conditioned decoding reduce autoencoding overhead, with cached features for incremental video processing.

02

Chunk-wise causal attention

Bidirectional interaction within each chunk; bounded preceding context across chunks. KV Cache reuses history during streaming.

03

Latent-to-pixel training

Latent trajectory learning and velocity consistency are followed by pixel-space L1 and DISTS supervision.

03 / LESS WAITING. MORE VIDEO.

Quality meets efficiency.

1080p video · single NVIDIA H20 GPU

Visual comparisonOne scene, six methods. Enlarge to compare fine textures and structural details.

Inference speed

FPS · higher is better ↑

Peak GPU memory

GB · lower is better ↓

Single H20 GPU · 1080p. Runtime excludes file I/O and model loading.

04 / MEASURED, NOT JUST SEEN

Across five benchmarks.

FastVR achieves the highest MUSIQ and CLIP-IQA scores across all five reported datasets using one-step diffusion. Full-reference and no-reference metrics capture different aspects of restoration quality.

SPMCS · SyntheticBestSecond best
Quantitative comparison on SPMCS

↑ Higher is better. ↓ Lower is better. Ties share the same rank at the reported precision. Values are reproduced from the paper.

A CLOSER LOOK

Visual comparisons

Original PNG comparisons. Click to inspect fine details.

Video comparisons

Choose a video to compare methods. Use fullscreen to inspect details.

BUILD ON FASTVR

Citation

If you find FastVR useful,
please consider citing our work.

Explore the code
BibTeX
@misc{chen2026fastvrefficientstreamingvideo,
  title         = {FastVR: Efficient Streaming Video Restoration with One-Step Diffusion},
  author        = {Xiaoxu Chen and Qin Yang and Haoran Bai and Sibin Deng and Ying Chen},
  year          = {2026},
  eprint        = {2609.36757},
  archivePrefix = {arXiv},
  primaryClass  = {cs.CV},
  url           = {https://arxiv.org/abs/2609.36757}
}

Toggle 100% to inspect the original pixels. Scroll to explore; Esc to close.