Video Generation · Efficient AI · Diffusion Models

Real-Time Video Generation: Search vs Memory vs Co-Design vs Distillation

Stream-T1, Echo-Infinity, SANA-Streaming and Causal Forcing++ all claim real-time video, but solve four different problems, and their FPS numbers are not comparable.

Real-Time Video Generation: Search vs Memory vs Co-Design vs Distillation

The four meanings of “real-time”

Every paper in this comparison says the word “real-time,” and none of them means the same thing by it. SANA-Streaming means a live broadcast: video-to-video editing at 1280x704 that keeps up with an incoming frame stream on a single RTX 5090. Causal Forcing++ means interactive: frames must stream out as fast as a user can react, which makes first-frame latency, not average FPS, the binding constraint. Echo-Infinity means sustained: not fast per frame, but able to roll out for 24 hours without the scene dissolving. Stream-T1 means something else again: it does not generate video at all. It is a search wrapper that spends extra inference compute to buy back quality on a frozen streaming model.

That matters because the headline FPS numbers in this space are routinely quoted side by side as if they were a league table. They are not. The hardware differs (RTX 5090 vs H100 vs A800), the task differs (editing vs text-to-video vs interactive action-conditioned), and the speed definitions differ (end-to-end pipeline vs DiT core alone vs a baseline’s evaluation setting). The table below keeps each number attached to its own setting on purpose.

Key numbers

MethodWhat it isSpeed, as reportedHardwareHeadline quality numbersSame harness?
SANA-StreamingHybrid linear-plus-softmax DiT for streaming V2V editing24 FPS end-to-end; 58 FPS DiT core alone1x RTX 50901280x704; “SOTA temporal coherence” stated without a published metric valueEditing task, no shared benchmark
Echo-InfinityAR DiT with learnable Memory Query tokens replacing fixed KV-cache schedules18.5 FPS1x H100VBench-Long 85.61 at 30s, 82.01 at 240s; 24-hour rollout over 1.3M+ framesVBench-Long; speed rivals LongLive 20.7, MemFlow 18.7, M&G 21.7 FPS
Causal Forcing++Frame-wise AR distillation of bidirectional diffusion to 1–2 steps14.1 FPS at 2 steps; 0.27s first-frame latencyWan2.1-1.3B studentVBench 84.14 total at 2 steps vs 84.04 for the 4-step baseline; VisionReward 6.661 vs 6.326Yes, steps ablation is same model, same benchmark
Stream-T1Training-free test-time search wrapper over frozen LongLiveEvaluated in LongLive’s 16 FPS, 832x480 setting; adds inference compute per chunkn/aVideoAlign MQ 0.350 → 0.629 at 5s; −0.002 → 0.226 at 30sYes, same backbone, before/after wrapper

Two rows deserve a second read. Echo-Infinity is the slowest-looking number in its own comparison (LongLive runs at 20.7 FPS against Echo-Infinity’s 18.5) and it is still the right choice for long rollouts, because its VBench-Long score at 240 seconds (82.01) holds where fixed-memory methods drift. Speed parity is the point: it buys stability, not throughput. And Stream-T1’s “16 FPS” is not its speed claim at all; it is the evaluation setting of the frozen baseline it wraps. Its real cost is the unreported multiplier from expanding and scoring candidate chunks per step.

The shared enemy: drift over long clips

All four methods exist because autoregressive streaming has one structural weakness. A streaming model generates video chunk by chunk, each chunk conditioned on a memory of the previous ones. Whatever the memory keeps, it keeps imperfectly; whatever it drops, the model hallucinates back. Errors compound: a small artifact in chunk three biases chunk four, the subject slowly morphs, and by 30 seconds the clip has visibly degraded. Stream-T1 measures this directly: its frozen baseline’s VideoAlign motion quality at 30 seconds is −0.002, effectively zero. Echo-Infinity’s own numbers show meaning eroding even when pixels hold: semantic score falls from 81.49 at 5 seconds to 59.53 at 240 seconds.

The four papers attack different points of that loop. Which one you want depends on where your bottleneck actually is.

The four approaches

Causal Forcing++ removes steps per frame. Bidirectional diffusion generates a whole clip at once and cannot emit frame one until the last frame is decided, which kills interactivity. The fix is autoregressive generation, but naive AR diffusion still burns many denoising steps per frame. Causal Forcing++ distills a bidirectional teacher into a frame-wise generator that needs only 2 steps per frame, hitting 14.1 FPS against 8.69 FPS for its own 4-step frame-wise variant and 10.4 FPS for the chunk-wise Causal Forcing baseline, with first-frame latency cut from 0.60s to 0.27s. The engineering win is training: replacing trajectory-precomputing ODE distillation with online causal consistency distillation cuts the few-step stage from 11,600 to 2,900 A800 GPU hours, about 4x, with no trajectory storage. The honest read from our paper page: VBench total moves only 84.04 to 84.14, within noise. This is an efficiency paper, not a quality leap, and its action-conditioned world-model variant stays at 4 steps chunk-wise, so “fully real-time” applies to prompt-driven generation.

Echo-Infinity replaces scheduled memory with learned memory. Prior infinite-video methods curate history with a fixed rule (a KV-cache window, a fixed compression ratio, a RoPE patch) that throws information away the same way regardless of content. Echo-Infinity maintains a small set of learnable Memory Query tokens; when a frame leaves the local attention window, its information is folded into the queries through attention and gating, and the queries are trained end-to-end with the diffusion transformer. Per-frame compute stays constant at any length. The result that justifies it is longevity: 24-hour rollouts exceeding 1.3 million frames in real time at 18.5 FPS on one H100, and a VBench-Long of 82.01 at 240 seconds. As our paper page notes, that long-horizon semantic score of 59.53 says meaning still erodes — the memory holds the scene together visually far longer, not semantically forever.

Stream-T1 spends compute instead of training. Given a frozen streaming model, how much quality can test-time search buy back? Stream-T1 runs three coordinated tricks per chunk: noise propagation by spherical interpolation from the previous chunk’s latents, reward pruning that expands several candidate continuations and scores them with frame-level and video-level reward models over a sliding window, and memory sinking that routes evicted KV-cache entries to discard, an EMA-smoothed sink, or a direct append at detected semantic boundaries. The gains are large exactly where drift is worst: VideoAlign motion quality goes from −0.002 to 0.226 at 30 seconds over the same LongLive backbone. The catch, from our paper page: search multiplies inference cost, which sits in unresolved tension with the “real-time” framing, and optimizing reward models can drift toward what the scorer likes rather than what looks right.

SANA-Streaming co-designs model, training, and kernels for editing. Its target is not open-ended generation but live video-to-video editing, which adds a constraint the others do not have: the output must keep up with a real input stream. Three pieces close the gap: a hybrid DiT that adds softmax attention to a subset of blocks in an otherwise linear-attention backbone, Cycle-Reverse Regularization that enforces temporal consistency by predicting source frames back from generated ones (sidestepping the nonexistence of paired long edited videos), and Blackwell-specific fused kernels plus mixed-precision quantization. The numbers: 24 end-to-end FPS at 1280x704 on a single RTX 5090, with the DiT core alone at 58 FPS. Note the gap between 24 and 58: the surrounding pipeline, not the model, is the bottleneck. Also note that the temporal-consistency SOTA claim comes without a published metric value, as flagged on our paper page.

When to use which

  • Interactive or game-like generation, tight budget: Causal Forcing++. It is the only one here that reports first-frame latency (0.27s) and per-step training cost, which is what you optimize when frames must stream out on reaction timescales.
  • Minutes-to-hours of coherent video: Echo-Infinity. Its FPS is mid-pack by design; it is buying 240-second VBench-Long stability and 24-hour rollouts.
  • Live editing of a real video feed on consumer hardware: SANA-Streaming. Nothing else here runs V2V at 1280x704 on one GPU.
  • A frozen streaming model you cannot retrain: Stream-T1. It is a wrapper; the baseline’s evaluation setting (16 FPS, 832x480) is where its numbers live, and its cost shows up in inference, not training.

Limits and open questions

No two of these papers share a benchmark, a hardware platform, and a task, so any single ranking across all four is fiction; the table above is deliberately annotated per row. The streaming quality metrics themselves are weak: Stream-T1’s own results show VBench consistency scores sitting at 97–99 on the baseline, leaving almost no headroom to measure real progress, while the reward-model metrics that do move are only as trustworthy as the reward models. Long-horizon semantics remain unsolved everywhere: 59.53 at 240 seconds is the best semantic number on this page, and it is far below short-clip scores. And “real-time” claims are fragile to definition: end-to-end versus core-only FPS, prompt-driven versus action-conditioned, and one GPU generation all change the answer.

FAQ

Which streaming video generation method is fastest?

It depends on what you count. SANA-Streaming reports 24 end-to-end FPS at 1280x704 on one RTX 5090 for editing; its DiT core alone runs at 58 FPS. Echo-Infinity reports 18.5 FPS on an H100, Causal Forcing++ 14.1 FPS on Wan2.1-1.3B. Different hardware, tasks, and definitions; not a league table.

How do streaming video models avoid drift over long clips?

Three of the four papers here fight drift directly: Echo-Infinity learns what to compress into memory tokens instead of evicting on a fixed schedule; Stream-T1 prunes candidate chunks with video reward models; SANA-Streaming trains consistency by reconstructing source frames. Causal Forcing++ fights latency instead, and inherits drift as an open problem.

Can video diffusion models run in real time on a single consumer GPU?

SANA-Streaming demonstrates 1280x704 video editing at 24 end-to-end FPS on one RTX 5090, though its kernels and quantization are tuned for Blackwell and portability is unproven. Causal Forcing++ runs on a 1.3B student, which is consumer-adjacent, but its reported numbers use datacenter A800 training hardware.

Does test-time scaling work for video generation without retraining?

Stream-T1 shows it does: wrapping frozen LongLive lifts VideoAlign motion quality from 0.350 to 0.629 at 5 seconds and from roughly zero to 0.226 at 30 seconds. The price is extra inference compute per chunk, which trades against the real-time goal.

What is the best method for very long video generation?

For measured long-horizon stability, Echo-Infinity: VBench-Long 82.01 at 240 seconds and demonstrated 24-hour rollouts over 1.3M frames. Its own long-horizon semantic score (59.53 at 240s) shows meaning still decays, so “best” here means “degrades slowest,” not “solved.”

Bottom line: “real-time video generation” is four different problems wearing one name: kill steps per frame, learn what to remember, search at test time, or co-design the whole pipeline. Pick the paper whose bottleneck matches yours. Read the Stream-T1, Echo-Infinity, SANA-Streaming and Causal Forcing++ paper pages for the full breakdowns.