World Models · Video Generation · Robotics · Vision-Language-Action

Interactive World Models: Kairos vs WALL-WM vs Gamma-World vs WorldCraft

Kairos runs a 4B world model on one RTX5090 (PAI-Bench 80.84); WALL-WM grounds robot actions in events (75.86 vs pi0.5's 55.64); Gamma-World runs four players at 24 FPS.

Interactive World Models: Kairos vs WALL-WM vs Gamma-World vs WorldCraft

Quick answer

“Interactive world model” now covers at least four different interfaces, and the 2026 papers behind each one barely share a benchmark. The question that separates them is not how good is the video but what can you actuate:

  • Kairos: a 4B native world model stack you can actually run: it leads PAI-Bench at 80.84 and generates a 480P rollout in 11.4s on a single RTX5090.
  • WALL-WM: a world action model for real robots: event-grounded training reaches 75.86 Task Progress on diverse manipulation versus 55.64 for pi0.5, with the widest margin on generalization (53.75 vs 18.50 for a from-scratch chunk baseline).
  • Gamma-World: NVIDIA’s multi-agent video world model: it simulates several players in one shared scene, streams a distilled student at 24 FPS, and runs four players after training on only two.
  • WorldCraft: object manipulation for camera-controlled world models: click an object, drag its path in world coordinates through a ~50M-parameter LoRA, and the camera controller survives (RPE rotation 0.131, where full fine-tuning blows it up to 0.237).

None of these four numbers is comparable to any other: PAI-Bench scores, real-robot Task Progress, FVD, and pixel trajectory error measure different things. What is comparable is the design move each paper makes and the failure mode it buys down. That is the structure of this page.

Key numbers

What you controlScale / costHeadline numberBenchmark type
KairosScene dynamics via action-conditioned rollout4B params; O(n) temporal attention; 11.4s per 480P rollout on one RTX5090, 3.0s on 4×A800PAI-Bench 80.84 (general), 80.03 (robot); WorldModelBench 8.94World-model leaderboards
WALL-WMReal robot end-effector, chunked at semantic eventsShared event-pretrained denoiser; unified mode adds a Qwen3.5-9B VLM75.86 Task Progress diverse manipulation vs pi0.5 55.64; generalization 53.75 vs 18.50Internal real-robot suites
Gamma-WorldMultiple keyboard-controlled players in one sceneDistilled block-causal student, KV cache, 24 FPS; trained on 32 GB200sFVD 280.0 vs Solaris 443.1 (consistency split); wins every FVD/FID column in Table 1Video generation metrics (FVD/FID)
WorldCraftCamera path plus click-and-drag object trajectories~50M-parameter SP-LoRA on a frozen camera backbone38.90px trajectory error vs DragAnything 39.86; camera RPE rotation 0.131Trajectory + camera-fidelity sets

Read the rightmost column first. Kairos and Gamma-World are graded on distributional similarity to reference video. WALL-WM is graded on whether a physical robot finished the task. WorldCraft is graded on whether the dragged object went where you dragged it. A single “best world model” ranking across these four would be meaningless; a “which interface do you need” ranking is the honest one.

Kairos: the deployability play

Kairos is built around one engineering bet: Hybrid Linear Temporal Attention, which splits temporal memory into three branches: sliding-window attention for local dynamics, dilated windows for mid-range structure, and gated linear attention with a delta-rule update as persistent global memory. That drops rollout cost from O(n²) to O(n) in horizon length, which is the entire reason a 4B world model can lead PAI-Bench (80.84, robot variant 80.03) against Cosmos 2.5 at both 2B and 14B while running on consumer hardware. A 480P rollout takes 11.4s on one RTX5090, 11.7s on one A800 (23.5GB), or 3.0s across four A800s.

The paper attaches a formal result (Theorem 2): if the global-memory branch’s gated delta update is contractive with factor ρ < 1, long-horizon excess risk stays bounded instead of compounding without limit. Read the condition, not just the theorem: the contraction factor is a learned, empirical property, and the paper does not measure it on trained checkpoints. The bound guarantees drift cannot explode; it does not guarantee drift is small. The same caveat applies to the benchmark wins: a cross-embodiment curriculum of roughly 100,000 hours of human behavioral data is a data moat, not an attention-math result.

Where Kairos sits relative to the others: it is the only one of the four aiming to be general substrate: a world model a policy rolls out against, closer in spirit to a world-action model than to a pure video generator. If your question is “can I afford to roll out a world model at all,” Kairos is the paper that says yes, with latency numbers instead of promises.

WALL-WM: supervision sliced at events, not clocks

WALL-WM attacks a different problem: most vision-language-action models predict fixed-length action chunks, which forces language, vision, and control into one window and turns training into short-horizon correlation fitting. WALL-WM cuts video, action, and caption at the same semantic event boundary, then trains a layer-coupled video-action denoiser that jointly predicts future video latents and end-effector tokens, with a sight-cone mask giving multi-view consistency.

The numbers that matter are on a real robot. Event mode averages 75.86 Task Progress on diverse manipulation against 55.64 for pi0.5, 39.97 for DreamZero, and 29.71 for LingBot-VA. The cleanest signal is the ablation, because the architecture is held fixed: event-mode supervision lifts reasoning manipulation from 32.6 to 71.6 and generalization from 22.0 to 53.75 against a fixed-length base. The honest caveat: on dexterous contact-heavy insertion the event mode barely clears its own from-scratch baseline (32.00 vs 31.25), and everything runs on internal suites with author-defined rubrics, so cross-paper comparisons are indicative rather than apples-to-apples. If your bottleneck is task decomposition and generalization rather than contact precision, the event-boundary recipe is the portable idea.

Gamma-World: past the two-player wall

Gamma-World is the only one of the four that is multi-agent by construction. Prior interactive video systems, like the two-player Solaris, bolted dense all-to-all attention over agent tokens plus learned per-slot ID embeddings: quadratic in player count and unable to add a player without retraining. Gamma-World replaces both: Simplex Rotary Agent Encoding places agents at vertices of a regular simplex in rotary-angle space, so every pair of players is permutation-equivalent by geometry rather than by learned identity, and Sparse Hub Attention routes cross-agent information through a small set of learnable hub tokens, cutting the dominant cross-agent cost from quadratic to linear.

The payoff is a property none of the others has: a model trained on two players runs four players at inference with no retraining. On the consistency split it reaches FVD 280.0 / FID 46.9 against Solaris’s 443.1 / 94.8 (roughly halving FVD), and it wins every FVD and FID column across all five protocols. A distilled block-causal student with KV caching streams at 24 FPS at near-teacher quality (FVD 239.7 vs the teacher’s 227.3). The limits: evaluation is essentially one prior system plus a weak frame-concat baseline on Minecraft-style scenes, four players is the largest count tested quantitatively, and 32 GB200s of training hardware is not a small-lab reproduction. Whether the geometry trick transfers to physical embodied agents is untested.

WorldCraft: object control without retraining the camera

WorldCraft starts from a camera-controlled world model (the class of systems that let you fly a camera but not touch anything) and adds click-and-drag object manipulation through three pieces. Normalized World Trajectory lifts your 2D drag into camera-invariant world coordinates using monocular depth, then re-projects it per frame (ablation: 30.82 trajectory error on small rotations vs 35.82 for pixel-space conditioning). Spatial-Pathway LoRA injects control through a ~50M-parameter adapter on a frozen backbone, the load-bearing choice, because full fine-tuning to add the same control wrecks camera tracking, blowing RPE rotation from 0.131 to 0.237. Trajectory-Anchored State Persistence keeps an object’s world state across off-camera excursions, holding combined RPE at 0.0233 over 253 frames.

The headline trajectory number is a modest win (38.90px error vs DragAnything’s 39.86 and Wan-Move’s 44.08, with PSNR 17.23 vs 15.97–16.42), and modest is the point: it shows object control can be added almost for free to an existing camera model. The caveats are structural: control runs at 16×16-pixel latent-token granularity so small objects cannot be addressed; monocular depth degrades at large camera rotations, exactly when world-space re-projection needs it most; and TASP persists only what you explicitly moved, it does not simulate off-screen dynamics, a different slice of the long-horizon state problem than the one memory-augmented world models attack.

When to use which

  • You need a world model as rollout substrate and care about cost per step. Kairos: linear-time temporal attention is the only architecture here with a complexity claim (O(n) vs O(n²)), and it is the only one with consumer-GPU latency numbers.
  • You are training a robot policy and your failures are compositional. WALL-WM: its widest margin is generalization (53.75 vs 18.50), and the event-ablation isolates the supervision change, so you can port the idea without adopting the whole stack.
  • You are simulating environments with several cooperating or competing agents. Gamma-World: it is the only system where player count is a runtime parameter, and the train-two-run-four result is the strongest generalization claim on this page.
  • You already have a camera-controlled video model and want manipulation. WorldCraft: the SP-LoRA recipe (~50M params, frozen backbone) is designed exactly for upgrading an existing system, and its ablation table tells you what breaks if you skip the adapter.

What none of the four gives you: a shared, public benchmark where they compete directly. FVD against pi0.5’s Task Progress is not a comparison, it is a category error. For the broader argument about when world models beat language-model reasoning as a substrate, see World Models vs LLMs.

Limits and open questions

Three cross-cutting gaps show up in all four papers. First, evaluation fragmentation: each system reports on its own suite (PAI-Bench, internal robot rubrics, Minecraft-style FVD splits, 50-clip trajectory sets), so the field has no answer to “which world model is best” that survives contact with the data. Second, long-horizon honesty: Kairos’s bound is conditional on an unmeasured contraction factor, WorldCraft’s state persistence covers only user-specified objects, Gamma-World’s four players is the tested ceiling, and WALL-WM’s event recipe buys nothing on contact-rich dexterity. Third, cost accounting asymmetry: Gamma-World tells you the training fleet (32 GB200s), Kairos tells you inference latency but not training cost, WALL-WM reports neither, and only WorldCraft discloses a data recipe (27,027 clips). Anyone budgeting a world-model project from these four papers is interpolating through missing numbers.

FAQ

Which interactive world model runs on a single consumer GPU?

Kairos: a 480P rollout takes 11.4s on one RTX5090, made possible by Hybrid Linear Temporal Attention, which is O(n) in horizon length instead of O(n²). That is the only consumer-hardware latency number among the four systems.

How much better is WALL-WM than pi0.5?

On the paper’s diverse real-robot manipulation suite, WALL-WM’s event mode averages 75.86 Task Progress against 55.64 for pi0.5. The margin is widest on generalization to randomized instructions and novel scenes: 53.75 versus 18.50 for a from-scratch fixed-length baseline. Both numbers come from internal suites with author-defined rubrics, so treat them as indicative.

Can a video world model handle more than two players?

Gamma-World demonstrates four players after training on only two. Simplex Rotary Agent Encoding is parameter-free and permutation-symmetric, so player count is not baked into the weights; on the consistency split the model reaches FVD 280.0 versus Solaris’s 443.1 while streaming at 24 FPS.

Does adding object manipulation to a camera world model hurt camera quality?

Barely, if it goes through a small adapter. WorldCraft’s ~50M-parameter SP-LoRA keeps camera RPE rotation at 0.131, while full fine-tuning to add the same control degrades it to 0.237. The measurement is camera-only, so the precise claim is that the camera controller survives, not that object motion is visually cost-free.

Why can’t these four papers be ranked against each other?

They report different metrics on different tasks: PAI-Bench leaderboards (Kairos), real-robot Task Progress (WALL-WM), FVD/FID on Minecraft-style scenes (Gamma-World), and pixel trajectory error with camera RPE (WorldCraft). Each is the right metric for its own interface question, and none is convertible into the others.