Multimodal Models · Open Models · Long Context

Keye-VL 2.0 vs InternVideo3: 256K-context MoE or an 8B agentic loop

Keye-VL 2.0 buys long video with sparse-attention infrastructure: 30B-A3B, 256K context, 74.1 on LongVideoBench. InternVideo3 bets on a dense 8B plus inference-time tool use for +2.7 on Video-MME.

Keye-VL 2.0 vs InternVideo3: 256K-context MoE or an 8B agentic loop

The two bets

Both papers solve long-video understanding with open weights, and they spend their complexity in nearly opposite places. Keye-VL 2.0 is a systems bet: a 30B-A3B mixture-of-experts with only about 3B active parameters, DeepSeek Sparse Attention adapted to a GQA-based multimodal architecture, custom kernels, and a 256K training context. The model is expected to swallow an hour of video in a single forward pass, cheaply enough to serve.

InternVideo3 is a budget bet: an 8B dense model from Shanghai AI Lab and Nanjing University that scores 73.8 on Video-MME in a single pass, and then, instead of growing the context window, wraps the model in an agentic reasoning loop called Multimodal Contextual Reasoning (MCR) that re-queries segments, runs ASR on spans it misheard, and verifies claims before answering. That loop lifts Video-MME from 73.1 to 75.8.

So the honest question is not “which model is better.” It is whether you would rather pay for long-context infrastructure at training time or for tool-using inference at serving time.

Key numbers

MeasurementKeye-VL 2.0 (30B-A3B)InternVideo3 (8B)Setting and sourceSame harness?
Active parameters~3B of 30B (MoE)8B denseBoth papers, architecture sectionsYes, by design
Training context256KNot stated; M2LA serves 768K tokens on one H200Keye-VL page; InternVideo3 pageNo (training target vs serving demo)
LongVideoBench74.1Not reportedKeye-VL page, its own table (Qwen3-VL-235B-A22B Thinking: 70.5)n/a
Video-MMEv2: 35.3 @ 64 frames, 42.4 @ 512 frames73.8 single-passDifferent protocol versionsNo: v1 vs v2, do not compare
MLVULarger baselines “in the same range” (no number)77.3InternVideo3 page (InternVideo2.5-7B: 72.8; Qwen3-VL-8B: 57.6)No
EgoSchemaNot reported76.6 (InternVideo2.5-7B: 63.9; Qwen3-VL-8B: 69.8)InternVideo3 pageNo
Agentic inferencetau2-Bench 82.6, VitaBench 33.1MCR loop: Video-MME 73.1 → 75.8 (+2.7)Different benchmarks entirelyNo
Serving efficiencyDSA: >3x prefill, >5x decode cheaper at 128K vs full attentionM2LA: 768K tokens on one H200 where base OOMs at 512KEach paper’s own stackNo
Frontier referencen/aGemini 3 Pro at 88.6 on Video-MMEInternVideo3 pageNo

Read the table for what it does not contain. There is no row where both models are scored on the same benchmark under the same protocol, so no cell in this table is a head-to-head result. The closest thing to a controlled comparison is each paper’s internal baseline column: Keye-VL 2.0 tables itself against Qwen3-VL-235B-A22B Thinking (74.1 vs 70.5 on LongVideoBench), and InternVideo3 tables its 8B against the 7B InternVideo2.5 and same-size Qwen3-VL-8B (73.8 vs 65.1 and 71.4 on Video-MME).

Why 35.3 is not worse than 73.8

The single most misleading comparison a reader can make here is lining up Keye-VL 2.0’s 35.3 to 42.4 “Video-MME” against InternVideo3’s 73.8. The Keye-VL page reports Video-MME v2, a revised, harder benchmark where its own numbers climb from 35.3 with 64 frames to 42.4 with 512 frames. InternVideo3’s 73.8 is the original Video-MME. Same family name, different protocol, different difficulty. Any page that subtracts one from the other is manufacturing a gap that the papers never claimed.

What you can say: InternVideo3’s 73.8 single-pass on v1 is a strong result for an 8B model, beating Qwen3-VL-8B’s 71.4 in its table, while the proprietary frontier (Gemini 3 Pro at 88.6) remains well ahead. Keye-VL 2.0’s strongest evidence is elsewhere: LongVideoBench 74.1, where it leads Qwen3-VL-235B-A22B Thinking’s 70.5, and all three TimeLens temporal-grounding subsets (58.5 on ActivityNet-TimeLens, 70.1 on QVHighlights-TimeLens, 58.4 on Charades-TimeLens).

What each model is really buying you

Keye-VL 2.0’s purchase is the context window. Adapting DeepSeek Sparse Attention to a GQA multimodal backbone, plus Chunk ViT processing, ViT-LM heterogeneous parallelism, and custom DSA kernels, is what makes a 256K video context servable: the report claims DSA-specific optimization cuts prefill cost by over 3x and decode cost by over 5x at 128K versus full attention. Post-training is staged on top: cross-modal multi-teacher on-policy distillation, then Context-RL and Video-RL targeting long-context retrieval, temporal grounding, and agentic behavior. The implicit claim is that most of what “agentic video understanding” needs can be baked into one long-context forward pass.

InternVideo3’s purchase is the opposite: keep the model small and spend at inference. Its MCR loop gains +2.7 on Video-MME (73.1 → 75.8) not from the weights but from the harness: extra passes that re-watch a segment, transcribe audio the visual path missed, or check an answer before committing. Its other structural piece, M2LA attention, is plumbing that stretches an existing GQA backbone to 768K tokens on a single H200 where the base architecture runs out of memory at 512K. Note the parallel to search-agent harnesses: the MCR loop’s gain belongs to the scaffold and tool budget, and crediting it to the 8B network understates what you would have to deploy.

When to use which

  • Hour-level video, one pass, controlled serving cost: Keye-VL 2.0. Its whole stack, from sparse attention and 3B active parameters to decode optimizations, exists so a long video fits one 256K context window without a giant dense model. The LongVideoBench and TimeLens numbers are the evidence.
  • Tight deployment budget, tolerable latency, tooling available at inference: InternVideo3’s 8B. A dense 8B is far cheaper to host than a 30B MoE with custom kernels, and the +2.7 MCR gain shows inference-time tool use can substitute for scale on at least some benchmarks.
  • You need the benchmark, not the philosophy: neither table is directly comparable, so match the benchmark to your workload. LongVideoBench and temporal grounding favor the Keye-VL evidence; Video-MME v1, MLVU (77.3), and EgoSchema (76.6) are where InternVideo3 publishes numbers.
  • You wanted a frontier model: both papers are candid that they are not it. InternVideo3’s page notes Gemini 3 Pro at 88.6 on Video-MME and that newer releases (Qwen3.5+, GLM, Kimi) already surpass it; Keye-VL 2.0’s page concedes that closed or larger open baselines still sit in the same range on Video-MME and MLVU.

Limits and open questions

The biggest gap is the missing shared evaluation: no protocol in either paper scores both models on the same benchmark version, so this comparison is structurally a juxtaposition of two evidence packages, not a race. Second, the two “context” numbers measure different things: 256K is Keye-VL 2.0’s training context, while InternVideo3’s 768K is a serving demonstration of M2LA on one H200; do not read one as larger than the other. Third, agentic claims live on different benchmarks: Keye-VL 2.0’s tool-use evidence (tau2-Bench 82.6, VitaBench 33.1) is a model-level capability number, while InternVideo3’s +2.7 is a harness-level delta on a single benchmark, and whether that delta holds on MLVU or EgoSchema is not shown. Finally, reproducibility: Keye-VL 2.0’s cost profile depends on its sparse kernels and serving stack, and InternVideo3’s MCR gain depends on the agent loop being deployed; in both cases, shipping only the checkpoint is not shipping the number.

FAQ

Is Keye-VL 2.0 better than InternVideo3?

There is no measured head-to-head, because the two papers report almost no overlapping benchmark versions. Keye-VL 2.0’s strongest numbers are LongVideoBench 74.1 and the three TimeLens temporal-grounding leads; InternVideo3’s are Video-MME 73.8, MLVU 77.3, and EgoSchema 76.6 for an 8B model. The fair claim is narrower: Keye-VL 2.0 is the better-evidenced choice for hour-level single-pass video, InternVideo3 for small-model deployment with an agentic loop.

Why is Keye-VL 2.0’s Video-MME score lower than InternVideo3’s?

Because they are different benchmarks with the same name. Keye-VL 2.0 reports Video-MME v2, a harder revision, scoring 35.3 with 64 frames and 42.4 with 512 frames. InternVideo3’s 73.8 is the original Video-MME. Comparing 42.4 against 73.8 is a protocol mismatch, not a capability gap.

What does InternVideo3’s agentic loop actually add?

InternVideo3’s MCR (Multimodal Contextual Reasoning) loop lifts Video-MME from 73.1 to 75.8, a +2.7 gain attributed to the harness, not the weights. The loop can re-query a video segment, run ASR on a span the visual path misheard, or verify a claim before answering. Whether +2.7 holds on the other long-video benchmarks is not shown in the paper.

How do Keye-VL 2.0 and InternVideo3 handle long context differently?

Keye-VL 2.0 builds long context into training: DeepSeek Sparse Attention adapted to a GQA multimodal backbone, 256K training context, and kernel-level optimizations that the report says cut prefill cost over 3x and decode cost over 5x at 128K versus full attention. InternVideo3 keeps an 8B dense model and instead stretches serving with M2LA attention, running 768K tokens on one H200 where the base architecture runs out of memory at 512K, a serving-level fix rather than a training-level one.