Compare
Comparisons
Head-to-head comparisons of methods and models, with numbers pulled from the source papers.
Multimodal Models · ARM vs SenseNova-U1 vs HYDRA-X
ARM (7B, discrete tokens) reaches GenEval 0.86 after RL; SenseNova-U1 hits 0.91 GenEval and 80.55 MMMU; HYDRA-X extends one-tokenizer unification to video (GEdit 7.17).
World Models · Interactive world models vs World models for agents
Kairos runs a 4B world model on one RTX5090 (PAI-Bench 80.84); WALL-WM grounds robot actions in events (75.86 vs pi0.5's 55.64); Gamma-World runs four players at 24 FPS.
Multimodal Models · Spatial reasoning vs MLLMs vs visual agents
Spatial reasoning for MLLMs on three axes: SpatialClaw's Python workspace (59.9% on 20 benchmarks), ReRe's synthesized second view (+8.5 VSI-Bench on 2B), and frozen-probe evidence video models carry geometry VLMs lack.
Video Generation · Causal Forcing++ vs AnyFlow
Both distill Wan2.1 video diffusion for cheap sampling. Causal Forcing++ pins 1-2 steps per frame for 14.1 FPS real-time streaming; AnyFlow keeps steps free so quality scales from 4 to 32 NFE.
Speech Recognition · Whisper vs Mega-ASR
Whisper bets 680,000 hours of messy web audio on one robust generalist; Mega-ASR synthesizes 2.4M degraded clips to cut WER to 45.69% vs 54.01% on VOiCES. The numbers are not on the same exam.
Text-to-Image · Qwen-Image 2.0 vs Qwen-Image-Flash
Qwen-Image 2.0 is Alibaba's full-capability image generator and editor: 1K-token instructions, native 2K, a 16x VAE. Qwen-Image-Flash distills it to 4 steps, and the recipe lets the student match its teacher.
Brain Decoding · MindEye vs Brain Diffuser vs MinD-Vis
Same NSD benchmark, three systems: MindEye dominates retrieval at 93.6% against 300 candidates, Brain Diffuser wins SSIM, and MinD-Vis reports a different dataset entirely.
Text-to-Image · DALL·E 2 vs Stable Diffusion vs Imagen
DALL·E 2 diffuses in CLIP-embedding space with a learned prior, Stable Diffusion in a cheap latent, Imagen behind a frozen T5, and SD3 swaps the U-Net for a rectified-flow transformer.
Long Context · Kimi K3 vs DeepSeek V4 vs Qwen3.8-Next vs Nemotron 3 Ultra
Four 2026 frontier open models target 1M context four ways: Kimi K3 runs 3:1 delta attention, DeepSeek V4 compresses sparse KV, Qwen3.8-Next runs 3:1 GDN, Nemotron 3 Ultra leans on Mamba-2.
Multimodal Models · Keye-VL 2.0 vs InternVideo3
Keye-VL 2.0 buys long video with sparse-attention infrastructure: 30B-A3B, 256K context, 74.1 on LongVideoBench. InternVideo3 bets on a dense 8B plus inference-time tool use for +2.7 on Video-MME.
Reinforcement Learning · Agentic RL vs APPO vs LongTraceRL vs SDAR vs SCOPE
GRPO gives a multi-turn agent one scalar reward per episode. APPO, LongTraceRL, SDAR and SCOPE fix credit assignment at four different pipeline layers; here are their numbers against matched GRPO-family baselines.
Retrieval-Augmented Generation · vector retrieval vs agentic retrieval
Vector retrieval compresses a corpus into one similarity score per passage. Five papers measure the cost and how grep agents, small readers, citation grading, and per-chunk configs each recover a layer.
Code Generation · AlphaCode vs Code Llama
AlphaCode converts up to a million sampled candidates into a top-54.3% Codeforces finish; Code Llama ships open 7B-70B weights that hit 67% HumanEval: one is search, the other is infrastructure.
Reinforcement Learning · Self-distillation in RLVR vs Privileged-context teachers vs Dense token-level rewards
RLVR gives one bit per trajectory. SDPG copies a hint-conditioned teacher on correct trajectories, SDAR gates the copy per token, AntiSD rewards disagreement: three ways to mine dense signal from one privileged context.
AI Agents · SkillOpt vs Ctx2Skill vs OpenSkill vs LatentSkill vs SkillAdaptor vs SkillsVote
Six 2026 methods teach LLM agents skills without retraining: SkillOpt edits one skill document for +23.5 points, Ctx2Skill mines skills by self-play, OpenSkill learns unsupervised from the open web, plus three more.
AI Agents · RHO vs Harness-1
RHO tunes an existing harness from unlabeled trajectories; Harness-1 designs the harness so RL only learns semantic decisions. Same premise, different problems, each with its own numbers.
Code Generation · SWE-Explore vs DeNovoSWE vs Claw-SWE-Bench
Three 2026 papers isolate what moves coding-agent scores: localization stalls at 0.15-0.20 line recall, training lifts a 30B model 8x, and the harness alone swings Pass@1 by 27.4 points.
Speech Synthesis · VALL-E vs NaturalSpeech 2 vs SwanVoice
VALL-E turns speech into codec-token language modeling with a 3-second prompt, NaturalSpeech 2 diffuses continuous codec latents for prosody and singing, and SwanVoice generates a whole 1-4 speaker dialogue in one pass.
Segmentation · Mask2Former vs Mask R-CNN vs SAM vs SAM 2
Mask2Former beats Mask R-CNN 43.7 to 37.2 COCO AP at ResNet-50 in one table; SAM gets 46.5 zero-shot COCO AP vs 51.0 supervised ViTDet-H; SAM 2 hits 76.8 J&F on SA-V val vs XMem 60.1, 6x faster than SAM.
Text Embeddings · Sentence-BERT vs SimCSE vs E5
Sentence-BERT cut 10k-sentence similarity from 65 hours to 5 seconds; SimCSE turned dropout into free training pairs for 76.25 unsupervised and 81.57 supervised STS; E5 beat BM25 zero-shot across 56 retrieval datasets.
Agent Memory · MemGPT vs MRAgent vs δ-mem vs EvoMem
MemGPT pages memory like an OS, MRAgent rebuilds a memory graph while reasoning, δ-mem compresses history into an 8×8 state, and EvoMem stores change patches: four answers to where agent memory should live.
Self-Supervised Learning · SimCLR vs BYOL vs MAE vs DINOv2
SimCLR hit 69.3 ImageNet linear with negatives and 4096 batches; BYOL dropped negatives for 74.3 and small-batch robustness; MAE owns fine-tuning at 87.8 but freezes poorly; DINOv2's frozen 86.3 beats OpenCLIP 86.2.
AI Theorem Proving Papers · AI theorem provers vs formal mathematics
AlphaGeometry, DeepSeek-Prover-V1.5 and MaxProof all report headline math scores, but they prove different kinds of theorems under different verifiers and search budgets. Here is what each number actually measures.
Video Generation · Stream-T1 vs Echo-Infinity vs SANA-Streaming vs Causal Forcing++
Stream-T1, Echo-Infinity, SANA-Streaming and Causal Forcing++ all claim real-time video, but solve four different problems, and their FPS numbers are not comparable.
Vision-Language-Action · RT-2 vs π0 vs Qwen-VLA vs MolmoAct2 vs World Pilot
RT-2 wrote actions as text tokens, π0 switched to ~50Hz flow-matching chunks, and 2026 models add 3D reasoning or world-model priors. Simulation is nearly saturated; real-robot margins and OOD shifts differentiate.
Small Language Models · Phi-3 vs SmolLM2 vs MobileLLM vs TinyLlama
Below 4B parameters there is no winner: phi-3-mini hits 69 MMLU on curated data, SmolLM2 beats Llama3.2-1B on general benchmarks but trails Qwen2.5-1.5B on math, and MobileLLM adds 2.7 to 4.3 percent at 125M to 350M.
Reinforcement Learning · DelTA vs AntiSD vs DVAO
RLVR grades a whole reasoning trace with one bit at the end. DelTA reweights the token update, AntiSD adds a per-token PMI reward, DVAO rebalances several rewards by group variance — three fixes at three layers.
Alignment · RLHF vs DPO
RLHF trains a reward model and runs PPO on top of it; DPO deletes both with a closed-form rewrite and a single classification-style loss. The numbers say 'as good or better, far simpler', with real caveats.
Diffusion Language Models · Diffusion language models vs Autoregressive language models
An 8B masked-diffusion LLM matches an 8B autoregressive one on MMLU, beats it on GSM8K with a sixth of the tokens, and fixes the reversal curse. Decoding speed and serving maturity still favor autoregression.
Reinforcement Learning · GRPO vs PPO
GRPO swaps PPO's learned value model for the mean reward of a sampled group. That frees one policy-sized model and made open reasoning RL affordable, at the cost of per-token credit assignment.
Fine-Tuning & Adaptation · LoRA vs Full fine-tuning
LoRA matches full fine-tuning within a point on standard tasks at 10,000x fewer trainable parameters, but it gets there through intruder dimensions that carry its forgetting and hurt under continual training.
Sequence Modeling · Mamba vs Transformer
Mamba swaps attention's quadratic compute and growing KV cache for a fixed-size selective state: 5x decode throughput and parity with 2x-larger Transformers at 3B, but exact recall stays with attention.
Language Models · Muon vs AdamW
Muon reaches AdamW's loss with about half the pretraining FLOPs and now trains DeepSeek-V4, Kimi K3 and Qwen3.8-Next; used for SFT of an AdamW-pretrained model it loses points, and in RL it collapses.
Long Context · Sparse attention vs Full attention
Sparse attention buys 7x to 28x less attention compute at 1M tokens for roughly half a point on long-context suites, but the wall-clock gain is far smaller than the FLOPs gain and most of it is invisible at 32K.
World Models · World models vs LLMs
An LLM reasons over rules and goals in language; a world model predicts what a scene does next. The 2026 evidence says they are complementary, and the hard part is knowing when a generated future is credible.