Multimodal Models · Text-to-Image · Vision Foundation Models
ARM vs SenseNova-U1 vs HYDRA-X: One Model That Reads and Draws
ARM (7B, discrete tokens) reaches GenEval 0.86 after RL; SenseNova-U1 hits 0.91 GenEval and 80.55 MMMU; HYDRA-X extends one-tokenizer unification to video (GEdit 7.17).
Quick answer
“Unified multimodal model” in 2026 means one network that both reads images and generates them, instead of a vision-language model bolted to a separate diffusion stack. Three recent papers take genuinely different routes to that goal, and they do share a few benchmarks, which makes this a real comparison rather than a vibe ranking:
- ARM: one 7B autoregressive model over a shared discrete visual tokenizer, covering understanding, text-to-image, and instruction editing under a single next-token objective. RL lifts GenEval from 0.79 to 0.86 and GEdit-Bench-EN from 5.75 to 6.68.
- SenseNova-U1: one transformer in the NEO-unify architecture that understands and generates in continuous pixel space, skipping the VAE entirely. Its 30B-A3B MoE scores 80.55 on MMMU and both variants hit 0.91 on GenEval.
- HYDRA-X: one native ViT tokenizer shared by images and videos across understanding, generation, and editing. Its editing stack reports GEdit overall 7.17, above BAGEL on the reported columns.
The headline tension: on shared benchmarks the continuous-token SenseNova-U1 reads far better (MMMU 80.55 vs ARM’s 40.2) and generates slightly better (GenEval 0.91 vs 0.86), while the discrete-token ARM argues one token space plus RL gets you three tasks at 7B. HYDRA-X’s evidence is mostly internal ablations, so treat its table position as directional. None of the three numbers below is interchangeable with another; the rightmost column of each table says what is actually being measured.
Key numbers
Benchmarks at least two of the three papers report, so you can line them up (with the caveat that each team ran its own harness):
| Benchmark | ARM (7B) | SenseNova-U1 | HYDRA-X (7B) | Comparability |
|---|---|---|---|---|
| GenEval overall (after post-training) | 0.86 (ARM-RL) | 0.91 (both 8B and A3B) | not reported | Same benchmark, different harnesses; ARM paper also lists Janus-Pro-7B 0.80 and Qwen-Image 0.87 |
| MMMU | 40.2 before RL, 41.0 after | 80.55 (A3B), 74.78 (8B) | not reported | Same benchmark; ARM paper lists continuous unified models Bagel 55.3 and Show-o2 48.9 for context |
| GEdit overall | 6.68 (ARM-RL) | not reported | 7.17 | Both papers cite BAGEL 6.52; same benchmark family, different harnesses |
| DPG / DPG-Bench | 86.00 (ARM-RL, reported as DPG) | 88.14 (A3B), 87.78 (8B, DPG-Bench) | not reported | Naming differs across papers, so treat as adjacent rather than identical |
Numbers only one paper reports, listed so you can see what each team chose to optimize (none of these are cross-paper comparable):
| Paper | Headline number | What it measures |
|---|---|---|
| ARM | GEdit-Bench-EN 5.75 to 6.68 after RL; WISE 0.50 to 0.56 | GRPO-style RL with GPT-4.1/GPT-o3 judges; reasoning-graded generation |
| ARM | Tokenizer ablation: 4 objectives give ImageNet zero-shot 80.2 and PSNR 19.6, vs 0.2 and 15.2 with caption+pixel only | Whether the shared discrete tokenizer is the load-bearing piece |
| SenseNova-U1 | CVTG-2K 0.940 average, LongText-Bench 0.979 English | Text-rich generation, its strongest niche |
| SenseNova-U1 | VSI-Bench: 8B scores 62.66, A3B scores 56.90 | Spatial reasoning, where the smaller dense model beats the MoE |
| HYDRA-X | DAVIS rFVD 11.19 (full attention 16.20, causal attention 14.05) | Tokenizer design ablation on video reconstruction |
| HYDRA-X | Editing ablation: source-target interaction lifts PSNR 20.74 to 27.65, ImgEdit 2.80 to 3.20 | What the interaction path inside the tokenizer buys |
Read the tables with one rule: GenEval and MMMU rows are cross-paper; everything in the second table is a within-paper ablation or a benchmark nobody else used.
How each model stays unified
The hard part of unification is the representation. Understanding wants semantics; generation wants fidelity. Each paper picks a different point on that trade.
ARM bets on discrete tokens. One visual tokenizer is supervised with four objectives at once (captioning, pixel reconstruction, a SigLIP-style sigmoid loss, and SigLIP2 feature distillation, with FSQ quantization), and the same tokens feed understanding, text-to-image, and editing. The ablation is the argument: caption plus pixel alone gives ImageNet zero-shot 0.2 and PSNR 15.2, while all four objectives give 80.2 and 19.6. On top of that backbone, a GRPO-style RL stage with GPT-4.1 and GPT-o3 as reward models moves GenEval 0.79 to 0.86 and GEdit 5.75 to 6.68, and the paper reports that RL on one task (text-to-image or editing) nudges the other while understanding stays flat (MMMU 40.2 to 41.0, POPE around 87).
SenseNova-U1 bets on continuous pixel space. The NEO-unify architecture generates with flow matching through an MLP decoder and no VAE, on the theory that skipping latent compression keeps generation fidelity. The cost shows up exactly where you would predict: the per-patch FFN and MLP head model each 32x32 patch independently, and the authors attribute visible grid artifacts to that choice. Its understanding numbers (MMMU 80.55, MMBench-EN 91.59, OCRBench 82.10 on the 8B) are the strongest of the three, and its one documented inversion is that the 8B dense model beats the 30B-A3B MoE on VSI-Bench spatial reasoning (62.66 vs 56.90), which the authors themselves frame as evidence that understanding and generation still compete for capacity.
HYDRA-X bets on one tokenizer across modalities, not just across tasks. A single ViT tokenizer handles images and videos, with tubelet attention plus hierarchical temporal patchify for temporal compression; the ablation says full spatiotemporal attention actually hurts reconstruction (DAVIS rFVD 16.20 vs 11.19 for tubelet). For editing, a source-target interaction path inside the tokenizer lifts reconstruction PSNR from 20.74 to 27.65 and ImgEdit from 2.80 to 3.20 in the controlled setup. It is the only one of the three that reports video understanding numbers at all (MVBench 59.1, Video-MME 60.0, LongVideoBench 59.5).
When to use which
- Pick SenseNova-U1 when the product is understanding-heavy (document OCR, chart QA, multimodal chat) but you still want solid text-to-image from the same weights. Its 0.91 GenEval with 80.55 MMMU is the best combined reading in this comparison, and its text-rich generation numbers (0.979 English on LongText-Bench) target a real product niche. Watch the grid artifacts on pixel-space output and the below-0.80 attribute-binding score on GenEval.
- Pick ARM when you want one small stack (7B) that also does instruction editing, and you can live with a discrete tokenizer. It does not top any single leaderboard (GenEval 0.86 sits under Qwen-Image 0.87, GEdit 6.68 under Step1X-Edit 6.70), but no specialist on those lists also reads images. The catch: the RL gains depend on proprietary GPT judges, so the exact 0.79 to 0.86 jump is hard to reproduce, and MMMU 40.2 is weak if understanding matters to you.
- Pick HYDRA-X when video belongs in the same token space as images. It is the only architecture here designed for that, and its ablations connect design choices to numbers cleanly. The catch: its strongest evidence is internal, there is no broad external replication yet, and dedicated proprietary video models still lead several understanding metrics.
If you only need best-in-class still-image editing today, none of these is the right tool; a diffusion editor such as BAGEL or Step1X-Edit scores higher on GEdit than any unified model here. The unified bet is about system count, not peak score.
Limits and open questions
All three papers share a structural limit: their headline generation numbers are graded against benchmarks that reward different things (GenEval’s object counting vs GEdit’s instruction fidelity vs DPG’s prompt alignment), and only GenEval, MMMU, and GEdit overlap at all. The ARM and HYDRA-X GEdit numbers come from different harnesses, so the 7.17 vs 6.68 gap should be read as “HYDRA-X claims a lead on its own setup,” not as a controlled head-to-head.
Reproducibility differs sharply. ARM’s RL stage requires GPT-4.1 and GPT-o3 access to reproduce its main gain. SenseNova-U1 documents its weak spots unusually honestly (grid artifacts, attribute binding, the VSI-Bench inversion) but the MoE scaling behavior suggests capacity competition the “synergistic” framing understates. HYDRA-X’s paper is strongest on ablation logic and weakest on external validity, with release details still pending at publication time. Finally, none of the three reports a controlled latency or serving-cost comparison against simply running a VLM next to a diffusion model, which is the honest baseline a unified stack has to beat.
FAQ
Which unified multimodal model is better at text-to-image, ARM or SenseNova-U1?
On the shared GenEval benchmark, SenseNova-U1 leads slightly: both its 8B and A3B variants score 0.91 overall, against 0.86 for ARM after its RL stage (0.79 before). ARM’s own paper lists Qwen-Image at 0.87 and Janus-Pro-7B at 0.80 on the same benchmark, so SenseNova-U1’s 0.91 is competitive with dedicated generators. The harnesses differ, but this is the closest thing to a direct answer the two papers offer.
Why does ARM’s MMMU score lag SenseNova-U1 by roughly 40 points?
ARM reports MMMU 40.2 before RL and 41.0 after, while SenseNova-U1 reports 80.55 (A3B) and 74.78 (8B). ARM’s paper itself lists continuous unified models far above it on the same benchmark (Bagel 55.3, Show-o2 48.9), so the gap is best explained as a discrete-token representation cost on understanding-heavy benchmarks, not as a training accident. If understanding quality drives your product, this row matters more than any generation number.
Is HYDRA-X better than BAGEL at image editing?
On the reported columns, yes: HYDRA-X lists GEdit overall 7.17 and ImgEdit 4.34, above BAGEL, while ARM’s independent table puts BAGEL at GEdit 6.52. Treat this as a claimed lead rather than a controlled comparison, since HYDRA-X’s numbers come from its own harness and the paper’s strongest evidence is its internal ablation (source-target interaction lifting editing PSNR from 20.74 to 27.65).
Do any of these unified models beat diffusion specialists?
Not on raw generation or editing scores. ARM-RL’s GenEval 0.86 sits under Qwen-Image’s 0.87, and its GEdit 6.68 sits under Step1X-Edit’s 6.70 and roughly ties BAGEL’s 6.52. The unified argument is efficiency of system count: one model that reads and draws at near-specialist quality, not one that tops every specialist table.
How reproducible are the RL gains in ARM and the data pipelines in these models?
ARM’s RL stage uses GPT-4.1 and GPT-o3 as reward models, so the 0.79 to 0.86 GenEval jump depends on proprietary judges and is hard to reproduce exactly. HYDRA-X had not published full release details at the time of its paper. SenseNova-U1 documents its pipeline and failure modes most openly, but its numbers still come from its own evaluation harness, as do everyone’s here.