Brain Decoding · Diffusion Models · Vision Foundation Models
MindEye vs Brain Diffuser vs MinD-Vis: fMRI-to-Image Decoding Compared
Same NSD benchmark, three systems: MindEye dominates retrieval at 93.6% against 300 candidates, Brain Diffuser wins SSIM, and MinD-Vis reports a different dataset entirely.
The three systems, one paragraph each
All three papers solve the same task: a person lies in an fMRI scanner viewing natural images, and the model must say which image it was (retrieval) or draw something close to it (reconstruction). That sounds like one problem. It is two, and the field’s best systems optimize them separately.
Brain Diffuser (Ozcelik & VanRullen, 2023) is the frozen-pipeline approach: ridge regression maps fMRI voxels to Versatile Diffusion’s VAE latents for a low-level initial guess, more ridge regressions predict CLIP vision and text embeddings, and the pretrained diffusion model does the rest. Only linear probes are trained; everything generative is off the shelf.
MindEye (Scotti et al., 2023) is the trained-pipeline approach: a 940M-parameter residual MLP maps voxels into CLIP ViT-L/14 space, with two parallel submodules, a contrastive projector for retrieval and a from-scratch diffusion prior that repairs the “modality gap” between brain and image CLIP embeddings. Trained on a single A100 in under 18 hours.
MinD-Vis (Chen et al., 2023) is the training-recipe approach: a masked autoencoder pretrains on ~136,000 fMRI segments (HCP plus GOD) with a 0.75 masking ratio, then a double-conditioned latent diffusion model fine-tunes on image pairs. Its headline numbers come from the GOD dataset, not NSD.
Key numbers
MindEye’s evaluation re-ran Brain Diffuser on the same NSD split, so this block is a genuine same-harness comparison, not numbers copied from three unrelated papers. NSD, four-subject average, 982-image test set.
| Method | PixCorr ↑ | SSIM ↑ | AlexNet(2) ↑ | AlexNet(5) ↑ | Inception ↑ | CLIP ↑ | Same harness? |
|---|---|---|---|---|---|---|---|
| MindEye | .309 | .323 | 94.7% | 97.8% | 93.8% | 94.1% | Yes, same NSD eval |
| Brain Diffuser | .254 | .356 | 94.2% | 96.2% | 87.2% | 91.5% | Yes, re-run by MindEye |
| Takagi & Nishimoto | n/r | n/r | 83.0% | 83.0% | 76.0% | 77.0% | Yes, same NSD eval |
| Gu et al. | .150 | .325 | n/r | n/r | n/r | n/r | Yes, same NSD eval |
Read the rows, not the winners. MindEye leads five of six columns. Brain Diffuser leads SSIM, the structural-similarity metric that rewards pixel-grid alignment, which fits its design: the VDVAE-initial-guess path preserves low-level layout instead of synthesizing plausible texture. That is the split these two systems are built around: MindEye buys semantic correctness, Brain Diffuser buys pixel structure. Neither number is “overall quality.”
The metrics are doing different jobs. AlexNet(2)/(5) and Inception feed the reconstruction through a pretrained classifier and ask whether the category survives; they measure semantics. PixCorr correlates pixel intensities; it measures low-level fidelity. A model can score high on one and visibly fail the other, which is why papers in this field report six metrics and why any demo that shows only the prettiest images deserves suspicion.
Retrieval, where the gap gets ridiculous
Reconstruction quality is contested. Retrieval is not. On NSD with 300 candidates where chance is 0.3%:
| Method | Image retrieval ↑ | Brain retrieval ↑ |
|---|---|---|
| MindEye | 93.6% | 90.1% |
| Brain Diffuser | 21.1% | 30.3% |
| Lin et al. (subject 1) | 11.0% | 49.0% |
The MindEye paper also reports 93.2% top-1 exact retrieval on the full 982-image test set for subject 1. This asymmetry is the most under-reported fact in brain-decoding coverage: the system whose reconstructions look more impressive is the one that fails to identify the stimulus four times out of five, while the system that retrieves at 93%+ treats generation as a rendering step. If a headline says “AI reads your mind and draws what you saw,” ask which of these two tables it is quoting.
There is also a subject-1 low-level sub-comparison worth knowing: MindEye’s low-level pipeline reaches PixCorr .456 / SSIM .493 on subject 1 versus Brain Diffuser’s .358 / .437 on the same subject. The four-subject average gap understates how far ahead the trained pipeline is on the best-sampled subject.
MinD-Vis’s numbers, and why they are not in the table
MinD-Vis does not report NSD numbers in its paper, and mixing its results into the table above would be exactly the kind of cross-dataset comparison that makes brain-decoding leaderboards meaningless. Its benchmarks are GOD and BOLD5000: 100-way top-1 accuracy of 0.212 on GOD subject 3 (a 66% relative improvement over Ozcelik et al. in the same protocol), FID 1.67, and five-sampling consistency of 0.1736 ± 0.029 on the 100-way task.
What MinD-Vis contributes is not a number to rank but a recipe everyone else adopted. Its ablation is stark: without the SC-MBM masked pretraining stage, accuracy collapses from 23.9% to 2.6–3.4%. Cross-attention-only conditioning drops to 15.6%. If you reproduce any decoder in this space, this pretraining stage is the component that matters most; that is MinD-Vis’s argument, not MindEye’s or Brain Diffuser’s.
When to use which
Identification is the goal → MindEye. The 93.6% retrieval figure against 300 candidates at 0.3% chance is a measured, same-harness result, and it is about identifying viewed images from a known stimulus set, not reading thoughts. Nothing else in this comparison comes close, and nothing else in this comparison was built to.
A reproducible reconstruction baseline → Brain Diffuser. Linear probes on frozen generators are cheap, interpretable, and its SSIM lead is real. If you need to stand up an fMRI-to-image pipeline on modest hardware and understand every component, this is the starting point.
A training recipe → MinD-Vis. Its pretraining ablation, 23.9% down to 2.6–3.4% without the masked stage, is the strongest single evidence in this field that how you pretrain the fMRI encoder matters more than which diffusion model hangs off the end. Use it as the first stage of whatever decoder you build.
Communication prosthetics → none of the above; check Brain2Qwerty instead. Brain2Qwerty decodes typed sentences from non-invasive MEG at 32% average character error rate. That is a fundamentally easier signal to wear, and still nowhere near assistive quality. It is the honest reference point for how far the deployable end of this field really is.
Limits and open questions
Subject-specific models. All three systems train a separate model per participant. The 93.6% retrieval figure comes from models fitted to hours of one person’s scanner data; it says nothing about decoding a stranger’s brain from a five-minute scan. Claims about general mind reading should be discounted until cross-subject numbers exist at this quality level.
The generator is doing a lot of the work. Every method here conditions a pretrained image model, Versatile Diffusion or Stable Diffusion’s VAE and prior. Reconstructions inherit the generator’s priors, which is why outputs look like plausible photographs even when the decoded signal is weak. This is not cheating; it is the design. But it means the images are “consistent with the brain data under a natural-image prior,” not “what the person saw.”
fMRI is the deployment wall. All three systems require a scanner, hours of per-subject training data, and laboratory conditions. None of the image systems here is closer to real-world use than the MEG typing decoder above.
Reconstruction demos are cherry-picked by physics. Given a metric suite where semantic classifiers dominate, papers select examples where the category is right. The honest evaluation is the retrieval number under a hard candidate pool, which is why this guide leads with it.
The missing controlled experiment. No paper here trains all three pipelines on the same subjects with the same compute and reports both retrieval and reconstruction. MindEye re-ran Brain Diffuser but not MinD-Vis; MinD-Vis never touched NSD. Until someone runs that study, ranking “best fMRI decoder” from the current tables is ranking partial evidence.
FAQ
Why does MindEye retrieve at 93.6% while Brain Diffuser only reaches 21.1%?
Different objectives. MindEye trains a dedicated contrastive submodule (BiMixCo plus SoftCLIP) to make brain embeddings discriminative, and treats generation as a separate rendering problem. Brain Diffuser invests nothing in retrieval; its ridge-regression probes target reconstruction latents directly. The 21.1% is not a bug: retrieval simply was not what Brain Diffuser was built to optimize.
Why is MinD-Vis missing from the NSD comparison table?
MinD-Vis was evaluated on GOD and BOLD5000, earlier datasets with different stimulus distributions and protocols. Its GOD numbers, 0.212 on 100-way top-1, cannot be compared with NSD PixCorr or SSIM, and the paper itself does not report NSD results. Putting it in the table would manufacture a comparison the data cannot support.
Brain Diffuser beats MindEye on SSIM: does that mean its reconstructions are better?
Only on the pixel-structure axis. SSIM rewards grid-level alignment, which Brain Diffuser’s VDVAE initial-guess path preserves. MindEye leads PixCorr and every semantic metric (AlexNet, Inception, CLIP), meaning its reconstructions land in the right category more reliably. For most uses, “what category of thing did the person see,” the semantic row is the one that matters.
Which system should I reproduce first?
Brain Diffuser: only ridge regressions are trained, everything else is frozen and pretrained, so it runs on modest hardware and every component is inspectable. Then add MinD-Vis’s masked pretraining stage before any diffusion fine-tuning; MinD-Vis’s own ablation shows that stage is responsible for most of the accuracy.
Can any of these systems read thoughts, imagination, or dreams?
No. All three decode images a person is currently viewing in a scanner, using per-subject models trained on hours of that person’s data. None of the three papers claims decoding of imagination, dreams, or inner speech. The retrieval numbers describe stimulus identification from a known candidate set, nothing more.