Segmentation · Vision Foundation Models
Mask2Former vs Mask R-CNN vs SAM vs SAM 2: COCO and Video Numbers
Mask2Former beats Mask R-CNN 43.7 to 37.2 COCO AP at ResNet-50 in one table; SAM gets 46.5 zero-shot COCO AP vs 51.0 supervised ViTDet-H; SAM 2 hits 76.8 J&F on SA-V val vs XMem 60.1, 6x faster than SAM.
The one-line answer
These four models are not four entries on one leaderboard, they solve three different problems. Mask R-CNN and Mask2Former are supervised models: you train them on your dataset with your classes, and they chase average precision. SAM is a foundation model: you prompt it with a point or box and it produces a mask on images it has never seen, no fine-tuning required. SAM 2 extends the same promptable idea to video with a streaming memory. The comparison that matters is therefore not “which has the highest AP” but “which problem are you solving”: a fixed taxonomy you can annotate, or arbitrary objects on arbitrary images and frames.
What each method actually changed
Mask R-CNN (2017) attached a small mask head to the Faster R-CNN detection framework: for every bounding box proposal, a tiny fully convolutional network predicts a binary mask, and the RoIAlign operation fixes the misalignment that earlier RoIPooling introduced when snapping boxes to the feature grid. It is the canonical detect-then-segment design, and its descendants dominated the COCO leaderboard for years.
Mask2Former (2021) belongs to a different family, mask classification. Instead of deriving masks from detected boxes, it predicts a fixed set of binary masks and assigns each a class label with a Transformer decoder, one architecture handling panoptic, instance, and semantic segmentation at once. Its specific contribution is masked attention: each decoder layer’s cross-attention is restricted to the foreground region of the mask predicted by the previous layer, so the queries see local detail instead of diluting attention over the whole image. The result is a universal model that beats specialized architectures on all three tasks.
SAM (2023) changed the task itself. Rather than training on a fixed category set, Meta defined a promptable segmentation task: given any prompt, a point, box, or mask, return a valid mask for the object the prompt refers to. The architecture splits into a heavy image encoder that runs once per image and a lightweight mask decoder that answers prompts in about 50 ms, so interactive use is practical. SAM was trained on SA-1B, 1.1 billion masks across 11 million images collected by a model-in-the-loop data engine, roughly 400 times more masks than any existing segmentation dataset at the time. It is explicitly not trained to win COCO, it is trained to segment anything, including categories nobody labeled.
SAM 2 (2024) took the same contract to video. A prompt on one frame must propagate through all frames, so the architecture adds a memory encoder, a memory bank, and memory attention that conditions the current frame on cached features of past frames and prompts. The data engine produced SA-V, 50.9K videos with 642.6K masklets, with 53 times more annotated masks than any prior video object segmentation dataset. SAM 2 is a single model that does promptable image segmentation, promptable video segmentation, and zero-shot video object segmentation.
U-Net (2015) and DeepLab sit upstream of this whole story: U-Net showed a contracting-expanding convolutional net with skip connections trains end-to-end from very few images and won the ISBI cell tracking challenge with 92.03 IoU against 83 for the second best method, while DeepLab introduced atrous convolution and atrous spatial pyramid pooling for dense prediction, reaching 79.7 mIoU on the PASCAL VOC 2012 test set. U-Net is the ancestor of every encoder-decoder segmenter, and the per-pixel classification philosophy it established is exactly what Mask2Former’s mask classification paradigm replaced.
Key numbers
| Measurement | Mask R-CNN | Mask2Former | SAM | SAM 2 | Setting and source | Same harness? |
|---|---|---|---|---|---|---|
| COCO instance val2017 AP, ResNet-50 | 37.2 | 43.7 | n/a | n/a | Mask2Former paper Table 2, 36 vs 50 epochs | Yes, one table, both trained on COCO train2017 only |
| COCO instance val2017 AP, ResNet-50, long recipe | 42.5 | 43.7 | n/a | n/a | Mask2Former Table 2, Mask R-CNN with 400-epoch multi-scale recipe | Yes, same table |
| COCO instance val2017 AP, ResNet-101 | 38.6 | 44.2 | n/a | n/a | Mask2Former Table 2 | Yes |
| COCO instance val2017 AP^boundary, ResNet-50 | 23.1 | 30.6 | n/a | n/a | Mask2Former Table 2, boundary quality metric | Yes |
| COCO instance val2017 AP, best model | n/a | 50.1 (Swin-L) | n/a | n/a | Mask2Former Table 2, QueryInst 48.9 in the same table | Yes |
| COCO panoptic test-dev PQ, Swin-L | n/a | 58.3 | n/a | n/a | Mask2Former Table II, MaskFormer 55.4 and K-Net 55.2 in the same table | Yes |
| ADE20K semantic val mIoU, Swin-L | n/a | 57.7 | n/a | n/a | Mask2Former Table 3, MaskFormer 55.6 in the same table | Yes |
| Zero-shot COCO instance AP, box prompts | n/a | n/a | 46.5 | n/a | SAM paper Table 5, SAM given ViTDet-H boxes | Yes, same eval; compare 51.0 for fully supervised ViTDet-H on COCO |
| Zero-shot edge detection ODS, BSDS500 | n/a | n/a | .768 | n/a | SAM paper Table 3, HED reaches .788 trained on BSDS | Yes, one table; SAM never trained on BSDS |
| Zero-shot VOS, SA-V val J&F | n/a | n/a | n/a | 76.8 (Hiera-B+) / 77.9 (Hiera-L) | SAM 2 paper, zero-shot semi-supervised VOS table, XMem 60.1 and Cutie-base+ 61.3 in the same table | Yes |
| Interactive image speed, 1-click 23-dataset average | n/a | n/a | baseline | higher accuracy at 6x speed (Hiera-B+) | SAM 2 paper Table 15 vs SAM ViT-H | Yes, one table |
The rows that justify the “Mask2Former beats Mask R-CNN” claim are the first four: they come from a single Mask2Former training table that retrained and re-evaluated Mask R-CNN baselines under a stated protocol, COCO train2017 only, single-scale inference. That is as close to a controlled comparison as this literature offers. The SAM rows carry the opposite message: SAM’s zero-shot numbers are deliberately compared against fully supervised specialist models in the same evaluation, and they land close but usually below, which is the point, SAM never saw the target data.
Why mask classification beat detect-then-segment
Mask R-CNN’s pipeline is sequential: find boxes, then segment inside each box. Errors in detection cap mask quality, because a missed box means a missed mask, and the mask head only ever sees the feature content of its box. Mask2Former’s queries instead compete to explain the whole image, and each query can produce a mask of any shape anywhere, decoupling mask quality from box quality. Restricting cross-attention to the previous layer’s predicted mask region is the detail that makes it work, the query zooms into its own object instead of averaging features over the scene.
The magnitude is visible in Table 2 of the Mask2Former paper. At ResNet-50, a standard 36-epoch Mask R-CNN reaches 37.2 AP; Mask2Former with the same backbone reaches 43.7 in 50 epochs. Even against a Mask R-CNN pushed to 400 epochs with multi-scale augmentation at 42.5 AP, the default Mask2Former still leads at 43.7, with better boundary quality too, 30.6 versus 28.0 AP^boundary on the long recipe and 23.1 at 36 epochs. One architecture also replaces three: the same design reaches 57.8 PQ on COCO panoptic and 57.7 mIoU on ADE20K semantic at the Swin-L tier, which is why the paper frames the win as reducing research effort by roughly three times, one model to maintain instead of three specialized ones.
SAM plays a different game
SAM’s paper contains a table that looks like a defeat and is actually the argument. On COCO instance segmentation, SAM prompted with ViTDet-H ground-truth-quality boxes reaches 46.5 zero-shot AP, while ViTDet-H, fully trained on COCO, reaches 51.0 with its own detection pipeline. SAM loses by 4.5 points, and that gap is the entire story: a model that never trained on COCO, at all, closes most of the distance to a specialist trained on it. The edge detection row repeats the pattern: 0.768 zero-shot ODS on BSDS500 against 0.788 for HED, a model trained on BSDS itself. Object proposals show the same shape: SAM reaches 59.3 mask AR@1000 against ViTDet-H’s 63.0.
What SAM buys you in exchange for those 4.5 points is the removal of every annotation requirement. If your application involves objects outside COCO’s 80 classes, medical imagery, satellite scenes, industrial parts, SAM’s zero-shot masks are available with a click, at about 50 ms per prompt. Mask R-CNN and Mask2Former give you higher precision but only for the taxonomy you paid to annotate. The models are answers to different questions, which is why “SAM replaces Mask R-CNN” headlines mislead in both directions.
SAM 2 extends the contract to video
Video breaks SAM’s implicit assumption that an image is processed once. A tracked object moves, occludes, and reappears, so SAM 2 adds a memory mechanism: past frames and their prompts are compressed by a memory encoder into a bank, and memory attention lets the current frame attend to that bank before predicting its mask. The result, in the paper’s zero-shot video object segmentation table, is a gap that dwarfs anything in the image tables: on SA-V val, specialized semi-supervised VOS models like XMem reach 60.1 J&F and Cutie-base+ 61.3, while SAM 2, zero-shot and not specialized for this task, reaches 76.8 with the Hiera-B+ model and 77.9 with Hiera-L.
The image side improves at the same time. SAM 2’s Hiera-B+ image-only variant beats SAM ViT-H on one-click point-to-mask accuracy averaged over SAM’s 23-dataset suite while running about 6 times faster, according to Table 15 of the SAM 2 paper, and mixing video data into training lifts the 23-dataset average to 61.4. Since SAM 2 also accepts image prompts and propagates them, an image-plus-video pipeline no longer needs two unrelated models.
When to use which
- You have a fixed label set, annotated training data, and a precision target on it: Mask2Former. The same-harness table is unambiguous, 43.7 against 37.2 at ResNet-50, and one architecture covers instance, panoptic, and semantic segmentation if your needs drift between them.
- You need bounding boxes as a deliverable, or your pipeline is detection-centric: Mask R-CNN. Masks come free from its box pipeline, it is simple, and every CV codebase still ships a prebuilt version. On raw accuracy it trails Mask2Former, but it remains the pragmatic default for classic detection-plus-segmentation products.
- Your objects are not in any dataset, annotation budget is near zero, or a human is in the loop clicking: SAM. Accept that it will land a few points below a fine-tuned specialist on any given benchmark; you are trading that gap for zero training.
- Your input is video, or a mix of images and video: SAM 2. The zero-shot video numbers, 76.8 J&F on SA-V val versus 60.1 for the best specialized baseline in the same table, are the largest single-model margin in this whole comparison, and it keeps SAM-level image ability with a 6x speedup at the B+ tier.
Limits and open questions
Almost every cross-model cell above is a claim placed side by side, not a head-to-head. The cleanest rows are the Mask2Former paper’s own Mask R-CNN retraining and the SAM paper’s same-evaluation zero-shot tables, and even those compare different training regimes, Mask2Former benefited from ImageNet-22K backbone pretraining at the Swin-L tier while its ResNet-50 rows are trained from COCO alone. SAM’s numbers, by design, compare a zero-shot model against supervised specialists, so “close to ViTDet” is not a license to skip evaluation on your own distribution; SAM’s paper itself documents weaker behavior on fine structures like text and on ambiguous prompts where its multiple-mask prediction picks one arbitrarily. Class-agnostic masks are also not the end of the pipeline: SAM gives you no category labels, so a detector or classifier still has to run somewhere downstream. SAM 2’s video memory is finite and its accuracy decays on long videos with heavy occlusion, and its SA-V val lead comes partly from training on SA-V-adjacent data, the zero-shot label means “no fine-tuning on the benchmark”, not “no related training data”. Finally, all numbers here are 2015 to 2024 snapshots; the segmenter you should deploy is the one you benchmark on your data.
FAQ
Which is better on COCO, Mask R-CNN or Mask2Former?
Mask2Former, in every row of the controlled comparison. In the Mask2Former paper’s Table 2, Mask2Former with a ResNet-50 reaches 43.7 instance AP on COCO val2017 against 37.2 for a standard 36-epoch Mask R-CNN with the same backbone, and even a 400-epoch multi-scale Mask R-CNN recipe stops at 42.5. At ResNet-101 the gap is 44.2 versus 38.6, and Mask2Former also wins the boundary quality metric at every backbone tier.
Can SAM replace Mask R-CNN or Mask2Former?
Only if your task tolerates class-agnostic masks and slightly lower precision. On COCO instance segmentation, SAM given ViTDet-H boxes reaches 46.5 AP zero-shot, against 51.0 for the fully supervised ViTDet-H and 50.1 for Mask2Former trained on COCO. SAM’s advantage is that it needs no training data at all and covers arbitrary objects, so it wins when your categories are not in any public dataset or when a human clicks objects interactively, and loses when you need maximum precision on a fixed, annotated taxonomy.
Is SAM 2 better than SAM?
On video, decisively: zero-shot on SA-V val, SAM 2’s Hiera-B+ reaches 76.8 J&F where the specialized XMem model gets 60.1, and SAM has no native video ability at all. On images, SAM 2’s Hiera-B+ beats SAM ViT-H on one-click point-to-mask accuracy averaged over 23 datasets while running about 6 times faster, and training with video data lifts the average further to 61.4. SAM 2 also accepts image prompts and propagates them through a video, so for mixed image-video workflows it strictly dominates keeping SAM.
Why is Mask2Former one architecture for three tasks?
Because it treats every task as mask classification: predict a set of binary masks and assign each a class label, where “instance” and “semantic” are just different label conventions over the same mask set. The mechanism that makes this work is masked attention, where each Transformer decoder layer attends only inside the previous layer’s predicted mask, letting queries refine their own object locally. The paper reports the same architecture reaching 50.1 AP on COCO instance, 57.8 PQ on COCO panoptic, and 57.7 mIoU on ADE20K semantic at the Swin-L tier.
What about U-Net and DeepLab today?
Both remain the right starting points in their niches. U-Net’s encoder-decoder with skip connections is still the default for medical and scientific imagery with small datasets, it won the 2015 ISBI cell tracking challenge with 92.03 IoU versus 83 for the runner-up trained on the same tiny data regime. DeepLab’s atrous convolution and ASPP remain standard components of dense prediction pipelines, and its 79.7 VOC 2012 mIoU defined the pre-Transformer state of the art. Neither competes with Mask2Former on COCO AP, but both train with far less data and far simpler machinery.