NeurIPS 2026 · Evaluations & Datasets Track
When Predicting Nothing Beats SAM 3:
Revisiting Evaluation in Video Object Segmentation
AIDAS Lab, Seoul National University
† Corresponding author
Can you beat SAM 3?
Click the highlighted cow whenever you see it, or do nothing. You have 12 seconds.
For keyboard play, use Mark target visible with Space or Enter.
Hint: patience might win this one.
Clicks mark presence only. Scores cover the full cow sequence, excluding the prompt frame.
| Predictor | Frame-wise J&F\(\mathcal{J}\&\mathcal{F}\) |
|---|---|
| Predict nothing | 82.8 |
| SAM 3 | 77.5 |
Why nothing wins
When targets are mostly absent, absence classification can dominate frame-wise J&F\(\mathcal{J}\&\mathcal{F}\).
Explore the frame-wise score
J&F = τ·p_1·s + (1−τ)·p_0\[\mathcal{J}\&\mathcal{F}=\tau p_1 s+(1-\tau)p_0\]
Fraction of target-absent frames correctly predicted empty.
Fraction of target-present frames correctly predicted non-empty.
Mean segmentation quality when both masks are non-empty.
Share of the score
At τ\(\tau\) = , absence makes up 83% of the score.
Volumetric J&F\(\mathcal{J}\&\mathcal{F}\)
Score whole mask volumes: jointly empty frames add nothing; false positives still cost.
From masks to a volume
Frame weighting
FaVOS benchmark
FaVOS-20 and FaVOS-40 average about 20% and 40% visible frames per object.
More results
Rankings are largely preserved at high visibility and robust across frame weightings.
FaVOS results
First-frame mask prompts; object-wise means on a 0–100 scale.
| Model | J&F ↑\(\mathcal{J}\&\mathcal{F}\uparrow\) | J&F_v ↑\(\mathcal{J}\&\mathcal{F}_v\uparrow\) |
|---|---|---|
| Empty predictor | 80.0 | 0.0 |
| STM | 71.9 | 35.5 |
| STCN | 50.6 | 26.0 |
| XMem | 63.8 | 36.8 |
| DeAOT-L | 59.9 | 38.0 |
| Cutie-B | 80.4 | 49.0 |
| SAM 2.1-L | 73.4 | 46.7 |
| SAMURAI-L | 60.8 | 37.3 |
| DAM4SAM-L | 69.6 | 44.3 |
| SAM2Long-L | 74.0 | 49.8 |
| SeC | 80.6 | 55.7 |
| SAM 3 | 78.7 | 56.0 |
| Model | J&F ↑\(\mathcal{J}\&\mathcal{F}\uparrow\) | J&F_v ↑\(\mathcal{J}\&\mathcal{F}_v\uparrow\) |
|---|---|---|
| Empty predictor | 59.9 | 0.0 |
| STM | 65.9 | 44.9 |
| STCN | 53.9 | 37.7 |
| XMem | 62.0 | 45.3 |
| DeAOT-L | 62.7 | 49.2 |
| Cutie-B | 76.7 | 58.6 |
| SAM 2.1-L | 80.1 | 60.5 |
| SAMURAI-L | 72.1 | 54.9 |
| DAM4SAM-L | 78.0 | 63.2 |
| SAM2Long-L | 78.3 | 64.8 |
| SeC | 84.0 | 70.9 |
| SAM 3 | 84.4 | 72.7 |
BibTeX
@inproceedings{hong2026favos,
title = {When Predicting Nothing Beats SAM 3: Revisiting Evaluation in Video Object Segmentation},
author = {Hong, Jihwan and Park, Woohyeon and Kim, Jaeik and Do, Jaeyoung},
booktitle = {Advances in Neural Information Processing Systems},
year = {2026},
note = {Evaluations \& Datasets Track}
}