Skip to content
AIDAS Laboratory Seoul National University

NeurIPS 2026 · Evaluations & Datasets Track

When Predicting Nothing Beats SAM 3:
Revisiting Evaluation in Video Object Segmentation

Jihwan Hong Woohyeon Park Jaeik Kim Jaeyoung Do†

AIDAS Lab, Seoul National University
† Corresponding author

Can you beat SAM 3?

Reference: the highlighted cow from Figure 1, among rocks and other cattle

Click the highlighted cow whenever you see it, or do nothing. You have 12 seconds.
For keyboard play, use Mark target visible with Space or Enter.

Hint: patience might win this one.

Clicks mark presence only. Scores cover the full cow sequence, excluding the prompt frame.

VideoPredict nothingSAM 3
SAM 3 follows another animal after the highlighted cow leaves.
This video · highlighted cow, full sequence
PredictorFrame-wise J&F
Predict nothing82.8
SAM 377.5

Why nothing wins

When targets are mostly absent, absence classification can dominate frame-wise J&F.

Explore the frame-wise score

J&F = τ·p_1·s + (1−τ)·p_0

Fraction of target-absent frames correctly predicted empty.

Fraction of target-present frames correctly predicted non-empty.

Mean segmentation quality when both masks are non-empty.

AbsenceSegmentation

Share of the score

Absence and segmentation contributions across temporal visibility With p0, p1, and s all 0.80, contributions cross at visibility 0.56. At visibility 0.20, absence contributes 83.3% and segmentation contributes 16.7%. 000.250.250.50.50.750.7511
Dashed line: τ* = p0 / (p0 + p1·s) = 0.56.

At τ = 20%, absence makes up 83% of the score.

Volumetric J&F

Score whole mask volumes: jointly empty frames add nothing; false positives still cost.

From masks to a volume

Final equation for the swimming pinniped excerpt: J&F_v = (J_v + F_v) / 2, approximately 0.927. J_v = 33,879 / 37,403, approximately 0.906; F_v approximately 0.947. Counts use 33 consecutive masks downsampled to 160 by 90; jointly empty time adds nothing.
Swimming pinniped · J_v ≈ 0.906, F_v ≈ 0.947, J&F_v ≈ 0.927 for this excerpt (33 masks, downsampled to 160 × 90). Surfaces smoothed for display; intro at 2×.

Frame weighting

Approaching train, FaVOS-40: the original linear volume forms a horn as union area grows 51-fold. Red is intersection; blue is non-overlap. An inset shows real early, middle and late silhouettes at one scale, with union areas relative to their median. This excerpt: video f91be6e652955638cfcd765284efbd64, object 1, frames 0–20, masks downsampled to 160 by 120. Surfaces smoothed for display.
Train · FaVOS-40 · 51× union-area growth. J_w for this excerpt (21 masks at 160 × 120). Silhouettes retained throughout; surfaces smoothed for display.

FaVOS benchmark

FaVOS-20 and FaVOS-40 average about 20% and 40% visible frames per object.

Drummer · FaVOS-20
Elephant sculpture · FaVOS-20
Large fish · FaVOS-40
Parrot · FaVOS-40
More results

Rankings are largely preserved at high visibility and robust across frame weightings.

FaVOS results

First-frame mask prompts; object-wise means on a 0–100 scale.

FaVOS-20
ModelJ&F ↑J&F_v ↑
Empty predictor80.00.0
STM71.935.5
STCN50.626.0
XMem63.836.8
DeAOT-L59.938.0
Cutie-B80.449.0
SAM 2.1-L73.446.7
SAMURAI-L60.837.3
DAM4SAM-L69.644.3
SAM2Long-L74.049.8
SeC80.655.7
SAM 378.756.0
FaVOS-40
ModelJ&F ↑J&F_v ↑
Empty predictor59.90.0
STM65.944.9
STCN53.937.7
XMem62.045.3
DeAOT-L62.749.2
Cutie-B76.758.6
SAM 2.1-L80.160.5
SAMURAI-L72.154.9
DAM4SAM-L78.063.2
SAM2Long-L78.364.8
SeC84.070.9
SAM 384.472.7

BibTeX

@inproceedings{hong2026favos,
  title     = {When Predicting Nothing Beats SAM 3: Revisiting Evaluation in Video Object Segmentation},
  author    = {Hong, Jihwan and Park, Woohyeon and Kim, Jaeik and Do, Jaeyoung},
  booktitle = {Advances in Neural Information Processing Systems},
  year      = {2026},
  note      = {Evaluations \& Datasets Track}
}