ZeroHour
AI model

PANORAMA

1 mentions in 7 days · 1 in 30 days · 1 total · first seen · last

Timeline

PANORAMA: Panoptic Grounded Captioning via Mask Proposal Selection

PANORAMA grounds vision-language caption phrases in pixel-level masks via mask proposal selection, achieving state-of-the-art grounding on the new PanoCaps benchmark.

The paper introduces panoptic grounded captioning, requiring VLMs to describe foreground and background regions while grounding each phrase with pixel-level masks. Contributions include PanoCaps, a human-annotated benchmark built from panoptic segmentation datasets with entity-level image-text alignment, a phrase-mask matching protocol, and a generalized Panoptic Quality metric. PANORAMA conditions a pretrained segmenter on contextualized phrase representations to select masks, achieving the best overall grounding on PanoCaps; code, data, and models are released.

Appears with

Entities are extracted by the model from each article. Watching an entity keeps it in this browser only (no account); the watchlist page and dashboard alerts use it.