ZeroHour

Search: “pixel”

2 stories in the last 24h

PANORAMA: Panoptic Grounded Captioning via Mask Proposal Selection

PANORAMA grounds vision-language caption phrases in pixel-level masks via mask proposal selection, achieving state-of-the-art grounding on the new PanoCaps benchmark.

The paper introduces panoptic grounded captioning, requiring VLMs to describe foreground and background regions while grounding each phrase with pixel-level masks. Contributions include PanoCaps, a human-annotated benchmark built from panoptic segmentation datasets with entity-level image-text alignment, a phrase-mask matching protocol, and a generalized Panoptic Quality metric. PANORAMA conditions a pretrained segmenter on contextualized phrase representations to select masks, achieving the best overall grounding on PanoCaps; code, data, and models are released.

DeepSeek-v4.1 Flash: Pushing the Limits of KV Cache Compression

DeepSeek-V4.1 Flash is a 552B-parameter multimodal MoE model with 1M-token context achieving 4x KV cache compression for long-horizon agent workloads.

A detailed analysis of the DeepSeek-V4.1 Flash technical report describes a 552B-parameter multimodal mixture-of-experts model supporting contexts up to 1 million tokens. Its Causal Encoder-Decoder (CED) architecture activates 8B parameters during prefill and 16B during decode, and reportedly delivers about 420 tokens/s. Joint optimization of architecture (CSA2 cross-layer compression), FP4 KV cache precision, and deployment strategy cuts runtime KV cache to roughly 1/4 and persistent KV cache to about 1/8 of DeepSeek-V4-Flash at the same sequence length, targeting storage and bandwidth bottlenecks in long-horizon agent serving. The author notes all DeepSeek-V4 Pro models were taken offline following the release.