FLAT: Resampling Image and Text into 1D Flexible-Length Aligned Transmodal Tokens for Retrieval and Generation
FLAT jointly trains a multimodal encoder with text-to-image and image-to-text decoders, producing flexible-length tokens that hit 83.1 GenEval on T2I after fine-tuning.
FLAT (Flexible-Length Aligned Transmodal representations) is a pre-training framework that jointly optimizes a shared multimodal encoder with T2I and I2T decoders, combining contrastive alignment with bidirectional cross-modal generative objectives. It maps visual and textual inputs into a unified continuous 1D sequence space and uses nested dropout over prefix-K tokens for dynamic output lengths. A single pre-training stage supports cross-modal retrieval and generation (71.1 GenEval), with task-specific fine-tuning reaching 83.1 GenEval on T2I, 40.5 BLEU-4 and 138.6 CIDEr on MS-COCO captioning, and strong Recall@5 on MS-COCO and Flickr30K.
Google’s new search redirects make links harder to check before you click
Google routes some search result links through encoded google.com/goto?url= redirects, breaking hover-to-verify link checks and raising scraping costs.
Google has started routing some search result links through opaque google.com/goto?url=... redirects using a custom Google-specific encoding, with the destination visible only via the redirect response's Location header. Google confirmed the rollout as an anti-abuse measure, most likely to make bulk extraction of destination URLs from search results more difficult and costly. Malwarebytes warns the change undermines the standard hover-before-clicking safety advice, while legitimate rank-tracking, SEO auditing, archival, and accessibility tools now face the same rate limits and costs as abusive scrapers.
Rare Not Random Using Token Efficiency for Secrets Scanning
Researcher proposes token efficiency (string length divided by BPE token count) as a better post-regex filter than entropy for secrets scanning, validated on CredData.
The post explores whether Byte-Pair Encoding tokenization can replace Shannon entropy as the primary filter for candidate secrets captured by regex in tools like Gitleaks. It defines 'token efficiency' as string length divided by token count under the cl100k_base tokenizer; secret-like strings such as GitHub tokens tokenize into many small tokens and score low, while natural text scores high. Evaluating labeled secrets from the CredData dataset shows a usable separation, with roughly 2.5 suggested as a minimum cutoff versus Gitleaks' 3.5 entropy threshold. The technique is positioned as a post-regex filtering step rather than a standalone detector.
YuE2 · Frontier Music with Symbolic Planning
YuE2, a 3.59B-parameter music generation model, scores 6.9632 on SongBench, beating Suno v5 via symbolic planning.
YuE2 is a music generation model of roughly 3.59B parameters and 28 layers supporting song creation, covering, and agentic editing through editable ABC symbolic scores. Its best-of-8 setting reaches 6.9632 on SongBench, the highest mean among 15 evaluated settings on WildSongBench (192 prompts), ahead of Suno v5 at 6.8721. The project also introduces MERT2, whose 632M-parameter encoders achieve state of the art on 14 of 15 MARBLE metrics, and SheetSage2, which transcribes beats, downbeats, key, chords, structure, and melody with SOTA on 10 of 13 benchmark metrics.
Generative Late-Interaction Embeddings For Visual Document Retrieval
GLIE compresses visual document retrieval embeddings to four vectors per page while retaining nearly 80% of uncompressed nDCG@5 accuracy.
Researchers analyzing late-interaction retrieval embeddings found they lie exactly on the unit sphere and concentrate near a manifold of intrinsic dimension five to six. GLIE learns a few k vectors per page that serve as a lightweight index and a basis to regenerate the full embedding set for exact rescoring of top candidates at query time. On ViDoRe v1 with four vectors per page, GLIE retains nearly 80% of uncompressed nDCG@5 versus 70% for the best prior post-hoc method, using a 415K-parameter network trained in under three GPU-minutes on 1,000 pages.
m-a-p/YuE2-3B — new model trending #30 on Hugging Face
M-A-P released YuE2-3B, an open music generation model that outperforms Suno v5 on WildSongBench and runs locally on a 24GB GPU.
The M-A-P (multimodal-art-projection) team released YuE2-3B, an open-weights music generation model that turns lyrics and a style prompt into full songs with vocals and accompaniment. It uses an AR-NAR Mixture-of-Transformers backbone with symbolic planning and flow matching through a VAE, and supports editable scores (melody and chords, including ABC notation) plus agentic editing workflows. On 192 WildSongBench prompts it reports a SongBench average of 6.9632 (best-of-8) versus 6.8721 for Suno v5, claimed as state of the art among evaluated open and proprietary models. It runs 48 kHz stereo inference locally on a single 24GB NVIDIA GPU without quantization, with companion releases including YuE2-Vae, MERT-v2 encoders, the WildSongBench dataset, and SheetSage2.
Why Is Video Still So Expensive? A Survey of Inference-Efficiency Mechanisms in Video and Audiovisual LLMs
A survey catalogs inference-efficiency techniques for video and audiovisual LLMs, mapping bottlenecks in sampling, encoding, token reduction, and LLM decoding.
This survey covers inference-efficiency mechanisms for visual and audiovisual video LLMs, reporting reductions in parameters, FLOPs, latency, memory, and token counts. It organizes methods by pipeline stage, covering frame sampling, modality encoding, connector-level token reduction, and LLM prefilling and decoding for systems built since late 2022. The authors compile accuracy-cost comparisons under shared host models and input protocols, identify gaps in audiovisual efficiency and standardized evaluation, and maintain a public repository.