ZeroHour

Search: “EmDash”

31 stories

Smart search ranks by meaning as well as keywords (one row per story, last 45 days).

JEPA-Anything: Learning Predictive Models across Different Worlds

Researchers introduce JEPA-Anything, a domain-agnostic predictive world-modeling framework using orthogonal predictive factorization, validated across seven domains including biology and weather.

JEPA-Anything extends joint-embedding predictive architectures through orthogonal predictive factorization (OPF), which decomposes latent targets into complementary factors learned via dedicated pathways. It is evaluated across seven domains: vision, biology, clinical trajectories, control, molecular dynamics, physical fields, and weather. Against matched JEPA baselines it improves metrics on all 10 dynamics tasks, cuts Interventional Pong single-intervention error by 34.8%, and achieves the lowest one-step and 100-step molecular errors across four systems. A factor-nominated biological intervention received experimental support in cell co-cultures, patient-derived organoids, tumor fragments, and mice.

🔬“We have foundation models for language, not for physics” — Anima Anandkumar, Bren Professor of Computing

Caltech professor Anima Anandkumar discusses Neural Operators and FourCastNet for physics modeling, arguing inductive biases beat pure token scaling.

Anima Anandkumar, Bren Professor at Caltech and co-founder of Accelerated Understanding, describes Fourier Neural Operators that learn in frequency and spherical-harmonic domains to model weather, fusion, and fluid or heat flow. Her team built FourCastNet 3, a global weather model competitive with physics-based simulations that runs on consumer-grade GPUs. She also introduced TorchLean, a framework for writing PyTorch-style networks inside the Lean proof assistant for formal verification, and was appointed to the United Nations Scientific Advisory Board. She argues physical domains resist scaling due to tiny datasets and context lengths in the hundreds of billions, so progress comes from built-in structure and physical priors.

Latent Space · 22d agoAI research1

Multi-Vector (Late Interaction) Embedding Models with Sentence Transformers

Hugging Face details building and using multi-vector late-interaction embedding models with Sentence Transformers for retrieval workloads.

Hugging Face published a guide on multi-vector, late-interaction embedding models (ColBERT-style) supported through Sentence Transformers. The post covers how practitioners can build and use these models for retrieval and RAG pipelines. It is a developer tooling and technique write-up, not a security advisory.

Hugging Face Blog · Aug 18, 2026AI tools & infra1

Everybody's Lost Their Minds

A veteran security engineer argues AI-driven vulnerability discovery is not making organizations safer because patching and basic security hygiene remain the real bottleneck.

The author, a long-time security practitioner, criticizes the industry's rush into AI-assisted vulnerability research, noting that Anthropic and OpenAI programs like Glasswing, Daybreak, Athena, and Akrites consumed millions of dollars of engineering time and surfaced thousands of vulnerabilities, only a fraction of which were reported upstream. He argues that finding vulnerabilities was never the bottleneck; patching, asset inventory, and attack surface management remain the true problems. The post also critiques AI hype, agentic workflows, anthropomorphic media coverage, and AI companies' calls for their own regulation.

Training and Finetuning Multi-Vector Embedding Models with Sentence Transformers

Hugging Face published a tutorial on training and finetuning multi-vector embedding models using the Sentence Transformers library.

Hugging Face's blog walks through training and finetuning multi-vector embedding models with Sentence Transformers. Multi-vector approaches store multiple vectors per document to support late-interaction retrieval. The post is a practical guide for developers building retrieval pipelines with the library.

Hugging Face Blog · 23d agoAI tools & infra1

RelateAnything: Real-Time Open-Vocabulary Relation Prediction From Any Inputs

RelateAnything is a 53M-parameter open-vocabulary relation prediction model running at 20 ms/frame, with 2.3-3.5x higher mean recall than comparable open-vocabulary methods.

RelateAnything predicts scored relations between image regions using any predicate vocabulary supplied at inference as text embeddings, with object labels never required as input, so region sources can change without retraining. Training covers 19,103 predicates using positive-unlabeled supervision; the authors release RA-4M (474k images, 4.3M geometrically verified relations over 10,102 free-text predicates) and the OV-SGG-Bench evaluation suite. The 53M-parameter model runs at 20 ms/frame and achieves 2.3-3.5x the mean recall of the strongest comparable open-vocabulary method across cross-dataset and zero-shot benchmarks. Model, corpus, and benchmark are public.

Hugging Face daily papers · 7d agoAI research1

LexFlip: A Dissociation Diagnostic for Legal Meaning Preservation Metrics

LexFlip releases 373 minimal perturbations of Quebec statutory French that reverse legal force while preserving tokens, exposing weaknesses in embedding-based meaning preservation metrics.

LexFlip provides 373 minimal perturbations of Quebec statutory French that reverse legal force while preserving 0.93 of tokens, creating dissociation items that break monotone token-overlap metric validation. The seven embedding and BERTScore metrics tested register only 0.022-0.039 of their identical-to-unrelated range on these edits, versus 0.670 for bidirectional NLI. Against FrJudge, with a measured human ceiling of r=0.597, a bare length feature outscores every semantic metric tested.

arXiv cs.AI / cs.LG / cs.CL · 13d agoAI research

Technical Manual for a Toolkit for Measuring Contextual Individuation in Transformer Language Models

An open methodology toolkit measures whether transformer language models contextualize fixed word forms across domains using bridge forms and layer-wise silhouette analysis.

The manual documents an open toolkit built around 'bridge forms' - identical written words recurring across two or more subject domains with a different sense in each - to test whether transformer language models individuate word occurrences by context beyond the embedding layer. It covers declarative specification of bridge forms, Wikipedia corpus acquisition, occurrence localization, layer-wise representation extraction, domain-pairwise silhouette measurement, and visualization, justifying each choice against failure modes such as sense contamination and subword-tokenization misalignment. It is a methodological and implementation reference and reports no empirical results.

arXiv cs.AI / cs.LG / cs.CL · 13d agoAI research

Augustinian BabyLM: What Ostensive Definition Can and Cannot Teach a Small Language Model

Study shows visually grounded token embeddings in a small masked LM persist through training and improve object-property knowledge, but escape standard BabyLM benchmarks.

The paper implements ostensive definition for a small DeBERTa masked language model trained on 10M words, seeding visually grounded tokens with embeddings derived from labeled image regions before training. Visual initialization leaves a persistent, seed-replicated advantage on object-property knowledge (COMPS) and a corpus-tailored Visual-Property Swap benchmark covering color, material, size, and shape, but has no effect on most BabyLM grammar benchmarks. Synthetic grounding of previously unseeded words causally transfers the advantage to exactly those words.

arXiv cs.AI / cs.LG / cs.CL · 7d agoAI research

Structured Claim-Level Discourse Representations for Dense Health Narratives

Researchers propose a claim-level discourse framework for health videos, finding 13.22 atomic claims per minute and that LLMs struggle with pragmatic profiling.

The paper introduces a structured framework for claim-level discourse analysis in dense health narratives on social media videos, modeling tuples that link atomic claims with thematic aspects, stance, and multidimensional pragmatic attributes. Analysis found an average of 13.22 atomic claims per minute in health video discourse. A benchmark spanning four health domains with 1,191 manually annotated claims from 60 videos shows current LLMs perform strongly on thematic categorization and stance prediction but struggle with high-dimensional pragmatic profiling, suggesting future systems need task decomposition and specialized inference strategies.

arXiv cs.AI / cs.LG / cs.CL · 1d agoAI research

Claude Fable 5.1's language is less "load-bearing" than its predecessor's

Arena.ai found Claude Fable 5.1 writes 30% longer answers with fewer em dashes, hedges, and validation phrases than Fable 5.

Arena.ai compared tens of thousands of high-reasoning Text Arena outputs from Anthropic's Claude Fable 5 and Fable 5.1. Fable 5.1's median answer length rose 30% to 414 words (from 319), remaining 21% shorter than Opus 5's 525. Em-dash usage fell from 16.2 to 11.0 per 1,000 words while semicolons rose to 6.09, hedges like 'perhaps' and 'arguably' dropped 36%, and praise appeared in 1.98% of responses versus 3.17%. Long content words fell from 42.6% to 38.6% and abstract nouns declined 25%, suggesting the successor's language is less 'load-bearing'.

The Decoder · 7d agoAI research1

Beyond the Turing threshold: Productive grammars generate essentially undecidable languages

A theoretical paper designs formal grammars that emulate Post's productive sets, generating languages that are provably beyond Turing decidability.

The paper elaborates on Emil Post's productive sets, which are not even semi-computable, and builds formal grammars that emulate their construction over natural numbers. The resulting languages are shown to be essentially undecidable, placing them beyond Turing decidability. This is pure computability and formal language theory with limited direct security relevance.

arXiv cs.CR · 8d agoResearch1

When EOS Tokens Disagree: Understanding Length Inflation in On-Policy Distillation

Researchers show EOS termination-token mismatch between base students and post-trained teachers drives length inflation in on-policy distillation across Qwen3, Llama, and Gemma.

The paper studies length inflation in on-policy distillation (OPD), where student responses become excessively long and can exhaust the generation budget. The authors identify termination-token mismatch between base students and post-trained teachers as an important cause, observing that Qwen3, Llama, and Gemma place stopping probability on different EOS tokens even with identical declared stopping sets. Treating functionally equivalent EOS tokens as a shared semantic stopping action substantially mitigates mismatch-induced inflation across all three families, while aligning decoding stopping sets alone is insufficient. A stage-wise analysis of K2-Horizon training shows termination preferences shift during training, and an implementation with the proposed corrections is released.

Introducing OlmoEarth embeddings: Custom embedding exports from OlmoEarth Studio for downstream analysis

AllenAI's OlmoEarth Studio adds custom embedding exports to support downstream geospatial analysis workflows.

A Hugging Face blog post from AllenAI introduces OlmoEarth embeddings, a feature allowing custom embedding exports from OlmoEarth Studio for downstream analysis tasks. Only the title was available, so no benchmark or performance details are provided. OlmoEarth is Ai2's open geospatial AI model family.

Hugging Face Blog · Aug 12, 2026AI tools & infra

Register Tokens for Bounded-State Reasoning in Diffusion Language Models

Register tokens let diffusion language models like LLaDA and Dream carry reasoning state across cleared chunks, gaining up to 19.5 points on code.

Researchers propose register tokens: dedicated fixed-position tokens whose continuous hidden states are trained to carry reasoning progress across generation chunks in masked diffusion language models. After decoding and clearing a chunk, the model continues from the prompt and the carried register state instead of retaining earlier text. On LLaDA and Dream, registers outperform discrete-text carry on every benchmark, with gains up to 8.5 points on math and 19.5 points on code. Registers are especially effective for bounded code generation and can be further refined with reinforcement learning on long-horizon reasoning tasks.

Hugging Face daily papers · 4d agoAI research

ENEAS: Embedding-guided Neural Ensemble for Adaptive Segmentation

ENEAS adds text prompting and semantic verification to video segmentation to keep tracking targets through occlusion and reject lookalike distractors.

ENEAS is a unified text-promptable method for instance tracking and open-concept semantic discovery in video, designed to fix temporal hallucinations, spatial fragmentation, and semantic misclassification seen in SAM 3-class foundation models. It extends the geometrically robust SeC architecture with a text-prompting adapter and temporal memory, and uses a verification layer combining fast visual embedding matching with conditional VLM refinement for ambiguous candidates. It targets 3D reconstruction pipelines where a single misclassified distractor corrupts the asset. Code and models are open-sourced.

Hugging Face daily papers · 15d agoAI research

MUSE: Benchmarking Large Vision-Language Models on Multi-Modal Understanding in Situated Education

Introduces MUSE, a twelve-task benchmark evaluating vision-language models on artistic image understanding in situated educational, Southeast Asian contexts.

MUSE is a benchmark assessing large vision-language models on artistic image understanding across twelve tasks spanning visual perception, semantic and affective interpretation, cultural understanding, and compositional reasoning. It decouples image annotation from question generation for controllable difficulty and curates images centering Singaporean and Southeast Asian multicultural contexts alongside Western art. Evaluations of open-source and proprietary models found substantial disparities, especially in affective interpretation and compositional reasoning.

arXiv cs.AI / cs.LG / cs.CL · 1d agoAI research1

Reason Through the Latent! Making Latent Visual Reasoning Necessary

Researchers introduce CVRR, forcing multimodal models to rely on recurrent latent computation rather than accessible image tokens, validated via causal interventions and benchmarks.

The paper presents Causal Visual Recurrent Reasoning (CVRR), which makes recurrent hidden-state computation the required image-conditioned path for prediction in vision-language models. Before decoding, visual states and the original multimodal KV cache are removed so only the final recurrent state carries image information to the answer. CVRR retains strong performance on V*, MMVP, BLINK, and MME-RealWorld-Lite while comparable latent reasoners fail under the same constraint. Causal interventions show predictions remain sensitive to recurrent content and that persistent visual evidence causally revises the recurrent trajectory.

Hugging Face daily papers · 12d agoAI research

Show HN: LLM Attention Visualization

A developer released a browser-based tool that visualizes which past tokens influence each LLM output token using aggregated, value-weighted attention scores.

A Show HN project presents a React application built on Transformers.js that renders per-token attention influence by aggregating attention weights scaled by value-vector magnitudes across all attention heads and layers. To expose internal tensors, the author instrumented the ONNX computation graph, hosted a modified model on Hugging Face, and pre-generated prompts to avoid long model downloads in the browser. Demos with a 600-million-parameter model show how verbatim copying draws heavily on source tokens and how single outputs blend information from multiple phrases.

Characterizing Language Generation in the Limit: Finite Witnesses and a Separation-Width Hierarch

New work characterizes language generation in the limit via finite witnesses, proves a full separation-width hierarchy, and formalizes all results in Lean.

The paper fully characterizes when language generation in the limit is possible for arbitrary families over a countable universe: each target must admit a finite positive witness such that targets activated by any finite sample share an infinite common intersection. It defines positive separation width and proves every level of the resulting hierarchy occurs, with countable families admitting singleton witnesses and unions of families with infinite common cores requiring unbounded finite witnesses. The characterization, a universal normalization, and a diagonal capture lemma are machine-checked in the Lean proof assistant, with the development maintained on GitHub.

arXiv cs.AI / cs.LG / cs.CL · 8d agoAI research1

CausalArena: Benchmarking Causal Discovery in the Foundation Model Era

Researchers introduce CausalArena, a unified benchmark revealing that causal discovery rankings shift substantially across structural causal model families and protocols.

The paper presents CausalArena, a unified and evolvable benchmark for causal discovery combining synthetic structural causal models, semantically grounded operational SCMs, formula-grounded scientific SCMs, and public real-world datasets. Experiments across classical, neural, and pretrained causal discovery foundation models show large ranking shifts between benchmark regimes. The authors identify pretraining-evaluation overlap and benchmark diversity as central evaluation challenges.

arXiv cs.AI / cs.LG / cs.CL · 7d agoAI research

Unifying Models of Intergroup Hostility in Online Discourse

Researchers unify six theories of intergroup hostility using 2.86 million TikTok, Truth Social, and Twitter/X posts, finding boundary and threat construction anchor rhetoric.

The study models mechanisms from six foundational theories of intergroup hostility - boundary construction, threat construction, scapegoating, negative evaluation, dehumanization, and action orientation - within a common empirical framework using 2.86 million posts from TikTok, Truth Social, and Twitter/X during the 2024 U.S. presidential election. Structurally, boundary and threat construction anchor the system; temporally, boundary construction, derogation, and action orientation appear early, dehumanization and threat construction later, and scapegoating last. The work aims to give computational social science a unified empirical basis for modeling hostile rhetoric beyond single-label detection.

arXiv cs.AI / cs.LG / cs.CL · 20h agoAI research

One Symptom, Three Levers: A Critical Review of On-Policy Self-Distillation

A review paper frames on-policy self-distillation collapse as governed by three levers: token weighting, privileged information, and guidance decay.

The paper critically reviews On-Policy Self-Distillation (OPSD), where a language model trains on its own generations scored token-by-token by a teacher conditioned on privileged information such as reference solutions or environment feedback. It identifies collapse, the progressive narrowing of producible reasoning paths, as the dominant failure mode and analyzes it through three levers: signal weighting, the nature of privileged information, and teacher dynamics. The review is restricted to mathematical reasoning, reports no new experiments, and offers a shared vocabulary separating settled findings from disputed ones.

Hugging Face daily papers · 23d agoAI research

Thought without systematicity? Evaluating reasoning models on rule induction tasks

Study finds reasoning models often fail on structurally equivalent variants of tasks they solve, suggesting their reasoning lacks systematicity.

The paper extends rule induction tasks from cognitive science using task isomorphisms such as recombination and substitution to test systematicity in reasoning models. Despite solving tasks correctly, models frequently fail on structurally equivalent variants of the same task. The authors conclude many model behaviors lack systematicity, making it difficult to establish cognitive abilities beyond the specific evaluation contexts.

Hugging Face daily papers · 6d agoAI research

Objective vs. Search: Decomposing What Makes a Good Tokeniser

New tokeniser study shows search procedure, not optimisation objective, drives bits-per-byte performance across model sizes, vocabulary sizes, and multilingual settings.

The paper disentangles BPE and UnigramLM along two axes: optimisation objective (compression vs log-likelihood) and search procedure (bottom-up merging vs top-down pruning). Two new algorithms, BottomUpLL and TopDownComp, complete the 2x2 design space, and trained language models are evaluated on bits-per-byte and BLiMP across model sizes, vocabulary sizes, and English-only vs multilingual domains. Bottom-up tokenisers consistently achieve lower bits-per-byte in most settings, while BLiMP shows no consistent relationship with design choice.

arXiv cs.AI / cs.LG / cs.CL · 1d agoAI research

On-Demand Attention: Language Models Know When to Recall

On-Demand Attention uses a lightweight recall head to selectively invoke global attention, cutting long-context decoding costs in vLLM with minimal quality loss.

Researchers introduce On-Demand Attention (ODA), a local-first decoding method in which a lightweight recall head predicts when global attention benefits the next token. Only the recall head is trained, leaving pretrained weights and the complete KV cache unchanged. GPU-side conditional execution implemented in vLLM converts reduced global reads into practical decoding speedups at long context lengths. Experiments across Qwen and Gemma models, including hybrid-attention backbones, show selective recall recovers most of the performance lost under local attention.

arXiv cs.AI / cs.LG / cs.CL · 20h agoAI research

Inoculation Midtraining with Learned Neologisms

Inoculation Midtraining confines unsafe LLM behavior to a neologism-marked context, reducing misalignment after unsafe post-training but leaking under nearby contextual cues.

The paper introduces Inoculation Midtraining, which teaches a base model during midtraining that unsafe behavior belongs to a context marked by a learned neologism token, then post-trains on unsafe data within that context. Across supervised fine-tuning and RL post-training regimes, the technique reduces misalignment while preserving transfer of benign properties like German or Shakespearean prose. However, it does not outperform standard Inoculation Prompting, is sensitive to training configuration, and produces a leaky boundary that nearby contextual cues can reactivate. The authors conclude it is not yet a load-bearing component of a developer safety framework.

You Can't Prefer Emotions You Don't Sample: Intensity Undershoot in DPO-Tuned LLMs

Study quantifies DPO-tuned LLMs undershooting requested emotional intensity, tracing the gap to candidate-pool extremity rather than conditioning format.

Conditioning an instruction-tuned LLM on continuous valence-arousal targets yields gain of only 0.26 for valence and 0.13 for arousal on Llama-3.1-8B, far below faithful control of 1.0. The authors attribute undershoot to neutral-heavy preference corpora like EmoBank and candidate pools lacking extreme affect, leaving DPO without extreme exemplars. Uniform target coverage with a hotter candidate pool raises valence gain to 0.40 on Llama-3.1-8B and 0.44 on Qwen3-8B, with modest in-distribution cost; arousal gains remain unstable across seeds.

arXiv cs.AI / cs.LG / cs.CL · 10d agoAI research

NCP-ArchPreview Technical Report: Moving towards Latent Space Language Models through Next Concept Prediction

An 8.9B-parameter latent-space language model using next-concept prediction matches OLMo-3-7B pretraining loss with only 51.3% of the training tokens.

NCP-ArchPreview augments next-token prediction with Next Concept Prediction over a product-quantized concept vocabulary built from hidden states, trained jointly end-to-end. The 8.9B model was trained on 5.73T tokens from the Dolma-3 dataset, the largest latent-space language model demonstration to date. It consumes 51.3% of the tokens to reach OLMo-3-7B's final pretraining loss and outperforms it by 2.45 points on the downstream macro-average, including a 5.99-point GSM8K gain. The learned latent space also enables lightweight domain adaptation via a 17M-parameter VQ module and improves speculative drafting accepted length by 4.17%.

Hugging Face daily papers · 9d agoAI research1

Generative Marketing Mix Modeling: A Causal Inference Framework Linking GEO and GEM to Business Impact

Paper proposes Generative Marketing Mix Modeling to causally estimate Generative Engine Optimization and Marketing effects on business outcomes.

The authors develop GMMM, a causal inference framework for measuring how often users see and notice a firm's name in generated answers, which standard marketing data ignore. For GEO it combines repeated generated answers with question counts, shares of generative-system usage and notice probabilities; for GEM it uses sponsored placement records with notice probabilities. The framework compares expected business responses under alternative treatment sequences, establishes identification conditions, and is evaluated on simulated product-recommendation answers in English and Japanese.

arXiv cs.AI / cs.LG / cs.CL · 7d agoAI research1

Molecular Déjà Vu: Digit-Level Retrieval of Published Values in Frontier Language Models

Audit of 22 frontier models finds widespread verbatim retrieval of published molecular property values, with higher reasoning increasing recall of memorized numbers.

An arXiv audit tests 22 frontier LLMs across 12 molecular regression benchmarks for verbatim retrieval of published values. More than 50% of the LLMs show verbatim retrieval on five datasets, and identical experiments are flagged 89% more often at a high reasoning level than at the lowest one. Suppressing retrieval moves model prediction errors closer together in relative terms, suggesting predictive capability is not determined solely by memorized values.

arXiv cs.AI / cs.LG / cs.CL · 13d agoAI research1