ZeroHour

News

31 stories in the last 24h

Google Research Introduces Retrieve-for-Train (R4T): An RL-Compiled Diffusion Retriever for 12× to 20× Faster Query Fan-Out

Google Research introduced R4T, an RL-trained fan-out pipeline distilled into a 53.9M-parameter diffusion retriever achieving 12x-20x faster query fan-out.

Google Research introduced Retrieve-for-Train (R4T), which trains a fan-out language model with GRPO plus soft PPO regularization, then distills query fan-out into a 53.9M-parameter diffusion transformer that generates all retrieval embeddings in a single non-autoregressive pass. A three-term reward (groundedness 0.6, diversity 0.2 via Vendi Score, alignment 0.2) prevents paraphrastic collapse and reward hacking during training. On the Polyvore dataset, Gemma3-4B R4T-FOLM averaged 49.1 versus 40.9 for Best-of-N, and the diffusion retriever cut fan-out latency from 1.46s to 0.07s at batch size 8, a consistent 12x-20x speedup over autoregressive methods.

MarkTechPost · 7h agoAI research1

Nunchux AI Introduces VC-Attention: A Training-Free Low-Bit Attention Kernel That Speeds Up Video Diffusion Transformers

Nunchux AI introduces VC-Attention, a training-free low-bit attention kernel that speeds up video diffusion transformers up to 3.58x.

Nunchux AI unveiled VC-Attention, a training-free attention kernel for video Diffusion Transformers combining V-Smooth (k-means value-token grouping with block-mean residual quantization) and ExpCast-FP8 (single multiply-add softmax exponentiation). Benchmarks on Wan2.2-T2V-A14B, LongCat-Video, HunyuanVideo-1.5, and MiniMax-H3 show 1.59x attention speedup on B200 at 8-bit and 3.58x on RTX 5090 at 4-bit, with end-to-end gains up to 1.70x. It beats SageAttention2 by 2.3 dB PSNR on Wan2.2 at 8-bit and SageAttention3 by up to 3.6 dB at 4-bit. No public kernel release yet; a proprietary extension runs in Nunchux's stack.

MarkTechPost · 12h agoAI research 2 sources1

HarnessTax: How Much Does the Harness Matter for Coding Agents?

HarnessTax is a research project measuring how much the harness, the scaffolding around LLMs, affects coding agent performance.

HarnessTax examines how much the harness — the scaffolding, prompts, and tooling wrapped around a large language model — contributes to coding agent results, as opposed to the underlying model itself. The project was posted on Hacker News on September 16, 2026, where it drew 42 points and 9 comments. Further details are available on the project's GitHub Pages site.

Stanford Researchers Release Paper2Agent: Turning Research Papers Into AI Agents That Reproduce Results and Run on New Data

Stanford researchers released Paper2Agent, a Nature-published pipeline that turns research papers into MCP servers agents can execute.

A Stanford team led by Jiacheng Miao and James Zou published Paper2Agent in Nature on 16 September 2026. Built on Claude Code's agent SDK, it converts a paper and its codebase into a Model Context Protocol server with validated tools, resources, and prompts. In benchmarks, the AlphaGenome agent built 22 tools in about 45 minutes for US$14, scored 100% on 15 novel queries versus 78.7% for Claude Code with repository access, and cut median runtime 1.9x. In scale tests, 74 of 100 bioRxiv papers were converted and 593 of 599 proposed tools passed validation.

MarkTechPost · 15h agoAI research1

Breaking the 1.58-bit Barrier for Ternary LLMs

An arXiv paper claims a method that breaks the 1.58-bit barrier for ternary large language models.

The arXiv preprint 2609.16338, titled 'Breaking the 1.58-bit Barrier for Ternary LLMs,' presents research on ternary-weight large language models, which use roughly 1.58 bits per weight. The source text contained only the title and Hacker News engagement data (56 points, no comments), so further technical details are not available.

Objective vs. Search: Decomposing What Makes a Good Tokeniser

New tokeniser study shows search procedure, not optimisation objective, drives bits-per-byte performance across model sizes, vocabulary sizes, and multilingual settings.

The paper disentangles BPE and UnigramLM along two axes: optimisation objective (compression vs log-likelihood) and search procedure (bottom-up merging vs top-down pruning). Two new algorithms, BottomUpLL and TopDownComp, complete the 2x2 design space, and trained language models are evaluated on bits-per-byte and BLiMP across model sizes, vocabulary sizes, and English-only vs multilingual domains. Bottom-up tokenisers consistently achieve lower bits-per-byte in most settings, while BLiMP shows no consistent relationship with design choice.

arXiv cs.AI / cs.LG / cs.CL · 19h agoAI research

A Zeroth-Order Paradigm for LLM Preference Alignment

Researchers propose ComPO, a zeroth-order comparison-based preference alignment method with convergence guarantees that mitigates likelihood displacement in LLMs.

ComPO extracts directional information from preference pairs via comparison oracles instead of optimizing a differentiable preference loss, addressing likelihood displacement in direct alignment methods. The paper establishes convergence guarantees for the offline scheme and introduces an online variant with reverse-KL control using unlabeled policy generations. Experiments on Mistral, Llama, Gemma-2, Gemma-3, and Qwen3 models show improvements over existing direct alignment methods, including length-controlled win rates.

PANORAMA: Panoptic Grounded Captioning via Mask Proposal Selection

PANORAMA grounds vision-language caption phrases in pixel-level masks via mask proposal selection, achieving state-of-the-art grounding on the new PanoCaps benchmark.

The paper introduces panoptic grounded captioning, requiring VLMs to describe foreground and background regions while grounding each phrase with pixel-level masks. Contributions include PanoCaps, a human-annotated benchmark built from panoptic segmentation datasets with entity-level image-text alignment, a phrase-mask matching protocol, and a generalized Panoptic Quality metric. PANORAMA conditions a pretrained segmenter on contextualized phrase representations to select masks, achieving the best overall grounding on PanoCaps; code, data, and models are released.

Dreaming the Sound of Contact: Leveraging Video and Audio Generation for Zero-Shot Force-Aware Manipulation and Data Generation

Pipeline pairs generated video with audio-derived force profiles to enable zero-shot, force-aware robot manipulation on Franka Panda for contact-rich tasks.

The paper leverages generated video and audio jointly: loudness of generated contact sounds shapes a bounded, time-varying desired-force profile from a natural-language task prompt. Trajectories execute on a Franka Panda robot with a closed-loop force regulator tracking the audio-shaped profile, succeeding where a kinematic-only baseline fails. The pipeline also serves as a data generation engine to train closed-loop manipulation policies.

arXiv cs.AI / cs.LG / cs.CL · 19h agoAI research

Exponential Hardness of Off-Policy Evaluation under History-Dependent Logging

Researchers prove off-policy evaluation under history-dependent logging requires exponentially many episodes, resolving a hardness question for model-based POMDP evaluation.

The paper constructs POMDPs with at most two latent states per stage, three actions, and a three-memory-state logger where evaluating a known deterministic target policy to accuracy 1/8 requires Θ((3/2)^H log(1/δ)) episodes for any horizon H≥3. Coverage and outcome-revealing conditions hold with constants independent of H, yet a reset erases the unknown transition that determines the target value. The authors characterize the resulting statistical experiment exactly, derive a matching optimal estimator, and validate predictions on a two-lane gridworld. This settles the history-dependent-logging, model-based case posed by Zhang and Jiang (arXiv:2503.01134).

arXiv cs.AI / cs.LG / cs.CL · 19h agoAI research

ScienceIDE: Turning World's Scientific Codebase into Agent Learnable Environments

ScienceIDE turns scientific code repositories into agent-trainable environments and trains PhAI-IDE models at 72B, 9B, and 4B scales.

ScienceIDE is infrastructure that transforms scientific repositories into executable environments supporting task generation, execution, and scientific verification, guided by expert-defined cases and acceptance criteria. Using verified interaction trajectories, the authors train PhAI-IDE-72B, PhAI-IDE-9B, and PhAI-IDE-4B. The model family improves held-out scientific-code repair and selected general benchmarks in code, reasoning, and knowledge, indicating positive transfer. Code is released on GitHub.

Cognitive Extensions for Dual-Process Language Agents: Memory and Self-Reflection in Interactive Environments

Researchers extend the SwiftSage dual-process agent with adaptive memory and self-reflection modules, improving scores in interactive environments.

The work adds an Adaptive Memory Module (AMM) for salience-gated episodic storage and trigger-driven retrieval, and a Self-Reflection Module (SRM) for bounded execution-time validation and corrective intervention, to the SwiftSage agent. Controlled ablations on ScienceWorld across four configurations show the full system achieves the best mean final score (64.62), success rate (43.17%), and successful-step efficiency (19.33 steps). SRM is the strongest standalone contributor, suggesting execution-time control is the dominant bottleneck while episodic memory helps once the runtime loop is stable.

arXiv cs.AI / cs.LG / cs.CL · 19h agoAI research

Affora: A Design System for Agent-Friendly Interfaces

Affora is a design system making interfaces legible to computer-use agents while preserving human workflows, with reusable components and executable checks.

Affora supports both human users and computer-use agents through a shared interface rather than a separate agent-only surface. Three controlled studies cover component implementations, visual variation, and interaction-design principles, producing guidance from individual components to complete sites with reusable implementations and executable checks. Evaluation on independently authored interfaces shows gains where agent-readability deficits exist, limited effects where they do not, and a workflow case gives preliminary evidence of reduced interaction cost.

arXiv cs.AI / cs.LG / cs.CL · 19h agoAI research

Playing log(N)-Questions over Wikipedia Abstracts: Communication Efficiency Between Paired Frontier Models

Six frontier models play a two-agent log(N)-Questions game; Claude Opus 5 lags with 28/68 wins while the top five are near-tied.

The study evaluates six frontier models on a two-agent game where a questioner must identify one of N Wikipedia lead paragraphs in exactly log2 N yes/no questions, run over 408 games at $363 total API cost. Claude Opus 5 wins 28 of 68 games versus 45-56 for GLM-5.3, GPT-5.6 Sol, Grok 4.6, Gemini 3.8 Flash, and Kimi K3. Pooled top-five win rates decline with set size (r=-0.973) and fit win = p^(log2 N) with per-round reliability p=0.928, and information per question correlates with win rate at r=+0.88.

How Model Growth, Recursion, and Boundary Operators Influence Scaling Exponents

Researchers show model growth via looped transformers improves scaling exponents; a 7.4B architecture matches GPT-3 13B with roughly 20x less compute.

The paper shows that architectural interventions, contrary to conventional wisdom, can modify pre-training scaling exponents and yield exponential performance gains with compute. Looped transformers with increasing loop counts provide a model growth mechanism; a 7.4B model-growth architecture matches GPT-3 13B on CORE with roughly 20x less compute, with efficiency gains that increase with scale. A boundary operator that normalizes and injects an earlier block also improves compute efficiency, and in data-constrained multi-epoch settings increasing loops with scale is compute-optimal.

arXiv cs.AI / cs.LG / cs.CL · 20h agoAI research

Evidence-Grounded Agentic Formulation Development in an Autonomous Laboratory

Andromeda 2, an agentic laboratory system, reaches a 50% high-performance hit rate for paclitaxel SEDDS formulations versus 17% for its predecessor and 2% for DoE.

Andromeda 2 is an agentic system that reasons over structured in-house experimental evidence and invokes computational and experimental tools to design and execute successive formulation batches for self-emulsifying drug delivery systems (SEDDS). For paclitaxel it achieved a 50% high-performance hit rate versus 17% for Andromeda 1 and 2% for a wet-lab DoE campaign, identifying 12 formulations meeting all four target product profile objectives versus 6 and 0. A selected full-TPP formulation reached approximately 19% w/w apparent paclitaxel loading, about 3.3-fold higher than a published paclitaxel S-SEDDS, and an ablation showed structured evidence access increased mean AUC by 34%.

arXiv cs.AI / cs.LG / cs.CL · 20h agoAI research

Prepared Or Unprepared? Evaluating Healthcare Workforce Readiness for Clinical Adoption of Artificial Intelligence in Nigeria

Survey of 761 Nigerian healthcare professionals finds high AI awareness (92.6%) but limited knowledge, preparedness, and major training and infrastructure barriers.

A cross-sectional study of 761 healthcare professionals across Nigeria, conducted from December 2025 to March 2026, found 92.6% awareness of AI in healthcare but 40.9% reporting low knowledge and only 63.0% feeling adequately prepared. Top barriers were lack of training (84.7%), poor infrastructure (71.1%), and high tool costs (61.0%). Willingness to adopt was strong, with 92.5% interested in training and 78.7% supporting AI in undergraduate curricula; preparedness differed significantly across geopolitical zones and professions.

arXiv cs.AI / cs.LG / cs.CL · 20h agoAI research

Reporting Practice Matters: The Impact of Reference Choice on Chest X-ray Report Evaluation

Study shows radiology reporting-style variations in reference reports can flip rankings of chest X-ray report generation models; releases MIMIC-CXR-Ext-ReRef dataset.

The paper quantifies how variations in radiologists' reporting practices distort evaluation of radiology report generation (RRG) models, introducing a radiologist-informed taxonomy and the ReRef method for rewriting reference reports while preserving clinical meaning. On MIMIC-CXR with RadCliQ-v1, condensing normal-findings discussion caused Libra to drop from first to second while CheXOne rose from third to first among nine models. The authors release MIMIC-CXR-Ext-ReRef, a radiologist-validated dataset of 120 original/alternative reference pairs, arguing metrics conflate clinical correctness with stylistic conformity.

arXiv cs.AI / cs.LG / cs.CL · 20h agoAI research

MUSE: Benchmarking Large Vision-Language Models on Multi-Modal Understanding in Situated Education

Introduces MUSE, a twelve-task benchmark evaluating vision-language models on artistic image understanding in situated educational, Southeast Asian contexts.

MUSE is a benchmark assessing large vision-language models on artistic image understanding across twelve tasks spanning visual perception, semantic and affective interpretation, cultural understanding, and compositional reasoning. It decouples image annotation from question generation for controllable difficulty and curates images centering Singaporean and Southeast Asian multicultural contexts alongside Western art. Evaluations of open-source and proprietary models found substantial disparities, especially in affective interpretation and compositional reasoning.

arXiv cs.AI / cs.LG / cs.CL · 20h agoAI research

A General Kernel Framework for Non-CND Distance Measures Using |D|-Dimensional Sparse Landmark Embeddings

Proposes the Sparse Landmark Embedding kernel, guaranteeing PSD kernels for arbitrary distances like geodesic and Wasserstein without CND requirements.

The paper introduces the Sparse Landmark Embedding (SLE) kernel, which embeds inputs via compactly supported bump functions at all |D| training points so any standard PSD kernel applies, removing the Hilbertian (CND) distance requirement that fails on manifolds and distribution spaces. Compact support controls sparsity, keeping kernel matrices well-conditioned despite high dimensionality. The authors prove PSD, sparsity, stability, and universal approximation guarantees, and show SLE matches or exceeds domain-specific baselines using geodesic and Wasserstein distances on accuracy and uncertainty quantification.

arXiv cs.AI / cs.LG / cs.CL · 20h agoAI research

Probabilistic Linear Explanations

Researchers introduce a unified probabilistic explainability framework using sparse anchored linear models that outperforms LIME and MAPLE on relevance error.

The paper proposes probabilistic explanations based on sparse, anchored linear models applicable to both binary classification and continuous regression. It proves that minimizing relevance error for neural-network models is NP-hard and relates it to a tractable fidelity-error surrogate. Solutions are computed via a mixed integer programming formulation with provably optimal empirical solutions and a polynomial-time iterative hard thresholding algorithm with approximation guarantees. Empirical evaluations show lower relevance error than LIME and MAPLE while satisfying anchoring and sparsity constraints by construction.

arXiv cs.AI / cs.LG / cs.CL · 20h agoAI research

Double descent is the principle of least action

A statistical mechanics analysis explains double descent: finite-time diffusion induces effective weight decay that regularizes models as parameters grow.

The paper models stochastic gradient-based training as a particle diffusing over the training-loss energy landscape at an induced temperature, sampling parameters via a Boltzmann distribution. Finite training time carries an effective weight decay, making every parameter a quadratic degree of freedom governed by the equipartition theorem. Adding parameters at fixed training loss lowers the temperature and the L2 norm of the stationary path, increasing effective regularization and explaining the double descent phenomenon.

arXiv cs.AI / cs.LG / cs.CL · 20h agoAI research

RLLBC-Lib: An Educational Code Library for Reinforcement Learning and Learning-Based Control

Researchers release RLLBC-Lib, an educational code library covering tabular and deep reinforcement learning with support for automated grading.

RLLBC-Lib is an educational code library aimed at lowering the entry barrier for students learning reinforcement learning in the context of learning-based control. It comprises a comprehensive library of tabular RL approaches, a deep RL library following the same design principles, and implementations contrasting RL with other learning-based control approaches. The library also serves as a basis for creating programming assignments with automated grading.

arXiv cs.AI / cs.LG / cs.CL · 20h agoAI research

Long-Lived Characters, Local Inference: Incremental Memory Maintenance for Game NPCs

Researchers present incremental KV-cache memory maintenance for long-lived game NPCs running locally on a quantized Qwen hybrid model.

The paper studies incremental memory maintenance for long-lived game NPCs deployed locally with a quantized Qwen hybrid recurrent-attention language model. The runtime removes superseded attention KV entries, computes replacement records at the true sequence tail, and preserves the continuing recurrent state and unchanged KV. Experiments across eight scripted maintenance rounds show true-tail updates preserve current-state and historical bindings, while slot-preserving alternatives repeat a double-subtraction error.

arXiv cs.AI / cs.LG / cs.CL · 20h agoAI research

Social Laws for Multi-agent Coordination in Stochastic Environments

Researchers extend social laws to stochastic, reward-based multi-agent environments, defining alpha-robustness and a verification method via Markov decision processes.

The paper extends the concept of social laws from deterministic, goal-based settings to stochastic, reward-based multi-agent environments. It introduces alpha-robustness, a measure of the guaranteed utility each agent retains while pursuing its optimal single-agent policy assuming all agents obey the social law. Robustness verification is reduced to solving a series of Markov decision processes, with empirical evaluations on toy environments.

arXiv cs.AI / cs.LG / cs.CL · 20h agoAI research

Comprehensive reconstruction of collider events with hypergraph representation learning and graph-conditioned diffusion

VyPER framework reconstructs collider events using hypergraph representation learning and graph-conditioned diffusion, outperforming existing reconstruction techniques across Standard Model processes.

Researchers present VyPER, a geometric learning framework that represents collider events as hypergraphs with a physics-inspired topology for particle event reconstruction. It combines supervised hyperedge classification for assigning measured jets and charged leptons to parent particles with a graph-conditioned diffusion model predicting unmeasured neutrino kinematics, optimized with a joint loss. Evaluated across several proton-proton collision processes, it demonstrates accurate reconstruction across Higgs, electroweak, and top-quark sectors.

arXiv cs.AI / cs.LG / cs.CL · 20h agoAI research

Higher-order pruning of experts in mixture-of-experts language models

New HOPE method uses second-order objectives to prune Mixture-of-Experts LLMs, outperforming REAP at 50% pruning on models up to 122B parameters.

Researchers derive HOPE (Higher-Order Pruning of Experts), a second-order pruning objective for Mixture-of-Experts language models that provably minimizes an upper bound on pruning error by modeling cooperative expert interactions. They show REAP, a state-of-the-art first-order method, is a special case of HOPE with interaction terms ignored. Across three frontier MoE models up to 122B parameters, two calibration sets, and benchmarks covering math, instruction following, coding, and agentic tasks, HOPE achieves the best average rank (1.58 of 5 at 50% pruning versus 2.42 for REAP), with gains up to +6.1% on agentic coding.

arXiv cs.AI / cs.LG / cs.CL · 20h agoAI research

Beyond Outcomes: Dual-View Relational Learning for Efficient Agent Benchmarking

DualViewEval compresses agent benchmarks by jointly modeling outcome and process signals, achieving 24x-40x compression with only 20 tasks on APEX-Agents and BFCL.

DualViewEval is an agent benchmark compression method that jointly exploits outcome and process relations from trajectories to learn exact-size minisets predicting full-benchmark scores. The authors analyze large-scale trajectories and identify six process signals systematically associated with final agent performance. Across five agent benchmarks and five baselines, it achieves the best results on all datasets: with only 20 tasks it reaches 24x-40x compression on APEX-Agents and BFCL, reduces MAE by 14.5%-28.2% over the strongest competitors, and improves Kendall's tau by up to 7.2% relative to EssenceBench on SWE-bench Verified.

arXiv cs.AI / cs.LG / cs.CL · 20h agoAI research1

How Much is a Human Right Worth? ECtHR-NPD: A Benchmark for Predicting Non-Pecuniary Damage Awards

ECtHR-NPD benchmark covers 14,575 European Court of Human Rights cases for predicting non-pecuniary damage awards; LLMs struggle with zero and high awards.

Researchers introduce ECtHR-NPD, described as the first benchmark for predicting non-pecuniary damage awards at the European Court of Human Rights from case information where no statutory formula exists. It contains 14,575 cases with case-level awards in nominal euros, chronological splits, and a protocol separating target construction from model input. Evaluations covering constant predictors, gradient-boosted trees, retrieval methods, fine-tuned encoder LMs, prompted decoder LMs, and knowledge-augmented agents show sophisticated LM approaches do not consistently outperform the strongest feature-based baseline. All model families struggle to identify zero awards and to calibrate high-award predictions, with further degradation on a Challenging test view.

arXiv cs.AI / cs.LG / cs.CL · 20h agoAI research

Structured Claim-Level Discourse Representations for Dense Health Narratives

Researchers propose a claim-level discourse framework for health videos, finding 13.22 atomic claims per minute and that LLMs struggle with pragmatic profiling.

The paper introduces a structured framework for claim-level discourse analysis in dense health narratives on social media videos, modeling tuples that link atomic claims with thematic aspects, stance, and multidimensional pragmatic attributes. Analysis found an average of 13.22 atomic claims per minute in health video discourse. A benchmark spanning four health domains with 1,191 manually annotated claims from 60 videos shows current LLMs perform strongly on thematic categorization and stance prediction but struggle with high-dimensional pragmatic profiling, suggesting future systems need task decomposition and specialized inference strategies.

arXiv cs.AI / cs.LG / cs.CL · 20h agoAI research

Fast Learning Rates for Physics-Informed Kernel Methods

Theoretical analysis proves finite-sample learning rates for physics-informed kernel estimators, showing differential observations can improve rates from n^-1/4 to n^-1/2.

The paper analyzes a physics-informed kernel estimator combining n value observations and m differential observations for a linear differential operator D, asking how much differential information improves prediction. The authors prove finite-sample bounds, supported by simulations, revealing a two-regime structure: when m is limited the rate depends jointly on n and m, and when m exceeds a problem-dependent threshold the rate saturates to the oracle rate. Examples in Sobolev spaces, including partial Laplacian constraints on the torus and gradient observations on bounded domains, illustrate improvements from the nonparametric n^-1/4 rate to the parametric n^-1/2 rate, plus physically consistent rates in a stronger norm.

arXiv cs.AI / cs.LG / cs.CL · 21h agoAI research