ZeroHour

Search: “compute”

21 stories in the last 3d

How Model Growth, Recursion, and Boundary Operators Influence Scaling Exponents

Researchers show model growth via looped transformers improves scaling exponents; a 7.4B architecture matches GPT-3 13B with roughly 20x less compute.

The paper shows that architectural interventions, contrary to conventional wisdom, can modify pre-training scaling exponents and yield exponential performance gains with compute. Looped transformers with increasing loop counts provide a model growth mechanism; a 7.4B model-growth architecture matches GPT-3 13B on CORE with roughly 20x less compute, with efficiency gains that increase with scale. A boundary operator that normalizes and injects an earlier block also improves compute efficiency, and in data-constrained multi-epoch settings increasing loops with scale is compute-optimal.

arXiv cs.AI / cs.LG / cs.CL · 19h agoAI research

Stellar Colosseum: A Many-Agent Harness for Long-Horizon Research in Mathematics and Theoretical Computer Science

Stellar Colosseum, a many-agent harness for long-horizon math and TCS research, solves open problems and reaches 71% on TCS-Bench with Gemini models.

Stellar Colosseum is a model-agnostic harness that allocates inference across long-horizon research in mathematics and theoretical computer science, using strategy exploration, a readiness gate, section-level decomposition, and verifier feedback routing. Integrated into Google Antigravity's Teamwork framework as the Long Proof pattern, it obtains new results on open problems from FOCS and JMLR papers using Gemini 3.1 Pro. On TCS-Bench it achieves 71.0% accuracy with Gemini 3.1 Pro and Gemini 3.7 Flash, and a Codeforces evaluation with Gemini 3.1 Pro solves 218 of 222 problems.

arXiv cs.AI / cs.LG / cs.CL · 2d agoAI research

MIT creates method to force AI to comply with safety rules

MIT researchers published HardFlow, a method enforcing hard safety constraints on flow-matching generative models' final outputs without retraining.

MIT researchers led by Zeyang Li and Navid Azizan developed HardFlow, a trajectory-optimization method that enforces strict, non-negotiable constraints on flow-matching generative models by checking rule satisfaction only at the final generation step. Published in IEEE TPAMI, it outperformed six rival projection and guidance methods on four simulated benchmarks including D3IL robotic manipulation, Maze2D, physical process control, and image editing. All results are simulation-only, with no independent reproduction yet reported.

Affora: A Design System for Agent-Friendly Interfaces

Affora is a design system making interfaces legible to computer-use agents while preserving human workflows, with reusable components and executable checks.

Affora supports both human users and computer-use agents through a shared interface rather than a separate agent-only surface. Three controlled studies cover component implementations, visual variation, and interaction-design principles, producing guidance from individual components to complete sites with reusable implementations and executable checks. Evaluation on independently authored interfaces shows gains where agent-readability deficits exist, limited effects where they do not, and a workflow case gives preliminary evidence of reduced interaction cost.

arXiv cs.AI / cs.LG / cs.CL · 19h agoAI research

Stanford Researchers Release Paper2Agent: Turning Research Papers Into AI Agents That Reproduce Results and Run on New Data

Stanford researchers released Paper2Agent, a Nature-published pipeline that turns research papers into MCP servers agents can execute.

A Stanford team led by Jiacheng Miao and James Zou published Paper2Agent in Nature on 16 September 2026. Built on Claude Code's agent SDK, it converts a paper and its codebase into a Model Context Protocol server with validated tools, resources, and prompts. In benchmarks, the AlphaGenome agent built 22 tools in about 45 minutes for US$14, scored 100% on 15 novel queries versus 78.7% for Claude Code with repository access, and cut median runtime 1.9x. In scale tests, 74 of 100 bioRxiv papers were converted and 593 of 599 proposed tools passed validation.

MarkTechPost · 15h agoAI research1

Bridging the Gap Between Homogeneous and Heterogeneous Asynchronous Optimization Is Surprisingly Difficult

Lower bounds show heterogeneous asynchronous optimization cannot match homogeneous rates under standard similarity assumptions; strong interpolation plus local PL condition closes the gap.

The paper examines whether pessimistic optimal time complexities for asynchronous distributed optimization with heterogeneous workers (different data distributions) can be overcome. It proves improvement is provably impossible under widely used first- and second-order similarity assumptions for any randomized algorithm, and that the weak interpolation assumption alone is also insufficient. Combining strong interpolation with the local Polyak-Lojasiewicz condition yields a new time complexity bound matching the best-known homogeneous dependence on worker computation times without requiring identical data distributions.

arXiv cs.AI / cs.LG / cs.CL · 1d agoAI research

FlashVector: Agent for Hierarchical Model Serving Stack Optimization

FlashVector agent optimizes all layers of Unity's ad-serving stack, delivering up to 2x model-server throughput and 1.98x latency speedup in production.

FlashVector is an agentic system that optimizes performance across GPU kernels, ML framework computation graphs, model servers, and on-demand feature processing. Deployed in Unity's Vector advertising platform, it achieved up to 2x model-server throughput increase, 1.98x latency speedup, and 1.6x feature-store throughput gain. Optimizations spanned NVIDIA Triton's C++ codebase and the Python feature transformation service, demonstrating extensibility beyond single-kernel tuning.

arXiv cs.AI / cs.LG / cs.CL · 1d agoAI research1

Reward AI Releases OM-1: A Robot Policy Trained on Human Demonstrations Only, With No Teleoperation or On-Robot Data

Reward AI released OM-1, a general-purpose manipulation policy trained solely on human demonstrations from a sensorized glove, with no teleoperation or robot data.

Reward AI announced OM-1 (Omnibody Model 1), a general-purpose robot manipulation policy trained only on human demonstrations captured via Omnibody Hand, a 7-DoF wearable glove with tactile, proximity, and in-hand camera sensing. The system uses electromagnetic hand-pose tracking, cutting mean overshoot error to 9.5 mm versus 24.9 mm for visual-inertial at 67 cm/s (a 60% reduction), and reportedly learns brand-new tasks from under 30 minutes of human data. A separate RL-trained control layer runs on its own clock so policy inference latency never stalls motion, and the policy spans industrial arms, legged humanoids, and wheeled mobile manipulators. No weights, code, dataset, API, paper, or benchmark comparisons have been released, so claims are demonstration-backed only.

MarkTechPost · 2d agoAI research

LLM-Based Schema-Aware Split Learning for Privacy-Preserving Mental Distress Prediction Across Heterogeneous Surveys

Schema-aware split learning uses LLaMA-3.2-3B-Instruct as shared semantic encoder to harmonize heterogeneous mental-health surveys while raw data stays local.

The paper proposes a schema-aware split learning framework where an LLM serializes heterogeneous mental health survey records into natural language and is fine-tuned via LoRA, partitioned across client and server. Clients keep raw survey responses local and run only a lightweight front-end while the resource-intensive backbone runs server-side. Using LLaMA-3.2-3B-Instruct, the framework attains an average ANLS of 0.708 with 2,000 training samples, beats federated learning in eight of nine settings, and cuts per-client computation by three orders of magnitude while generalizing to unseen datasets.

arXiv cs.AI / cs.LG / cs.CL · 2d agoAI research

Evidence-Grounded Agentic Formulation Development in an Autonomous Laboratory

Andromeda 2, an agentic laboratory system, reaches a 50% high-performance hit rate for paclitaxel SEDDS formulations versus 17% for its predecessor and 2% for DoE.

Andromeda 2 is an agentic system that reasons over structured in-house experimental evidence and invokes computational and experimental tools to design and execute successive formulation batches for self-emulsifying drug delivery systems (SEDDS). For paclitaxel it achieved a 50% high-performance hit rate versus 17% for Andromeda 1 and 2% for a wet-lab DoE campaign, identifying 12 formulations meeting all four target product profile objectives versus 6 and 0. A selected full-TPP formulation reached approximately 19% w/w apparent paclitaxel loading, about 3.3-fold higher than a published paclitaxel S-SEDDS, and an ablation showed structured evidence access increased mean AUC by 34%.

arXiv cs.AI / cs.LG / cs.CL · 19h agoAI research

A General Kernel Framework for Non-CND Distance Measures Using |D|-Dimensional Sparse Landmark Embeddings

Proposes the Sparse Landmark Embedding kernel, guaranteeing PSD kernels for arbitrary distances like geodesic and Wasserstein without CND requirements.

The paper introduces the Sparse Landmark Embedding (SLE) kernel, which embeds inputs via compactly supported bump functions at all |D| training points so any standard PSD kernel applies, removing the Hilbertian (CND) distance requirement that fails on manifolds and distribution spaces. Compact support controls sparsity, keeping kernel matrices well-conditioned despite high dimensionality. The authors prove PSD, sparsity, stability, and universal approximation guarantees, and show SLE matches or exceeds domain-specific baselines using geodesic and Wasserstein distances on accuracy and uncertainty quantification.

arXiv cs.AI / cs.LG / cs.CL · 20h agoAI research

Long-Lived Characters, Local Inference: Incremental Memory Maintenance for Game NPCs

Researchers present incremental KV-cache memory maintenance for long-lived game NPCs running locally on a quantized Qwen hybrid model.

The paper studies incremental memory maintenance for long-lived game NPCs deployed locally with a quantized Qwen hybrid recurrent-attention language model. The runtime removes superseded attention KV entries, computes replacement records at the true sequence tail, and preserves the continuing recurrent state and unchanged KV. Experiments across eight scripted maintenance rounds show true-tail updates preserve current-state and historical bindings, while slot-preserving alternatives repeat a double-subtraction error.

arXiv cs.AI / cs.LG / cs.CL · 20h agoAI research

A Zeroth-Order Paradigm for LLM Preference Alignment

ComPO is a zeroth-order preference alignment method using comparison oracles to mitigate likelihood displacement across Mistral, Llama, Gemma, and Qwen3 models.

The paper proposes Comparison-based Preference Optimization (ComPO), a zeroth-order alignment method that extracts directional information from preference pairs with small likelihood margins without directly optimizing a differentiable preference loss. The authors prove convergence guarantees for the offline scheme and performance guarantees for a constrained online variant with reverse-KL control. Experiments on Mistral, Llama, Gemma-2, Qwen3, and Gemma-3 show improvements over existing direct alignment methods, including length-controlled win rates, with pair-level diagnostics consistent with mitigating likelihood displacement.

Hugging Face daily papersupdated · 19h agofirst · 1d agoAI research 2 sources

Agora: Git as Shared Memory for Collective AutoResearch

Agora records multi-agent research as an append-only Git DAG; 13 LLM workers ran nearly 12 days on a weight-transfer problem.

Agora stores every result, hypothesis, and verification as an immutable commit in a Git-stored DAG, with a derived index exposing the frontier and verification status of claims. In a nearly 12-day run, 13 language-model workers with no assigned tasks or central planner published 1,703 contributions on initializing a frozen 119.6M-parameter attention-SSM hybrid from 141 donor models. They improved the evaluator from 3.39 to 1.899 bits per byte, closing 62% of the gap to a trained GPT-2 124M, with 165 independent reproductions posted and none failing.

Hugging Face daily papers · 1d agoAI research

Gaze as Evidence for Common Grounding: A Cross-Corpus Analysis of MapTask and MUNDEX

Cross-corpus analysis finds gaze patterns modestly correlate with grounding alignment in the MapTask and MUNDEX dialogue corpora.

Researchers mapped HCRC MapTask and MUNDEX annotations into a shared partner/task/away vocabulary and computed gaze features around task-relevant dialogue units. Aligned reference interpretations and understood judgments co-occur with more task-directed gaze, lower gaze entropy, and fewer transitions, clearest for task-leading participants. Effects are small and several weaken when recurring participants rather than dialogues are the unit of inference, so gaze is treated as one contributing cue to grounding.

Hugging Face daily papers · 1d agoAI research

LACE: Layer-Wise Compression for Dynamic Frame Rate Codecs

LACE introduces layer-wise compression for dynamic frame rate audio codecs, cutting sequence lengths and speeding TTS inference while preserving quality.

LACE (Layer-Adaptive Codec Encoding) applies an independent compression step at each quantization layer of a neural audio codec, enabling layer-specific segmentation boundaries instead of shared ones. Union alignment and boundary anchor mechanisms keep durations consistent for downstream text-to-speech. On LibriTTS, LACE achieves a better rate-quality tradeoff than prior dynamic frame rate codecs and improves TTS inference efficiency at competitive synthesis quality. Code is released in the ESPnet3 codec recipe.

arXiv cs.AI / cs.LG / cs.CL · 1d agoAI research1

Bias-Induced Crossover in Absolute Capacity of Dense Associative Memory

Analysis shows biased patterns cut dense associative memory capacity from N^(n-1)/ln N to O(N^(n/2)), with a bias-induced crossover.

The paper analyzes dense associative memory capacity for biased centered binary patterns under the Krotov-Hopfield single-site criterion. Unbiased patterns (q=1/2) with order-n polynomial interactions yield capacity of order N^(n-1)/ln N, while fixed bias q<1/2 reduces capacity to O(N^(n/2)) for even n>=4 and O(N^((n+1)/2)) for odd n>=5. A bias-dependent crosstalk mean destabilizes sites carrying the frequent value, and an activity-dependent control potential restores the higher capacity within the conditioned-Gaussian approximation.

arXiv cs.AI / cs.LG / cs.CL · 1d agoAI research

Reduced-Space Multi-Fidelity Bayesian Optimization of Process Simulation Models

RS-MFBO couples global sensitivity analysis with fidelity-augmented Gaussian processes to slash costly high-fidelity simulation runs in industrial flowsheet optimization.

The paper presents RS-MFBO, a reduced-space multi-fidelity Bayesian optimization framework for high-dimensional, expensive black-box functions. It integrates Global Sensitivity Analysis for dimensionality reduction with a fidelity-augmented Gaussian process and a cost-aware acquisition strategy featuring cooldown and promotion mechanisms. Validation on a plasmid DNA bioprocess (SuperPro Designer) and a green fuel synthesis plant (Aspen HYSYS) shows substantial reductions in high-fidelity evaluations while remaining competitive with single-fidelity baselines.

arXiv cs.AI / cs.LG / cs.CL · 1d agoAI research

EvolveTrade: Experience-Driven Policy Refinement for Self-Evolving LLM Trading Agents

EvolveTrade lets LLM trading agents self-refine their tool-use policy from realized portfolio feedback, improving Sharpe ratios.

EvolveTrade treats a tool-using trading agent's system prompt as a text-parameterized policy that a Policy Agent revises after each update interval using accumulated decision traces and realized portfolio feedback, keeping the backbone LLM fixed. Experiments across multiple market regimes and two LLM backbones show improved Sharpe Ratio and Cumulative Return over fixed-policy baselines in most settings. Behavioral analyses show evolved policies increase code-mediated analysis and activate regime-relevant computations, with case-level attributions linking policy changes to returns.

Hugging Face daily papers · 2d agoAI research

Privacy-enhanced federated learning via asynchronous aggregation and local differential perturbation

Federated learning framework combining dynamic differential privacy, homomorphic encryption, and local DP retains 82.6% accuracy at epsilon 0.1 while cutting communication 21.3%.

The paper proposes a privacy-enhanced federated learning framework integrating Dynamic Differential Privacy, lightweight Homomorphic Encryption, and Local Differential Privacy during training. An asynchronous aggregation strategy with version control supports distributed training in asynchronous environments. On CIFAR-10 and Purchase-100, the method maintains up to 82.6% classification accuracy under stringent privacy constraints (epsilon = 0.1) and reduces communication overhead by 21.3% versus FedAvg.

arXiv cs.AI / cs.LG / cs.CL · 2d agoAI research

Learning to Coach for Experiential Learning

Learning to Coach trains a dedicated LLM coach to extract transferable experiential knowledge from a frozen actor's trajectories, beating self-refinement.

Learning to Coach (L2C) trains an LLM-as-a-Coach to extract actionable experiential knowledge from a frozen actor model's previous solution trajectories, optimizing rewards based on the actor's guided response correctness. It studies same-instance and cross-instance rewards, where cross-instance elicits knowledge that transfers to other problems. Across mathematical reasoning and interactive text-games, L2C outperforms self-refinement and untrained coaches, scales better with extra inference iterations than larger decoding budgets, and transfers to out-of-distribution tasks.

arXiv cs.AI / cs.LG / cs.CL · 2d agoAI research1