HypoEvolve: Genetic Algorithms Enable Multi-Agent LLMs to Discover Scientific Hypotheses
HypoEvolve couples a generational genetic algorithm with specialized LLM agents to generate drug-repurposing hypotheses, beating six baselines on DepMap selectivity (0.171 vs 0.115).
HypoEvolve coordinates specialized LLM agents through a generational genetic algorithm in which scientific judgments and new proposals reshape a hypothesis population. Evaluation centers on drug repurposing, linking mechanistic explanations to target-level biological claims assessed via external measures adapted from DepMap and Open Targets. Across 34 cancer types, HypoEvolve scores highest against six baselines on both measures, with DepMap selectivity of 0.171 versus 0.115 for the strongest baseline, and gains generalize to held-out cancer types.
HypoEvolve: Genetic Algorithms Enable Multi-Agent LLMs to Discover Scientific Hypotheses
HypoEvolve uses a generational genetic algorithm coordinating specialized LLM agents to generate scientific hypotheses, outperforming six baselines on cancer drug repurposing.
HypoEvolve is a framework that coordinates specialized LLM agents through a generational genetic algorithm to produce, revise, and retain scientific hypotheses with explicit collaboration roles. It evaluates drug repurposing hypotheses against external evidence from DepMap and Open Targets across 34 cancer types. It achieves the highest scores against six baselines, reaching DepMap selectivity of 0.171 versus 0.115 for the strongest baseline, and gains over single-pass generation generalize to held-out cancer types.
Algorithmic stability via ensembling
Theoretical work derives a general framework quantifying stability guarantees for averaging-based ensembles under arbitrary data perturbations via covariance operator norms.
The paper develops a framework for quantifying algorithmic stability of ensembling strategies defined via averaging, for varied types of data perturbation. The main result bounds the stability of the ensembled algorithm in terms of the norm of a covariance operator describing the ensembling process. The framework yields interpretable insights across practical perturbation examples and provides sharper guarantees than those derived from differential privacy considerations.
Guiding Worker Self-Selection in Crowdsourcing Contests: An LLM-Augmented Algorithmic Approach
Researchers introduce GRAF, a greedy framework for crowdsourcing contest self-selection, and LLMScore, an LLM-driven method that auto-designs its scoring algorithm.
The paper studies self-selection in Tullock contests (SSTC), where workers choose contests and then compete within them. GRAF is a greedy polynomial-time framework that orders workers by a score vector with zero worker regret and platform optimality guarantees in special cases. LLMScore is an LLM-driven evolutionary framework that produces human-readable, inspectable scoring code, jointly optimizing platform utility and worker satisfaction. Across 1,000 synthetic instances in four settings, GRAF with LLMScore achieves high-quality, often near-optimal outcomes with low worker regret, transferring from small training instances to larger, structurally different settings.
LLM-Driven Algorithm Design for Quantum Circuit Synthesis based on Binary Decision Diagrams
QuantumEvo uses an LLM to evolve BDD variable-ordering heuristics, achieving a 70.9% tie-or-win rate on quantum circuit cost versus baseline methods.
The QuantumEvo framework uses an LLM as a heuristic generator for quantum-cost-aware BDD variable ordering in reversible circuit synthesis, searching over heuristics initialized from multiple families and selecting them by downstream quantum circuit cost. The discovered heuristic HGA-QE modifies the sifting step inside a genetic algorithm and achieves a 70.9% tie-or-win rate against the per-function best baseline, with strict wins on 13.5% of functions. Advantages are clearer on benchmark suites not used for heuristic discovery.
5 useful things you'll learn in my new post-training textbook (shipping now!)
Nathan Lambert's new RLHF and post-training LLM textbook covers PPO, GRPO, GSPO, CISPO and related techniques, freely available online.
Nathan Lambert's book 'Reinforcement Learning from Human Feedback: Aligning and Post-training LLMs' is now shipping from Manning. It covers policy-gradient algorithms including PPO, GRPO, GSPO, CISPO, and RLOO, plus loss aggregation, truncated importance sampling, asynchronous RL systems, and post-training topics like rejection sampling, outcome reward models, and on-policy distillation. The book is freely available online with a 12-hour course, codebase, and exercises.
AI for Military Support
Study of 2,015 Israeli military personnel found algorithmic aversion toward AI targeting decision support, reduced when explainable AI features were added.
The paper 'Black Box Warfare' reconstructed a real-world military AI decision-support system used in targeting and tested a high-fidelity replica in two experiments with 2,015 Israeli military personnel. Contrary to automation-bias fears, participants showed strong algorithmic aversion, especially in high-collateral-damage scenarios. Integrating explainable AI features reduced aversion and promoted more thoughtful evaluation of algorithmic recommendations. The authors conclude that trust in military AI is dynamic and that human agency remains central in high-stakes decisions.
Import AI 470: No rights for machines; automating environment generation with SPADE; and building better GPU kernels with Hawkeye
METR analysis finds AI accelerating cyber vulnerability discovery, while SPADE self-play environment generation improves Qwen3 reasoning benchmark scores at 30B scale.
Import AI 470 discusses a METR research note reporting differential acceleration from AI: major acceleration in reported cyber vulnerabilities (cURL, OpenSSL, Firefox, Microsoft, NVD, OSV), minor acceleration in mathematics, and no measurable acceleration in AI-research optimization benchmarks. It also covers SPADE, a self-play framework from a multi-university team (University of Washington, Stanford, MIT, CMU, and others) that co-evolves executable training environments and agent capability using Environment Designer and Reasoning Agent roles with hint-based regret rewards. Trained on Qwen3-4B-Instruct-2507, Qwen3-8B, and Qwen3-30B-A3B-Instruct-2507 via GRPO (400 rollouts of 25 environments), SPADE lifted the 30B-A3B game-environment suite average to 58.3, +8.1 over base, and improved tool-use results across backbones. The issue also references Hawkeye for building better GPU kernels.
Objective vs. Search: Decomposing What Makes a Good Tokeniser
New tokeniser study shows search procedure, not optimisation objective, drives bits-per-byte performance across model sizes, vocabulary sizes, and multilingual settings.
The paper disentangles BPE and UnigramLM along two axes: optimisation objective (compression vs log-likelihood) and search procedure (bottom-up merging vs top-down pruning). Two new algorithms, BottomUpLL and TopDownComp, complete the 2x2 design space, and trained language models are evaluated on bits-per-byte and BLiMP across model sizes, vocabulary sizes, and English-only vs multilingual domains. Bottom-up tokenisers consistently achieve lower bits-per-byte in most settings, while BLiMP shows no consistent relationship with design choice.
Online Learning with LLM Experts from Limited Feedback
Paper proposes bandit algorithms for adaptively routing prompts to LLM experts, minimizing regret under limited feedback budgets.
The paper formulates adaptive prompt routing to K LLM experts as a contextual bandit problem with d prompt features over T rounds. Proposed algorithms strategically select actions and observe rewards, achieving O(dT/m) regret in the full-information setting and O(dTK/m) in the bandit setting, where m is the feedback budget. Experiments demonstrate efficient learning of high-quality routing strategies across diverse LLMs from limited feedback.
Why don't machine learning research agents overfit?
Amazon researchers explain why ML research agents avoid benchmark overfitting, attributing generalization to compressibility of successful strategies.
Amazon Science summarizes the paper "What fits (into few tokens) doesn't overfit: Compression and generalization in ML research agents," which investigates why benchmark hill-climbing loops, whether run by human communities or LLM research agents, do not produce rampant overfitting. The explanation formalizes Occam's razor via a counting argument: successful ML strategies are highly compressible, so short descriptions lack room to memorize benchmark data and must capture real structure. LLM-based agents, being resettable and controllable, allow this hypothesis to be tested empirically.
Unsolved Problem by Fields Medalist Breached by Two High School Students
Two high school students used Claude Opus 5 and GPT-5.6 Sol to help solve an open Lorentzian polynomials problem, posting a 75-page arXiv proof.
Aayush Bathija and Prince Rohatgi of Oak Park High School, mentored by UCLA postdoc Daniel Soskin, published the 75-page paper 'Bounded Ratios for Lorentzian Polynomials' (arXiv 2609.05341), solving an open problem in Fields Medalist June Huh's Lorentzian polynomial theory. The main structural theorem extends bounded coefficient-ratio characterization from quadratic to arbitrary-degree polynomials via discrete convexity conditions. The students used Claude Opus 5 and GPT-5.6 Sol for exploration and proof ideas but independently verified all arguments; the result follows an open letter from 25 Fields Medalists voicing concerns about AI's impact on mathematical rigor.
On the Regularization Landscape for the Linear Recommendation Models
Study shows leading linear recommendation models reduce to nuclear-norm or Frobenius-norm regularization, with two new closed-form low-rank solutions proposed.
The paper unifies top-performing linear recommendation algorithms under a single regularization framework, showing they effectively apply either nuclear-norm or Frobenius-norm regularizers. Nuclear-norm solutions have a rigid structure, are low-rank, and have closed form, while Frobenius-norm solutions are more expressive but full-rank or require hard-to-tune procedures such as ADMM. The authors derive two new low-rank, closed-form solutions that combine the advantages of both regularization families.
A positive resolution of the gap-entropy conjecture
New proof resolves the gap-entropy conjecture for Gaussian bandits, bounding optimal best-arm identification samples by H(log(1/delta)+Ent(I)) up to constants.
A paper proves the gap-entropy conjecture for fixed-confidence best-arm identification with independent unit-variance Gaussian arms, means in [0,1], and a unique optimal arm. It shows the optimal expected sample count, averaged over arm-label permutations, is within absolute constant factors of H(log(1/delta)+Ent(I)), where H sums squared gaps and Ent(I) is the instance's gap-entropy. It also gives an instance-independent algorithm bounded by a constant multiple of this quantity plus a g^-2 loglog(e^e/g) term for the smallest gap g.
Distributed and Private Textual Data Synthesis from Embeddings
Researchers propose a distributed differentially private text synthesis method combining DP summaries and secure protocols, removing the need for a trusted curator.
The paper presents a differential privacy and cryptography co-design for synthesizing textual training data without a trusted curator or tightly synchronized user participation. It releases a one-time DP summary in embedding space, identifying frequent semantic regions and their DP centroids to enable training-free offline text synthesis, with semantic support protection to avoid exposing rare user texts. A custom secure protocol enforces end-to-end DP guarantees over distributed user data. Across four benchmarks the approach achieves utility comparable to the state-of-the-art centralized DP synthesis method.
Learning Sparse Decision Trees via Transformer Variational Auto-Encoders
TREVIS uses a Tree Transformer VAE latent space to learn decision trees matching near-optimal predictive performance while improving structural sparsity.
TREVIS learns decision trees optimized for complex objectives by exploring the latent space of a Tree Transformer Variational Auto-Encoder (TTVAE). Mapping trees to continuous latent representations replaces the discrete search space with a continuous one, enabling gradient-based optimization through a differentiable surrogate model. Experiments show TREVIS matches the predictive performance of near-optimal algorithms while improving structural sparsity, targeting high-stakes contexts needing transparent decision logic.
Probabilistic Linear Explanations
Researchers introduce a unified probabilistic explainability framework using sparse anchored linear models that outperforms LIME and MAPLE on relevance error.
The paper proposes probabilistic explanations based on sparse, anchored linear models applicable to both binary classification and continuous regression. It proves that minimizing relevance error for neural-network models is NP-hard and relates it to a tractable fidelity-error surrogate. Solutions are computed via a mixed integer programming formulation with provably optimal empirical solutions and a polynomial-time iterative hard thresholding algorithm with approximation guarantees. Empirical evaluations show lower relevance error than LIME and MAPLE while satisfying anchoring and sparsity constraints by construction.
Bridging the Gap Between Homogeneous and Heterogeneous Asynchronous Optimization Is Surprisingly Difficult
Lower bounds show heterogeneous asynchronous optimization cannot match homogeneous rates under standard similarity assumptions; strong interpolation plus local PL condition closes the gap.
The paper examines whether pessimistic optimal time complexities for asynchronous distributed optimization with heterogeneous workers (different data distributions) can be overcome. It proves improvement is provably impossible under widely used first- and second-order similarity assumptions for any randomized algorithm, and that the weak interpolation assumption alone is also insufficient. Combining strong interpolation with the local Polyak-Lojasiewicz condition yields a new time complexity bound matching the best-known homogeneous dependence on worker computation times without requiring identical data distributions.
Coupled Calibration and Learning: Mitigating Teacher Bias in LLM Distillation without Target-Domain Reward Feedback
CCL couples teacher calibration with student updates via token-level branching, provably removing teacher bias in LLM distillation.
The paper proposes Coupled Calibration and Learning (CCL), an LLM distillation algorithm that alternates teacher calibration using source-question reward feedback with student training on target questions under covariate shift. Each iteration calibrates the teacher on source feedback, trains the student on target questions, and lets the updated student inform subsequent calibration. The authors prove the student's expected KL divergence to the oracle student converges to zero at a polynomial rate, and show regularized direct matching error can remain bounded away from zero.
Sakana AI Researchers Introduce PC-ALM, a Layer-Local Alternative to Backpropagation That Trains 1000-Layer Networks
Sakana AI's PC-ALM adds per-layer Lagrange multipliers to predictive coding, matching backprop on networks up to 1000 layers with layer-local updates.
Sakana AI researchers propose Augmented Lagrangian Predictive Coding (PC-ALM), a training method that keeps every update layer-local while recovering backprop-aligned credit signals. The team proves multipliers converge to exact backprop adjoints in linear networks and trains 1000-layer residual MLPs on MNIST within about 2 points of backprop accuracy. PC-ALM matched backprop across a width/depth grid from 8 to 128 on MNIST and Fashion-MNIST where standard predictive coding failed in deep, narrow networks, and improved over PC on ResNet-18 with CIFAR-10 and Tiny ImageNet. An MIT-licensed JAX reference implementation reproduces the results on CPU.
Safe Meta-Reinforcement Learning via Information Space Reachability
Safe meta-RL framework reasons about safety in information space, learning a safety value function used for safety filtering and constrained policy optimization.
The paper proposes safe meta-RL that reasons about safety in information space, capturing both physical state and the agent's belief over the underlying task. A safety value function measures the probability of avoiding unsafe regions indefinitely and satisfies a self-consistency condition and Bellman equation, making it learnable via meta-RL. The resulting algorithm uses the learned function for safety filtering and constrained policy optimization, with effectiveness demonstrated on meta-RL benchmarks.
Quenched Ensemble Sampling
Quenched Ensemble Sampling generalizes nested sampling's hard energy constraint to repulsive potentials, traversing first-order phase transitions where tempering fails.
Quenched Ensemble Sampling generalizes nested sampling's hard energy constraint into a family of repulsive potentials at the energy boundary, preserving monotone energy descent while making the constrained target amenable to scalable gradient-based kernels. On synthetic phase-transition models it estimates marginal likelihood and draws posterior samples across first-order transitions where popular alternatives such as tempering fail. Applications include marginal likelihood estimation for Bayesian neural network architecture comparison and partition function estimation in a high-dimensional continuous lattice field theory.