arXiv cs.AI / cs.LG / cs.CL·4d agoBeyond Repeated Sampling: Learning Search Policies for LLM Reasoning#llm#reasoning#test-time-computeAI research
Hugging Face daily papers·13d agoWhen Agents Slow Down: Understanding LLM Agents' Test-Time Strategies via Elo-per-token Analysis#benchmarks#elo#evaluation5
Latent Space·18d ago[AINews] OpenAI reports Navier-Stokes singularity find in 88 hours using Astra-next, roughly 10,000 agents and 130B tokens (>$40M), a contender for second ever Millennium Prize awarded#agentic-ai#millennium-prize#multi-agent-rl 15 min1
Hugging Face daily papers·18d agoStudying Without a Syllabus: Task-Agnostic Environment Preprocessing#benchmarks#environment-preprocessing#llm-agents
Hugging Face daily papers·18d agoAn Open Recipe for IMO Gold: Training Nemotron for Olympiad Mathematics#mathematical-reasoning#nemotron#olympiad-mathematics
arXiv cs.AI / cs.LG / cs.CL·19d agoA*-Thought-V2: Efficient Latent Reasoning via Geometric Dynamics of LLM#chain-of-thought#efficiency#hidden-state-dynamicsAI research
arXiv cs.CR·23d agoRethinking Indirect Prompt Injection as a Test-Time Search Problem#adversarial-attacks#agentic-ai#ai-evalsAI safety & security
Hugging Face daily papers·24d agoWhat Else Needs Fixing? Exploring Cost-Effective Test-Time Compute for Revision Propagation in Artifacts Generated Through Conversation#artifact-revision#benchmark#code-generation