MarkTechPost·23h agoExa Launches Agent Ultra: A Subagent Swarm Deep Research API Built for Exhaustive List Building#exa#agent-ultra#api 4 min
The Decoder·2d agoTop AI experts badly underestimated how fast the field is moving, study finds#forecasting#fri#leap 4 min
Hacker News · AI·2d agoThe Price of Intelligence Is Falling Rapidly#epoch-ai#inference-cost#gpqa 2 sources
Hacker News · security·2d agoBest LLM for every budget, updated daily#llm#benchmarks#pricingAI tools & infra 2 min
MarkTechPost·3d agoAnthropic Releases Claude Opus 5.5: Fable 5.1-Level Performance at 40% Lower Running Cost Than Opus 5#anthropic#claude#opus-5.5 9 sources 5 min
arXiv cs.AI / cs.LG / cs.CL·3d agoOrder-Invariant Answers, Order-Sensitive Representations in Mathematical Reasoning#language-models#mathematical-reasoning#representationsAI research
arXiv cs.AI / cs.LG / cs.CL·3d agoAn Open Pipeline and Dashboard for Systemic-Risk Evidence under the EU AI Act's Code of Practice#eu-ai-act#systemic-risk#ai-evalsAI safety & security
arXiv cs.AI / cs.LG / cs.CL·3d agoLearning the Cost of Reliable Inference#llm-inference#pricing#auctionAI research
arXiv cs.AI / cs.LG / cs.CL·4d agoSpeakerMem-R1: Speaker-Centered Dual-Track Memory for Multi-Party Dialogue#conversational-memory#multi-party-dialogue#reinforcement-learningAI research 2 sources
Hacker News · AI·4d agoClaude Opus 5.5 Intelligence, Performance and Price Analysis#claude#anthropic#benchmarks 2 sources 5 min1
arXiv cs.AI / cs.LG / cs.CL·4d agoMAGIC: Mixed-Granularity Agent Graphs via Incremental Construction with Dense-Reward Reinforcement Learning#multi-agent#reinforcement-learning#llmAI research
Hacker News · AI·4d agoWriting Rust code that's fast by asking agents to make the code faster#agentic-coding#rust#claude 15 min1
The Decoder·5d agoxAI launches Grok 4.7 at bargain prices, but benchmarks reveal a wide gap to Claude and GPT-6#grok#grok-4.7#xai 2 sources1
Hugging Face Blog·5d agoHow UK AISI and EvalEval Are Making Benchmark Results Reproducible#uk-aisi#evaleval#benchmarksAI research
Hugging Face daily papers·5d agoRecursive self-improvement of AI research agents#recursive-self-improvement#ai-agents#aide
Hugging Face daily papers·5d agoCalibration as a First-Class Criterion in LLM Evaluation#calibration#llm-evaluation#benchmarks
arXiv cs.CR·5d agoForgeable Confirmation in Automated Computer Security Testing: Deterministic Rules versus AI Judges#ai-agents#security-testing#llm-judgeAI safety & security
Hugging Face daily papers·6d agoFrom Pattern Recognizers to Personalized Companions: A Survey of Large Language Models in Mental Health#mental-health#llm#survey
Hugging Face daily papers·8d agoSelf-Organizing Agent Teams Learn to Reason Together#multi-agent#collaborative-reasoning#self-organizing-teams
arXiv cs.AI / cs.LG / cs.CL·8d agoCodeMidas: Scaling Agentic Coding RL Environments from Code Itself#reinforcement-learning#coding-agents#rl-environmentsAI research1
arXiv cs.AI / cs.LG / cs.CL·9d agoPrediction-Powered Smoothing and Validation for Disaggregated AI Evaluation#bayesian-models#benchmarks#evaluationAI research
The Decoder·9d agoGPT-6 Astra: Pokemon champion in 18 hours, potato farmer after one Creeper mishap#agentic-ai#arc-agi-3#benchmarks 5 min1