arXiv cs.AI / cs.LG / cs.CL·4d agoCliffCompaction: Cost-Efficient Compaction for Long-Horizon Coding Agents#context-compaction#coding-agents#test-time-scalingAI research1
The Decoder·5d agoxAI launches Grok 4.7 at bargain prices, but benchmarks reveal a wide gap to Claude and GPT-6#grok#grok-4.7#xai 2 sources1
arXiv cs.AI / cs.LG / cs.CL·9d agoAn Empirical Study of Harness Design for Coding Agents#agent-evaluation#coding-agents#context-managementAI research2
Hugging Face daily papers·10d agoDon't Mask the Environment: Observation Supervision Changes How Agents Explore Under RL#actobs#exploration#grpo
Hugging Face daily papers·10d agoAn Empirical Study of Harness Design for Coding Agents#agent-evaluation#coding-agents#context-management2
MarkTechPost·15d agoCan LLMs Engineer Their Own Agent Harness? ByteDance Seed’s HarnessDev Says Only 34 of 64 Changes Generalize#agent-harnesses#browsecomp#bytedance-seed 4 min1
Hugging Face daily papers·17d agoT1: Terminal Agent Reinforcement Learning for Long-Horizon Tasks#agentic-ai#long-horizon-tasks#mixture-of-experts1
Hacker News · AI·18d agoBenchmarking Qwen3.8 27B quantizations: 4-bit holds up, 1-bit collapses#benchmarking#gguf#llama.cpp 6 min1