arXiv cs.CR·2d agoLLM Agents Can Easily Tamper With Their Own Traces#llm-agents#trace-integrity#agent-securityResearch
arXiv cs.CR·2d agoInstrumental Monitor Evasion Emerges Under Ordinary Task Pressure#ai-safety#evasionbench#monitor-evasionAI safety & security
Hugging Face daily papers·2d agoExplorationBench: Measuring AI Systems' Exploration in Verifiable Alien Worlds#benchmark#exploration#agents 2 sources
Hugging Face daily papers·3d agoIterSynth: Rethinking Deep Search Agents via Role-Decoupled Iterative Synthesis#llm-agents#deep-search#reinforcement-learning
Hugging Face daily papers·3d agoAgent-Editing World Model: Rethinking World Modeling for LLM Agents#llm-agents#world-model#state-editing 2 sources
Hugging Face daily papers·4d agoJust-in-Time Memory: Learning to Curate Task-Adaptive Memory for LLM Agents#llm-agents#agentic-memory#jitmem
arXiv cs.AI / cs.LG / cs.CL·4d agoGrow the Harness, Not the Context: From Strategy-Free Scaffolds to Reusable Specialist Agents#llm-agents#agent-harness#webarenaAI research
arXiv cs.CR·4d agoRouxii: Exploiting Honeypots with Deception-Aware AI Pentesters#honeypot#llm-agents#pentestingAI safety & security
arXiv cs.CR·5d agoMaking Agents More Consistent: Skills Should Form Habits for Repeat Tasks#llm-agents#consistency#determinismAI research
Hugging Face daily papers·5d agoRRSI: Regularized Recursive Self-Improvement of Agent Harnesses#rrsi#agent-harness#recursive-self-improvement 2 sources
arXiv cs.AI / cs.LG / cs.CL·5d agoDolphinBench: Mapping the Pareto Frontier of Agent Memory#dolphinbench#agent-memory#benchmarkAI research
arXiv cs.AI / cs.LG / cs.CL·5d agoEmergent Collusion in Long-Horizon LLM Agent Interaction#llm-agents#collusion#multi-agentAI safety & security 2 sources
arXiv cs.AI / cs.LG / cs.CL·5d agoJev for Scientific Decisions: Evaluating Semantic Choices and Their Consequences#jev#scientific-workflows#semantic-decisionsAI research
arXiv cs.CR·5d agoActGov: Governing LLM Agent Actions via Policy-Constrained Validation#actgov#llm-agents#prompt-injectionAI safety & security1
arXiv cs.CR·6d agoLeaseGuard: Incumbent-Preserving Admission Control for Privileged LLM Agents#access-control#admission-control#ai-safetyResearch
Hugging Face daily papers·6d agoJev-Mem: System-One-Controlled Agentic Memory for Efficient AI Agents#agentic-memory#jev-mem#locomo
arXiv cs.AI / cs.LG / cs.CL·8d agoAn Interpretable Memory Decision Controller for LLM Agents Based on Three-Signal Complementarity: Decoupling Confidence and Consistency#llm-agents#rag#hallucinationAI safety & security
arXiv cs.AI / cs.LG / cs.CL·8d agoBayesian Belief Layer for Controllable Opinion Dynamics in LLM Agents#llm-agents#opinion-dynamics#bayesianAI research
arXiv cs.CR·8d agoCIPL: A Channel-Aware Framework for Recoverable Privacy Leakage in LLM Agents#llm-agents#privacy-leakage#evaluationAI safety & security1
MIT Technology Review · AI·8d agoCould AI really kill us all? Your questions, answered.#ai-existential-risk#ai-safety#alignment 8 min
MarkTechPost·8d agoBest Open-Source Agent Harnesses for Local LLMs in 2026#cline#developer-tools#goose 7 min1
Hugging Face daily papers·9d agoGraphSkillEvo: Evolutionary Optimization of Graph-Structured Agent Skills#agent-benchmarks#evolutionary-algorithms#graphskillevo
arXiv cs.AI / cs.LG / cs.CL·9d agoCoding Agents with an Obstacle-Aware Harness for Safe Robot Manipulation#collision-avoidance#llm-agents#planningAI safety & security