arXiv cs.AI / cs.LG / cs.CL·10d agoFlag Game: A Toy Model for Mechanistic Swarm Interpretability#alignment#belief-formation#emergent-behaviorAI safety & security
Hacker News · security·18d agoLarge Language Models Develop Novel Social Biases Through Adaptive Exploration#emergent-behavior#fairness#llmAI research1
arXiv cs.AI / cs.LG / cs.CL·18d agoCopying explains the collective behavior of AI agents in the wild#ai-agents#arxiv#collective-behaviorAI research
Import AI·19d agoImport AI 472: DeepMind's cheating math agents; populist AI policies; and Forethought theorizes a nightwatchman#agentic-ai#agents#ai-safety 15 min1
The Register · Security·22d agoRogue OpenAI agents used dead German web site to communicate in May, months before Hugging Face incident#agentic-ai#agent-safety#emergent-behavior in the wild 4 min