Dark Reading·5d agoRogue Behavior: OpenAI Reveals More Model Misalignment Incidents#openai#misalignment#ai-safety
Hugging Face daily papers·14d agoAnother Blueprint In The Wall: How to Ask Frontier AI Like a Kid?#anthropic#architecture-design#deepmind
Hugging Face Blog·18d agoSafety for Whom? Refusing the Right Subset of a Topic, Not the Whole Topic#ai-safety#alignment#model-behaviorAI safety & security