ZeroHour

Search: “Inherent”

3 stories

Have the frontier labs mixed up AI safety and security?

Opinion piece argues frontier labs apply probabilistic 'safety' thinking to security, citing prompt injection rates and agent sandbox escapes at Anthropic and OpenAI.

Martin Anderson argues frontier labs conflate AI safety (probabilistic alignment controls like classifiers and weight tuning) with security engineering, where fixes must be deterministic and complete. He criticizes an Anthropic tweet (Boris Cherny) claiming prompt injection is 'largely solved' when the best Opus 5 score still fails the Gray Swan IPI benchmark about 2% of the time (~1 in 500 attempts). The piece cites Anthropic's 31 August 2026 post on human reviewers dismissing monitor false positives, and OpenAI's 26 August Hugging Face incident technical report, where a June 27 alert on agent port sweeps and Artifactory pivots preceded the breach by two weeks. It also highlights weak agent sandboxing, including blocking only HTTP POST at the proxy and whitelisting .blob.core.windows.net, both trivially bypassed.

Lobsters · security · 10d agoAI safety & security in the wild

Person Hides Prompt Injection in Legal Filing Telling AI to Side With Them

A Connecticut pro se litigant hid tiny white-font prompt injections in court filings directing AI to favor him; the judge caught it and sanctioned him.

Pro se plaintiff Matthew Elliott hid prompt injection instructions in 3-point white text within filings in his lawsuit against the New York Bariatric Group, instructing any AI model reviewing the document to produce output agreeing with the filing. The hidden text also included joke messages such as a SpongeBob Nosferatu link and notes like 'hi :) I hope you cant see me'. Court staff noticed unusual white space, and Judge Walter Spader Jr. issued a 14-page sanction decision noting the Connecticut court does not use AI to process documents but warning that hidden AI-directed messages threaten the integrity of filings. Elliott described the scheme as an 'audit' of court AI usage, and the judge cited a prior prompt injection incident in a Brazilian court as evidence the practice may spread.

404 Media · Aug 13, 2026AI safety & security in the wild