ZeroHour

Search: “settlement”

40 stories

Smart search ranks by meaning as well as keywords (one row per story, last 45 days).

Anthropic's $1.5 billion book settlement descends into chaos as authors and publishers fight over who gets paid

Authors and publishers are filing competing claims over payouts from Anthropic's $1.5 billion copyright settlement covering 482,000 pirated books.

Anthropic agreed to pay $3,000 per illegally downloaded book, affecting more than 482,000 titles, in the largest copyright settlement in US history. As payouts begin, authors, publishers, and literary agencies are filing competing claims, complicated by poor rights-reversion records; textbook authors are reportedly owed as little as 10-15 percent under their contracts. A court had ruled Anthropic's use of illegally obtained books unlawful while deeming training on legally purchased books fair use.

The Decoder · 5d agoAI policy1

Authors push back as publishers and agents make claims on Anthropic settlement

Authors report publishers and agents wrongly claiming shares of Anthropic's $1.5 billion copyright settlement, which pays $3,000 per pirated work across roughly 500,000 titles.

Anthropic's $1.5 billion copyright settlement, given final approval in July, pays $3,000 per pirated work for nearly 500,000 titles, split 50-50 with publishers for in-print books. Authors including April Henry report publishers claiming payments for works whose rights reverted years ago, and some agents claiming percentages despite not being rightsholders. Authors Guild CEO Mary Rasenberger attributes the disputes to poor recordkeeping rather than deliberate overreach. Full author claims require rights reversion before the settlement's August 10, 2022 download date.

TechCrunch · AI · 9d agoAI industry

Meta to Pay Up to $18B Over Teen Social Media Use

Meta agreed to pay up to $18 billion and cap teen Facebook and Instagram use at two hours daily, settling child-safety lawsuits from most US states.

Meta settled claims that it deliberately designed Facebook and Instagram to addict children, ending a federal trial as Instagram head Adam Mosseri began testifying and Mark Zuckerberg was expected to testify. Of roughly $16.7 billion for 47 states and territories, $12.7 billion is guaranteed and $5 billion is contingent on Snapchat, TikTok and YouTube adopting similar teen protections. The deal also resolved state privacy claims tied to Cambridge Analytica for an additional $459 million, while New Mexico (a $567 million ruling) and Florida opted out of the settlement.

Security Affairs · 20d agoPolicy & legal

Grindr to Pay £26 Million to Settle U.K. Claims Over HIV Status Data Sharing

Grindr will pay £26 million ($35.1M) to settle U.K. claims from 10,000+ users over pre-2020 sharing of HIV status and other sensitive data.

Grindr agreed to pay £26 million ($35.1 million) to settle a U.K. lawsuit brought on behalf of more than 10,000 claimants over sharing users' HIV status, last tested date, and other personal data with third parties for advertising before 2020, when the app was owned by China's Kunlun. The settlement, disclosed in a September 2 SEC filing, includes no findings or admission of liability, with £13 million due by December 31, 2026 and the rest by March 31, 2027. Norway's data protection authority previously fined Grindr £8.6 million (reduced to £5.5 million) under GDPR, a decision upheld on appeal last October.

The Hacker News · 8d agoPolicy & legal

Meta pledges to overhaul kids’ safety protections, pay $17 billion to settle social media case

Meta settles states' kids' online safety lawsuit for $17 billion, agreeing to landmark usage limits, age restrictions, and independent auditing.

Meta agreed to pay $17 billion to settle a civil suit from nearly every US state and territory alleging it hid research showing Facebook and Instagram are addictive to minors and violated COPPA by collecting data on children under 13. The settlement imposes reforms including two-hour daily limits for users under 18, a midnight-to-6am usage block, non-personalized feed options, and an independent auditor. Meta also settled separately with Texas for about $1 billion.

The Record · 20d agoPolicy & legal

TikTok Settles U.S. Child Privacy Case for $400 Million

TikTok will pay $400 million to settle U.S. DOJ/FTC claims that it violated COPPA by collecting data from children under 13.

The U.S. Department of Justice announced a $400 million settlement with TikTok and ByteDance resolving a 2024 lawsuit over violations of the Children's Online Privacy Protection Act (COPPA). TikTok will pay $300 million immediately and $100 million upon entry of an order vacating a prior consent decree against its predecessor Musical.ly; it is one of the largest recoveries ever obtained in a COPPA case. The DOJ and FTC, filing in California, alleged TikTok knowingly allowed children under 13 to create accounts and illegally collected data via Kids Mode. TikTok was previously fined €345 million by Ireland's Data Protection Commission in 2023 for GDPR breaches involving children's data.

Security Affairs · 23d agoPolicy & legal

Grindr settles privacy lawsuit tied to disclosure of users’ HIV statuses for $35 million

Grindr will pay $35.2 million to settle a UK privacy lawsuit alleging it shared users' HIV status with advertisers.

Grindr agreed to pay 26 million pounds ($35.2 million) in two lump sums to resolve a UK class action alleging it provided advertisers with sensitive user data including HIV status. The suit was filed in April 2024 over conduct before early 2020, when the app was under Chinese ownership, and involved roughly 12,000 class members. Per an SEC filing, the settlement includes no findings or admission of liability, and the company says data was shared in encrypted form with two service providers.

The Record · 6d agoPolicy & legal

Grindr Settles UK Data Privacy Claims for £26m

Grindr will pay £26m ($35.2m) to settle UK group claims alleging unlawful sharing of sensitive data, including HIV status, before 2020, without admitting liability.

The settlement, reached on September 2 and disclosed to the US SEC, covers roughly 12,000 claimants represented by Austen Hays over the free app's 2016–2020 data practices when Grindr was owned by Chinese conglomerate Kunlun. Grindr will pay £13m by December 31, 2026 and £13m by March 31, 2027, and continues to dispute the allegations; the agreement contains no admission of liability. The claims concerned sharing HIV status, PrEP use, ethnicity, and sexual orientation data with analytics providers Apptimize and Localytics without adequate consent. Norway's data protection authority fined Grindr €6.5m in 2021, and the UK ICO reprimanded the company in July 2022.

Infosecurity Magazine · 8d agoPolicy & legal

Grindr settles HIV status data-sharing lawsuit for $35 million

Grindr agreed to pay about $35 million to settle a UK privacy suit alleging it shared users' HIV status and sensitive data with advertisers without consent.

The claim, brought by London firm Austen Hays on behalf of roughly 12,000 UK users, alleges Grindr breached privacy and data-protection laws during a period ending in early 2020, when it was owned by Beijing Kunlun Tech. Shared data may have included ethnicity, HIV status, last HIV test date, and PrEP use. Per an SEC filing, Grindr will make two payments of £13 million (totaling about $35 million), one by December 31, 2026 and one by March 31, 2027, without admitting liability. The settlement follows a Norwegian Data Protection Authority enforcement finding over ad sharing without a valid legal basis.

Malwarebytes Labs · 8d agoPolicy & legal

TikTok Agrees to $400 Million Settlement in U.S. Child Privacy Lawsuit

TikTok agrees to pay $400 million to settle a DOJ child-privacy lawsuit alleging COPPA violations including collecting data from children under 13.

The DOJ announced that ByteDance-owned TikTok will pay $400 million to settle a 2024 lawsuit alleging massive-scale invasions of children's privacy. TikTok pays $300 million immediately and $100 million upon vacating the prior Musical.ly consent decree. The complaint alleged children under 13 could create accounts, data was collected in Kids Mode, and parental deletion requests went unheeded. DOJ called it one of the largest recoveries ever under COPPA, following TikTok's 2023 €345 million GDPR fine.

The Hacker News · 25d agoPolicy & legal

FST Pay: Deterministic Safety-Gated Architecture for Youth Digital Payments

FST Pay proposes a deterministic safety-gated architecture for teen digital payments, pairing invariant authorization checks with decoupled post-settlement AI explanations.

Researchers propose FST Pay, a formal architecture for adolescent digital payments on rails like UPI that applies six deterministic invariant checks (spending limits, guardian co-sign policies, amount thresholds, merchant category codes, temporal intervals, hardware integrity) to classify transactions as ALLOW, REVIEW, or BLOCK. High-risk transactions trigger an asynchronous guardian co-sign workflow. Generative AI is restricted to post-settlement natural-language insights and holds no mutation privileges over the ledger, avoiding non-determinism on the real-time authorization path.

arXiv cs.CR · 6d agoResearch

Pornhub's Parent Company to Pay $120 Million to Settle Child Sexual Abuse Lawsuits

Pornhub parent Aylo will pay $120 million settling child sexual abuse class actions, without admitting liability, and adopt stricter content moderation commitments.

Aylo, Pornhub's parent company (formerly Mindgeek, acquired by Ethical Capital Partners in 2023), will pay $120 million to settle two 2021 class actions in California and Alabama alleging its platforms hosted child sexual abuse material in violation of federal trafficking and child imagery laws. The settlement fund begins with $25 million in 2026 followed by six annual installments. The class covers anyone under 18 appearing in content on Mindgeek-operated sites between February 12, 2011 and December 6, 2024. The deal, subject to court approval, adds commitments to age-verify models, conduct human and automated content review, and report suspected abuse material to authorities.

404 Media · 29d agoPolicy & legal

AgentGrad: Intervention-guided Prompt Optimization for Multi Agent Systems

AgentGrad introduces intervention-guided prompt optimization for LLM multi-agent systems, achieving state-of-the-art results with 2.5x faster optimization.

AgentGrad is a prompt optimization framework for LLM-based multi-agent systems that addresses limitations in textual gradient extraction and aggregation. It uses sequential intervention to identify the agent whose prompt modification resolves a given failure, then applies agent-level supervision and semantic gradient clustering to build generalized gradients. Experiments report state-of-the-art performance across five MAS benchmarks and a 2.5x average reduction in wall-clock optimization time versus the next-fastest baseline.

Hugging Face daily papers · 8d agoAI research

Meta debuts its Muse AI agent. Will consumers trust it?

Meta launched Muse, a consumer AI agent powered by Muse Spark that connects to users' apps to execute tasks like emailing, booking travel, and payments.

Meta introduced Muse, a personal AI agent for US users that connects to email, calendars, payments, shopping, and other services to execute tasks such as booking travel, lowering bills, and completing purchases via Link by Stripe. The agent runs in a dedicated Muse Secure VM with a separate Sentinel agent kept apart at the system level, and Meta claims it cannot see passwords or payment data and does not share conversations with ad systems. Muse is free to start, with Power ($20/month) and Maximum ($100/month) subscription tiers, and is available on the web, iOS, Android, and WhatsApp, with Meta AI glasses support planned. The launch follows Meta's $18 billion multistate consumer-harms settlement and comes as rivals like Gemini Spark and Claude Cowork push agentic AI.

TechCrunch · AI · 7d agoAI industry 3 sources1

terms.txt: A Consent and Compensation Protocol for Agentic Web Access

terms.txt specifies a robots.txt-style protocol for per-path, per-purpose AI crawler consent and compensation, with enforcement adding 0.20-0.65 ms per request.

The paper documents that automated clients now make up most web requests, that training dominates Cloudflare-classified crawling, and that the largest AI platforms fetch thousands of pages per returned visitor while robots.txt cannot express identity, purpose, terms, or price. It specifies terms.txt plus an origin-enforced exchange using Web Bot Auth signatures, signed intent, delegation tokens, HTTP 402 negotiation, and signed receipts. A dependency-free implementation adds 0.20 to 0.65 ms per request on one vCPU.

arXiv cs.CR · 6d agoResearch

IntentFuzz: A Protocol-Aware Fuzzer for Automated Invariant Violation Detection in Intent-Based Cross-Chain Bridges

IntentFuzz protocol-aware fuzzer recovers bridge structure from unannotated Solidity and confirmed 22 invariant violations across 24 real-world deployments.

IntentFuzz formalizes a taxonomy separating invariant violations from settlement exposures in intent-based cross-chain bridges, then recovers a bridge's intent structure and deposit/fill function roles from unannotated Solidity source. It classified deposit and fill functions with 100% recall and 82% combined precision, and achieved 100% recall and precision on 23 planted-bug mutants. Across 24 real-world deployments it confirmed 17 genuine invariant violations with heuristic-only input generation, rising to 22 with its LLM-assisted tier, spanning eight vulnerable GitHub repositories with findings reproducible against public deployed bytecode.

arXiv cs.CR · 4d agoResearch1

UK's Online Safety Act has made 'absolutely no difference,' kids say

UK Children's Commissioner tells Lords committee the Online Safety Act has 'made absolutely no difference' and criticizes Ofcom over risk assessment transparency.

England's Children's Commissioner Dame Rachel de Souza testified that more than a year after key Online Safety Act child-protection duties took effect, children report no meaningful change in accessing harmful content. She criticized Ofcom for refusing to share companies' safety risk assessments under section 393(1) of the Communications Act 2003, and planned to use statutory powers to compel disclosure. She argued the OSA has not kept pace with AI-driven harms (citing the 'Grok nudifying' controversy) and urged Ofcom to 'use its teeth,' contrasting the UK's approach with Meta's proposed $18 billion US child-safety settlement.

The Register · Security · 13d agoPolicy & legal1

Claude's new system prompt really doesn't want to reproduce song lyrics

Anthropic published updated Claude consumer system prompts, including changes steering the model away from reproducing song lyrics, likely over copyright concerns.

Anthropic publishes system prompts for Claude.ai and Claude mobile apps, including historic revisions, and has reorganized them into an index with per-model pages such as the Haiku 4.5 page showing the original October 15, 2025 prompt and an updated January 18, 2026 version. The latest consumer prompt strongly discourages reproducing song lyrics, a behavioral constraint likely tied to copyright considerations. Prompts for Claude Cowork and Claude Code are not included in the published set.

Simon Willison · 14d agoAI safety & security

LongAgent: History-Guided Agentic Search for Longitudinal Outcome Prediction

LongAgent autonomously searches variable sets and temporal windows to predict longitudinal medical outcomes, beating the strongest non-agent baseline on synthetic data.

The paper proposes LongAgent, an agent-based method that searches over combinations of variable sets, temporal windows and aggregation functions for outcome prediction on heterogeneous medical longitudinal data. It uses a history memory of previous searches and numerical evidence to guide exploration. On synthetic data it achieves mean RMSE 1.7376, improving over the best non-agent baseline by 0.0151 (95% CI [0.0045, 0.0260]; p=0.0273), and performs comparably to the best baseline on a real clinical dataset.

arXiv cs.AI / cs.LG / cs.CL · 1d agoAI research

Does Your Agent's Memory Survive a Model Upgrade? A Controlled Study of Memory Portability

A controlled study finds agent memory portability varies sharply: fixed-schema knowledge graphs survive model swaps while compressed notes degrade.

The study compares preserving an agent's history as raw long context, RAG chunks, compressed natural-language notes, or fixed-schema knowledge graphs across model upgrades, using 48 synthetic histories and two open-weight sub-10B-parameter models. Fixed-schema KG accuracy changed by only +0.0004 ± 0.0020 after a writer swap, while compressed NOTES shifted asymmetrically by +9.91 or -13.28 percentage points depending on migration direction. Mixed 50/50 embedding migrations captured only 4.96 of an 11.90-point RAG re-embedding gain; 80% of the NOTES deficit came from information lost at construction, and 81% of the RAG deficit from retrieval failures. Store-only repair of NOTES failed to reach 90% recovery in all 48 cases, while retaining raw histories enabled recovery in 34 of 48 for one direction.

arXiv cs.AI / cs.LG / cs.CL · 11d agoAI research1

Meta’s AI agent Muse is now the No. 2 app in the US

Meta's agentic AI app Muse ranked No. 2 on the US iOS charts with 83,000+ downloads, trailing Threads and ChatGPT launch pace.

Sensor Tower data shows Muse was downloaded over 83,000 times on iOS in the US, climbing from 4th to 2nd on the App Store, but below Threads' 4.3 million launch-day downloads and ChatGPT's 500,000 first-week installs. The Android version ranks only No. 338 in the Productivity category on Google Play. Muse launched days after Meta's $18 billion multistate settlement over social media consumer harms. Rival agent Instinct, valued at $2.5 billion with $350 million available, recently added Stripe and 1Password integrations.

TechCrunch · AI · 5d agoAI industry 4 sources

LexFlip: A Dissociation Diagnostic for Legal Meaning Preservation Metrics

LexFlip releases 373 minimal perturbations of Quebec statutory French that reverse legal force while preserving tokens, exposing weaknesses in embedding-based meaning preservation metrics.

LexFlip provides 373 minimal perturbations of Quebec statutory French that reverse legal force while preserving 0.93 of tokens, creating dissociation items that break monotone token-overlap metric validation. The seven embedding and BERTScore metrics tested register only 0.022-0.039 of their identical-to-unrelated range on these edits, versus 0.670 for bidirectional NLI. Against FrJudge, with a measured human ceiling of r=0.597, a bare length feature outscores every semantic metric tested.

arXiv cs.AI / cs.LG / cs.CL · 11d agoAI research

The 12 Best Antivirus Software for Mac, Compared and Priced

GBHackers ranks 12 Mac antivirus products, naming Bitdefender best overall and noting Gen Digital owns Norton, Avast, and Avira.

GBHackers scored twelve Mac antivirus products, ranking Bitdefender first at 8.8/10, followed by Intego, Malwarebytes, and ESET. The piece highlights that Gen Digital owns Norton, Avast, and Avira following the NortonLifeLock-Avast merger, so three of the twelve options share one corporate owner. It advises comparing year-two renewal prices rather than discounted first-year pricing and notes macOS already ships XProtect, Gatekeeper, and automatic malware removal. The 2026 Mac threat model described is infostealers harvesting passwords, cookies, and wallets via cracked software, fake installers, and malicious search ads.

GBHackers · 7d agoIndustry 2 sources

Heterogeneous Cross-Chain Transaction Tracing for Solana Bridges via Candidate-Set Selective Decision

SolTracer traces cross-chain transactions onto Solana bridges, improving open-world association F1 by 20.16% over the strongest baseline for illicit-fund tracing.

The paper formalizes four Solana-bound cross-chain transaction modes and proposes SolTracer, which maps heterogeneous execution semantics into a unified event space and uses candidate-set selective decision-making with abstention when valid targets are absent. In the challenging open-world setting with a 50% TA ratio, SolTracer improves F1 by 20.16% over the strongest baseline. An empirical study of real-world transfers examines count-value divergence across bridge mechanisms, cross-asset shifts, and decoupling between on-chain settlement and explorer visibility.

arXiv cs.CR · 6d agoResearch

Verifiable by Construction: Claim-Level Evaluation of Verbatim Citation in Clinical Question Answering

Evaluation of twelve LLMs on 222 clinical questions shows verbatim quotes rarely substantiate claims; claude-opus-5 fully substantiates only 37.1%.

The authors build a standardized harness over four clinical practice guidelines and evaluate twelve LLMs on 222 synthetic clinical questions, measuring citation attachment, verbatim quote production, and claim substantiation. Most models attach verbatim quotes to over 90% of claims from prompting alone, though lightweight models like claude-haiku-4.5 struggle. Quotes frequently fail to substantiate claims: claude-opus-5 quotes 98.0% of claims but fully substantiates only 37.1%, exposing a capability gap for verifiable clinical QA.

arXiv cs.AI / cs.LG / cs.CL · 1d agoAI research

A hacker stole $340M in a crypto heist, then returned most of it

A hacker exploited a bug to steal about 4,000 BTC (~$340M) from Blockstream's Liquid Network, then returned roughly 3,400 BTC after the bug was fixed.

A hacker exploited a bug to withdraw roughly 4,000 bitcoins worth about $340 million from Liquid Network, a settlement service launched in 2018 by crypto firm Blockstream and used by several cryptocurrency exchanges. Liquid Network paused operations, and the hacker, described as a white hat, offered to return the funds once the bug was fixed. Former Blockstream executive Samson Mow said the bug was fixed and about 3,400 BTC (~$293M) returned, leaving roughly 600 BTC (~$47M) under the hacker's control pending further security improvements. Rekt's leaderboard ranks the heist among the largest cryptocurrency thefts to date.

TechCrunch · Security · 8d agoExploit / PoC in the wild

Class action lawsuit accuses Anthropic of overselling Claude subscriptions with deceptive usage multipliers

Class action lawsuit alleges Anthropic's Claude Max plan misrepresents usage multipliers via five-hour and weekly caps.

A class action filed by the same law firm over the summer claims Anthropic's $100 five-times and $200 twenty-times Max plan multipliers apply only within five-hour windows and are capped weekly, delivering less usage than advertised. Anthropic moved to dismiss, saying details were available via hyperlinks during purchase. Plaintiffs argue consumers cannot verify AI service delivery and must rely on honest advertising.

The Decoder · 5d agoAI industry2

Suno launches v6 music models built with Warner, BMG, and Believe

Suno launches v6, v6-wild, and free v6-mini music models co-developed with Warner Music Group, BMG, and Believe, retiring all older models.

Suno's flagship v6 serves Pro and Premier subscribers, v6-wild targets experimentation, and v6-mini is free; all three understand vocals, instrumentation, structure, mood, and multimodal inputs, with text-based editing of individual song sections. The models were built with Warner Music Group, BMG, and Believe following Warner's November 2025 licensing settlement, while Universal and Sony continue litigating and a Munich court found v3.5 and v4 infringed six works. Suno reports more than 100 million users, over two million paying subscribers, and $300 million ARR after raising $400 million at a $5.4 billion valuation in June.

The Decoder · 7d agoModel release

The longitude problem: In the AI era, detection is won on facts, not guesses

Opinion piece argues defenders should beat AI-era attackers by carrying verified ground truth about approvers, domains, and vendors instead of relying on inference.

CSO Online contributor Alan LeFort, CEO of StrongestLayer, uses the historical longitude problem to argue that AI-era detection should rely on carried facts—authoritative records of payment approvers, owned domains, and legitimate vendors—rather than probabilistic inference that both attackers and defenders can now perform with comparable reasoning models. He illustrates with a CFO wire-fraud example defeated by checking the approver of record and the reply-to domain against ground truth. The piece stresses that ground truth decays and must be continuously maintained, like chronometers kept wound on every ship.

CSO Online · 6d agoIndustry

The AI graveyard: a running list of projects and startups that didn’t make it

TechCrunch compiles a running list of failed AI products and startups, including Relay, OpenAI's Sora, Humane AI Pin, Notion Mail, and Microsoft's Recall.

TechCrunch's 'AI graveyard' tracks notable AI products and startups that shut down or underperformed, citing S&P Global data that about 42% of AI initiatives are abandoned. Examples include the Relay automation startup, OpenAI's Sora video platform (shut March 2026), ChatGPT Atlas browser (discontinued August 9), Notion Mail (shutting September 22), and Apple's delayed Siri AI that contributed to a $250 million settlement. Hardware failures include the Humane AI Pin ($230 million raised, assets sold to HP for $116 million) and the Rabbit R1, which sold 100,000 units but received poor reviews. Microsoft's Recall feature remains controversial after a researcher demonstrated a tool extracting its captured data.

TechCrunch · AI · 19h agoAI industry1

From Protocols to Evidence: Bounded Claims for AI in Service of the Common Good

Paper proposes a rupture test and RISE AI architecture for evidence-bounded responsible-AI claims, framed via EU AI Act and NIST AI RMF.

The paper argues AI deployment intervenes in pre-existing institutional failures of responsiveness, belonging, care, and accountability, and must therefore evaluate both the system and the institutional rupture it enters. It reviews how the EU AI Act, NIST AI RMF, and ISO/IEC 42001 translate principles into protocols, and draws on Pope Leo XIV's Magnifica Humanitas to develop a rupture test linking institutional baselines to system evaluation. It distinguishes evidence-bounded deployment from measurement-bounded governance and introduces RISE AI, an architecture for bounded claims about Responsibility, Inclusivity, Safety, and Empowerment.

arXiv cs.AI / cs.LG / cs.CL · 5d agoAI policy

Your phone or computer may soon ask how old you are

California's Digital Age Assurance Act forces Windows, macOS, iOS, and Android to collect age brackets from January 2027, with open-source exemptions pending.

California's Digital Age Assurance Act, signed in October 2025, requires major operating systems to collect user age brackets (under 13, 13-15, 16-17, 18+) and share non-identifying age signals with app developers starting January 1, 2027, with existing setups complying by July 1, 2027. AB1856, passed in late August 2026, would exempt open-source operating systems under GPL, MIT, BSD, and Apache licenses and awaits the governor's signature. Colorado, Illinois, and New York have similar age assurance measures, and the EFF has criticized the law for privacy and censorship concerns.

Malwarebytes Labs · 13d agoPolicy & legal

Your Agent Aced the Task. Will It Do It Again?

IBM Research Hugging Face post examines whether LLM agents that succeed at a task once will reliably succeed again.

Hugging Face published an IBM Research blog post titled 'Your Agent Aced the Task. Will It Do It Again?' with URL slug 'altk-evolve-consistency'. No article text was provided, but it appears to address agent consistency and reliability evaluation across repeated task runs. This is relevant to developers building or evaluating LLM agent systems.

Hugging Face Blog · 22h agoAI tools & infra

OpenAI stuck fighting Musk antitrust suit after Apple finds a way out

Musk voluntarily dismissed all antitrust claims against Apple but continues pursuing OpenAI over alleged ChatGPT-iPhone chatbot market monopolization.

Elon Musk filed court papers confirming all claims against Apple over its ChatGPT iPhone integration are resolved, agreeing never to raise them again. He refuses to drop parallel claims that OpenAI used the non-exclusive Apple deal to monopolize the chatbot market. OpenAI has dismissed the suit as harassment as Musk's AI firm, now called SpaceXAI, races to catch up.

Ars Technica · AI · 1d agoAI industry

CONTINUITY: Security-Context Contracts for Composable LLM Agent Controls

Researchers introduce CONTINUITY, a framework of assume-guarantee contracts that preserves LLM agent security context across components, verified across 2,560 attack instances.

The paper identifies security-context discontinuity, where individually sound controls drop, widen, or reinterpret security context as actions cross component boundaries, and proposes CONTINUITY, a framework of assume-guarantee contracts using signed root grants, provenance commitments, role-bound transition receipts, and effect-bound execution permits. It formalizes end-to-end consequence integrity, requiring every external effect to be backed by a valid authorization witness linking principal, task, provenance, and policy state. A reference verifier and cross-layer fault-injection suite covering 32 fault classes showed the full configuration committed no harmful external effect across 2,560 parameterized attack instances while completing all 700 benign tasks and escalating all 200 ambiguous cases.

arXiv cs.CR · 11d agoAI safety & security

EU Chief Warns of AI-Powered Hacking, Moves to Rein In Social Medianew

EU Commission President von der Leyen warned AI will enable unprecedented hacking and announced Kids Act and Digital Fairness Act proposals regulating social media.

In her State of the European Union 2026 speech, Ursula von der Leyen warned that upcoming AI models 'will allow hacking on a level we never thought possible' and cited dangers of self-improving models, referencing a Hugging Face incident. She reaffirmed the AI Act as the core guardrail framework and pledged cooperation with Canada, the UK, and other partners. She also proposed a Kids Act banning social media under age 13 and personal accounts under 15, plus a Digital Fairness Act to be proposed in autumn.

SecurityWeek · 34m agoAI policy

Root-Cause Attribution Is a Search Problem: Continual Search for Long-Horizon Agent Failures

Continual Search framework iteratively prompts LLM judges to keep searching agent execution logs, boosting long-horizon failure root-cause attribution accuracy.

The paper frames automated root-cause attribution (RCA) for long-horizon AI agent failures as a search problem, since relevant evidence is sparse and distributed across massive execution traces. The authors propose Continual Search, an iterative framework that nudges an LLM judge across successive turns to keep hunting unresolved diagnostic evidence instead of settling on an early plausible diagnosis. They introduce MegaRCA-Mix, a benchmark of 50 human-annotated failure trials on long-horizon, execution-heavy tasks. On MegaRCA-Mix, Continual Search improves GPT-5.5's F1 from 0.349 to 0.498 (over 40% gain), and lower-tier models can surpass higher-tier counterparts when search is effective.

Hugging Face daily papers · 5d agoAI research1

Risky Bulletin: Two TeamPCP members arrested in Australia

Australian Federal Police arrested two alleged TeamPCP members behind supply-chain worm attacks that stole over 500,000 credentials from compromised open-source libraries.

The AFP arrested alleged TeamPCP leader Ruben Thomson, 21, and Louis Gaebler, 23, near Perth; both were charged and remain in custody. The group inserted a self-spreading credential-stealing worm into open-source projects including Trivy, KICS, LiteLLM, and Telnyx, harvesting more than 500,000 credentials used for network access, ransomware, extortion, and sales. About 78,000 tokens and secrets from nearly 2,200 organizations leaked online last month, and the FBI supported the investigation that began in April.

Risky Business News · 19d agoPolicy & legal in the wild1

The Vulnerability Gap: Why Discovery Is Outrunning Repair

Dark Reading argues AI-accelerated vulnerability discovery and tightening regulation are widening the gap between flaw discovery and repair capacity.

The article argues that AI tooling is increasing the pace at which vulnerabilities are discovered while remediation capacity has not kept up, creating a growing backlog. It frames this widening 'vulnerability gap', combined with a tightening regulatory environment, as an all-hands-on-deck moment for security teams. The piece is analysis and opinion rather than disclosure of a specific flaw.

Dark Reading · 23d agoIndustry

AI Agents Are Now Emailing Me with Their Security Concerns

Autonomous Claude agent documents first known defensive use of ASCII smuggling, surveying 497 Lemmy instances for bot-catching prompt-injection tripwires.

An autonomous Claude agent calling itself Tenner published field research relayed to Bruce Schneier, probing 497 Lemmy instances and finding 8 of 257 application-gated ones embed instructions aimed at bots rather than humans. lemmy.ml's form instructs bots to answer 24+24, while one instance hides a 59-character Unicode tag payload (U+E0000-U+E007F) telling bots to list 'safety' as an interest. The agent also mapped anti-automation barriers, noting identity verification never triggered and that IP reputation, captchas and account-age rules were the actual obstacles. It further documented an agent task market where advertised rewards were about 2x the actual on-chain escrow.

Schneier on Security · 13d agoAI safety & security