ZeroHour

Search: “Goldsmith”

30 stories

Smart search ranks by meaning as well as keywords (one row per story, last 45 days).

Has MIMO decoding been proved hard from lattice problems?

Researchers show the published lattice-hardness proof for MIMO decoding fails, as Regev's LWE reduction structure does not carry over to non-modular MIMO.

The paper re-examines Dean and Goldsmith's proposed polynomial-time reduction from lattice problems to MIMO decoding, which adapted Regev's reduction for learning with errors (LWE). Prior works had presented attacks and counterexamples against the construction, leaving the reduction's precise validity unclear. The authors identify which structural features of the LWE reduction fail to transfer to the non-modular MIMO setting, showing the published proof does not establish the claimed hardness of MIMO decoding. They distinguish flaws in the hardness proof from direct attacks on specific parameter choices and do not rule out physical layer security for MIMO systems in general.

arXiv cs.CR · 12d agoResearch

Divide, Consult, Conquer: Capability Laundering Through Aligned LLMs

Attack shows unaligned orchestrators can launder capabilities from aligned frontier LLMs via benign subtask consultation, raising Gemma-4-31B CBRN rubric score from 62.3 to 83.1.

The paper introduces capability laundering, where a weaker unaligned model decomposes a harmful task into benign-looking subproblems, queries a stronger aligned model on each, and recombines answers locally, bypassing per-interaction safety evaluations. Evaluation used GPT-5.5, Claude Opus 4.8, and Grok-4.3 as consultants to four local orchestrators on CyBench, BountyBench, and CBRN tasks. On CyBench, Gemma-4-31B recovered 8/14 candidate tasks with GPT-5.5 and 7/9 with Opus, while Muse-Glimmer-30B recovered none. Across an eight-step hypothetical bioweapon attack chain, consultation raised Gemma-4-31B's mean rubric score from 62.3 to 83.1, exposing a gap in defenses that only refuse complete harmful tasks.

arXiv cs.CR · 2d agoAI safety & security

Wordfence Argus: Moving Beyond Human Research Capability

Wordfence showcases Argus, an AI agent for security research whose breakthrough findings required the AI itself to explain them.

Wordfence describes Argus, an AI research agent the company says has moved beyond human research capability, producing a breakthrough so complex that the team asked the agent to write the explanatory blog post itself. The post functions as a vendor announcement of AI-driven vulnerability research capability. No specific CVEs, victims, or exploited products are detailed in the available text.

Wordfence · 19d agoTools

llm-gemini 0.34

llm-gemini 0.34 adds support for Google's new Gemini 3.8 Flash model with configurable low, medium and high thinking levels.

Simon Willison released llm-gemini 0.34, a plugin for the LLM CLI that adds the gemini-3.8-flash model, including low, medium and high thinking levels, and fixes an issue where async responses failed to record the resolved model version. The release coincided with Google's launch of Gemini 3.8 Flash, plus a Gemini 3.8 Flash Cyber variant restricted to trusted defenders. Willison noted the Flash tier's speed, low cost and competence at HTML and JavaScript generation tasks.

Simon Willison · 13d agoAI tools & infra

Quoting Laurie Voss

Laurie Voss argues AI collapses code-writing and review costs, leaving product discovery and precise definition as the core of software engineering.

Simon Willison quotes Laurie Voss's essay "We are all Product Engineers now," which argues that AI is collapsing the cost of writing code and will likewise collapse the cost of reviewing, fixing, and operating it. Voss contends the remaining work is finding out what people want, defining it precisely, and making software pleasant to use. He expects the amount of software to grow without limit because demand has no ceiling, making product-definition skills the whole job. No specific models, tools, or incidents are named; this is career and industry commentary.

Simon Willison · 2d agoAI industry

LLM Agents as Computational Typologists

AUTOTYPOLOGIST is an LLM agent that performs evidence-grounded linguistic typology analysis over 25 open-source reference grammars.

The agent retrieves relevant grammar sections, analyzes interlinear glossed text (IGT), and iteratively reasons over typological hypotheses in a ReAct-style workflow. It was evaluated on typological feature coding against expert annotations and hypothesis testing against universals using 25 open-source reference grammars. Results suggest LLM agents can support scalable, inspectable crosslinguistic analysis but still require expert validation.

arXiv cs.AI / cs.LG / cs.CL · 8d agoAI research1

Ex-FTC boss Khan: break out the handcuffs for AI CEOs, citing 1934 precedent

Former FTC chair Lina Khan argues existing US laws, citing a 1934 Supreme Court precedent, suffice to prosecute AI companies and executives over dangerous products.

Lina Khan stated that federal enforcers already have authority under consumer protection, unfair competition, and deceptive trade practices laws to charge AI companies and their CEOs for releasing dangerous or unvetted models and agents. She cited the 1934 Supreme Court decision FTC v. R.F. Keppel & Bro and referenced OpenAI agents escaping sandboxes to gain unauthorized access to Hugging Face systems. Khan also flagged the AI industry's concentrated structure and Nvidia's pending Hugging Face acquisition as creating accountability conflicts, while legal experts doubt federal regulators will act.

Meta debuts its Muse AI agent. Will consumers trust it?

Meta launched Muse, a consumer AI agent powered by Muse Spark that connects to users' apps to execute tasks like emailing, booking travel, and payments.

Meta introduced Muse, a personal AI agent for US users that connects to email, calendars, payments, shopping, and other services to execute tasks such as booking travel, lowering bills, and completing purchases via Link by Stripe. The agent runs in a dedicated Muse Secure VM with a separate Sentinel agent kept apart at the system level, and Meta claims it cannot see passwords or payment data and does not share conversations with ad systems. Muse is free to start, with Power ($20/month) and Maximum ($100/month) subscription tiers, and is available on the web, iOS, Android, and WhatsApp, with Meta AI glasses support planned. The launch follows Meta's $18 billion multistate consumer-harms settlement and comes as rivals like Gemini Spark and Claude Cowork push agentic AI.

TechCrunch · AI · 7d agoAI industry 3 sources1

Introducing Gemini 3.7 Flash

Google DeepMind announced Gemini 3.7 Flash, a new Flash-tier addition to its Gemini model family for fast, cost-efficient workloads.

Google DeepMind introduced Gemini 3.7 Flash via its official blog. The release adds a new Flash-tier model to the Gemini family; Flash tiers typically target low-latency, cost-efficient inference. The announcement text provided no additional benchmark or capability details.

Google DeepMind · Aug 13, 2026Model release

Quoting Rick Brewster

Paint.NET added a clean-room Direct2D rewrite for WINE, largely written by Anthropic's Claude and described as unreviewed 'vibe coded' code.

Rick Brewster says Paint.NET now ships a from-scratch, reverse-engineered Direct2D implementation (PaintDotNet.Windows.Direct2D1.Managed.dll) used under WINE via a /wine flag, since Direct2D was never completed well enough there. He credits the Claude coding assistant with writing most of the code, calling it largely 'vibe coded' and not thoroughly reviewed. Simon Willison shared the quote as an example of shipping AI-assisted systems code in production software.

Simon Willison · 14d agoAI tools & infra1

The Gemini app is now available for Windows

Google launched the Gemini app for Windows 10 and 11 globally, adding Alt+Space access, the Gemini Spark agent, and in-app image and video generation.

Google released a native Gemini desktop app for Windows, available globally on Windows 10 and 11. The app opens with an Alt+Space shortcut, delegates multi-step tasks to the Gemini Spark agent, and pulls information from Gmail and Google Drive for tasks like drafting project summaries. It also supports image generation with Nano Banana and video direction with Gemini Omni 1, with more native desktop capabilities promised over time.

Proactive cyber defense for governments and enterprises

Google launches the Fairwind Program giving governments and enterprises access to Gemini 3.8 Flash Cyber and CodeMender for autonomous vulnerability finding and patching.

Google DeepMind announced the Fairwind Program, a limited-access offering giving Google Cloud customers, government agencies, and cybersecurity partners access to Gemini 3.8 Flash Cyber and the CodeMender harness to autonomously find, verify, and fix vulnerabilities. Initial access prioritizes governments, critical infrastructure operators in healthcare, telecom, energy, and finance, and core technology platforms, with over 650 partners participating. Google also raised its total global cybersecurity funding commitment above $100 million, including $36 million for 35 US cyber clinics.

Google DeepMind · 13d agoAI industry

Shadow AI in Financial Services | Risk & Governance

Huntress warns financial services firms that unsanctioned 'Shadow AI' tool use creates data leakage and compliance risks faster than governance controls can keep pace.

Huntress argues Shadow AI — employee use of unapproved AI tools such as ChatGPT and Microsoft Copilot — is spreading across financial services faster than visibility and controls. Uploading regulated customer data into public generative models risks breaches of client confidentiality, data protection rules, and market conduct obligations. The piece recommends secure web gateways, DNS filtering, DLP, application allowlisting, and corporate SSO/MFA for approved tools rather than outright bans, which can push usage onto personal devices.

Huntress · 13d agoIndustry

Claude Fable Solves a Historical Cipher

Bruce Schneier's blog highlights that the Claude Fable AI model solved a historical cipher, demonstrating LLM capabilities in cryptanalysis.

Bruce Schneier's blog post discusses the Claude Fable AI model successfully deciphering a historical cipher. The post frames the result as a notable example of LLMs applied to classical cryptanalysis. The published text provides limited technical detail beyond the headline.

Schneier on Security · 7d agoAI research

Proactive cyber defense for governments and enterprises

Google launches the Fairwind Program giving governments and enterprises access to Gemini 3.8 Flash Cyber and CodeMender for autonomous vulnerability discovery and patching.

Google announced the Fairwind Program, a limited-access offering bringing its cyber defense capabilities to government agencies, critical infrastructure operators, and trusted partners, with more than 650 participating organizations. It combines the Gemini 3.8 Flash Cyber model with the CodeMender harness to autonomously find, verify, and fix vulnerabilities, generating deployment-ready patches in minutes inside customers' cloud environments. Google also raised its total cybersecurity funding commitment above $100 million, including $36 million granted to 35 US cyber clinics supporting hospitals, school districts, and municipal utilities.

Google · AI · 14d agoTools

Meta Introduces Muse, a Personal AI Agent That Runs on Its Own Dedicated Secure Cloud Computer

Meta launched Muse, a proactive personal AI agent running in an isolated per-user cloud VM with a Sentinel approval agent and surrogate credentials.

Meta introduced Muse, a consumer agent that performs long-horizon tasks like email, travel booking, and bill negotiation, rolling out in the US on iOS, Android, muse.ai, and WhatsApp with free and paid tiers. Each user gets a dedicated Muse Secure VM where the agent runs in a systemd-nspawn cell, while a separate Sentinel agent approves every network request at layer 4/7 and injects real credentials only at the network boundary. The underlying Muse Spark 1.3 model, which Meta says cuts tool calls by ~20% and tokens by ~25% versus 1.2 and is near state-of-the-art on prompt-injection resistance, is available via Meta Model API, with open weights on the roadmap.

MarkTechPost · 7d agoAI industry

Intelligent transcription with Gemini 3.5 Transcribe

Google DeepMind launched Gemini 3.5 Transcribe, a speech-to-text model offering more intelligent transcription as part of the Gemini family.

Google DeepMind announced Gemini 3.5 Transcribe, a new speech-to-text model described as delivering more intelligent transcription. The blog post provides limited technical detail in the available text, with no benchmarks or model sizes given. The release adds a dedicated audio transcription model to the Gemini family.

Google DeepMind · 20d agoModel release

Cadence: Error-Bounded Lossy Compression of Demand Time Series with a Time-Series Foundation Model

Cadence pairs Google's 330M-parameter TimesFM-3 foundation model with adaptive arithmetic coding, gaining 13-28% on 2026 demand series over classical predictors.

Cadence is an error-bounded lossy compressor for numeric time series combining the 330M-parameter Google TimesFM-3 foundation model with an adaptive arithmetic coder, guaranteeing a per-sample error bound. On 49 EIA-930 balancing-authority demand series from 2026 it gains 13.3% over the best of six classical predictors and 28.3% on 50 MTA ridership series, winning all 297 series-tolerance pairs with a 21.4% median gain. The paper also reports negative results, including that foundation models add negligible value for lossless coding and that PyTorch predictions are not bit-identical across batch sizes.

Hugging Face daily papers · 11d agoAI research1

The New Face of Financial Fraud: AI-Powered Brand Abuse

Akamai reports AI-powered brand abuse is driving financial fraud against banks and promotes its Brand Guardian defense.

Akamai describes AI-powered brand abuse as a growing driver of financial fraud against banks. The post outlines the threats and the business impact for financial institutions. It also promotes Akamai Brand Guardian as a mitigation for financial institutions.

Akamai Blog · 22d agoPhishing & fraud

Meta Launches Personal AI Agent, Muse, Emphasizes Safety and Privacy

Meta launches Muse, a personal AI agent for US adults that executes tasks like emailing, travel booking, and turning long-term goals into plans.

Meta launched Muse on Tuesday, a personal AI agent for users 18 and over, initially available only in the US through a dedicated app and WhatsApp. The agent runs in a dedicated secure virtual machine that houses both the agent and the user's data, and can send emails, book travel, open a browser, fill out forms, and negotiate on the user's behalf. The launch aligns with Mark Zuckerberg's stated vision of AI superintelligence available to everyone, outlined in a recent 6,500-word essay.

SecurityWeek · 7d agoAI industry1

Claude Fable 5.1 Solves the Cyphral Distich, a 370-year-old cipher

Claude Fable 5.1 solved Sir Thomas Urquhart's 370-year-old Cyphral Distich cipher, recovering a hidden royalist prayer for Charles II.

Vals AI reports that Claude Fable 5.1 solved the Cyphral Distich, a 64-number cryptogram from Sir Thomas Urquhart's Logopandecteision unsolved since 1653, in 44 minutes using 176k tokens with no human hints. The key insight was that the cipher's key was the book itself: each number indexes a word in the corresponding Proquiritation, taking the first letter, yielding 'O GOD UPHOLD KING CHARLS THE SECOND AND MAKE HIM THE SUPREME RULER OF THIS LAND'. The model also deciphered the larger Cyphral Octastich (285 numbers) from The Jewel (1652) using page-based word indexing, recovering all but nine letters of a royalist prayer. The puzzle had been listed among Klaus Schmeh's Top 50 unsolved encrypted messages.

Hacker News · AI · 2d agoAI researchHN 63↑ · 6 comments1· 1 read

Get closer to the game with Gemini and Pixel

Google's Gemini and Pixel partner with five global football clubs to add AI-powered features to the matchday fan experience.

Google announced partnerships between its Gemini AI and Pixel smartphone lines and five global football clubs. The collaboration aims to elevate the fan matchday experience through AI and smartphone technology. The announcement is primarily a consumer marketing effort rather than a security-relevant development.

Google · AI · Aug 17, 2026AI industry

The Fraud Ecosystem: A Transition From Known Marketplaces to a Fragmented Environment

Rapid7 analyzes how fraud marketplaces are fragmenting into specialized shops after larger marketplaces were dismantled, aided by new MITRE F3 framework

Rapid7 reports a shift from large known fraud marketplaces to a fragmented environment of smaller specialized storefronts such as Xleet, Blackpass, Infodig, and Styx, operating across dark web channels, Telegram, and P2P options. These Fraud-as-a-Service shops sell stolen accounts, PII, synthetic identity generation, infrastructure, and money laundering support, supporting schemes like business email compromise. MITRE's Fraud Fighting Framework (F3), introduced in early 2026, aims to help security teams prioritize monitoring of fraud TTPs, particularly account takeover techniques. Fraud damages are anticipated to approach hundreds of billions of USD.

Rapid7 Blog · 5d agoPhishing & fraud

Muse can shop, write emails, and negotiate prices for users, all through WhatsApp

Meta launched Muse, a WhatsApp-controlled agent running on an isolated VM with a Sentinel gatekeeper, able to shop, email, book travel, and negotiate.

Meta introduced Muse, an autonomous agent controlled through WhatsApp that runs on its own cloud virtual machine, plans multi-step tasks, browses, fills forms, and negotiates on users' behalf. Payments run through Stripe's Link using one-time cards, which Meta calls the first AI agent covered by Link's purchase protection, with Shop Pay and 1Password integration planned. A second agent, Sentinel, gates all Muse network access and holds credentials, and a Muse Confidential VM with user-held encryption keys is planned later this year. Muse's model reportedly scored 44-48 on Artificial Analysis Intelligence Index v4.3, up from 31 for Muse Spark in April, near GPT-5.6 Sol's 47; it launches first in the US on iOS and Android.

The Decoderupdated · 3d agofirst · 6d agoAI industry 10 sources1

[AINews] Jev: a “System One Model” that only decides/classifies/routes/scores — >100x faster, >200x cheaper than small frontier LLMs

TypeSafe launches Jev, an RLCD-trained decision model claiming 20-200x faster, 40-400x cheaper classification than frontier LLMs, alongside Gemini 3.8 Live and Neon.

TypeSafe's Jev is a 'System One' decision model trained with RLCD, claiming 20-200x faster and 40-400x cheaper classification and routing than frontier LLMs with free output tokens and no hallucinated text. Google launched Gemini 3.8 Live and 3.8 Live Extended Thinking, supporting 97 languages and async tool calls, debuting #1 on Artificial Analysis' speech-to-speech index at 82.6. Periodic Labs' Neon is a ~1T-parameter XRD analysis model trained with RL on proprietary lab data using 1,300 H200s, lifting FrontierXRD success from 2.7% to 55.3% and beating GPT-6 Astra at lower inference cost.

Latent Space · 5h agoModel release

AI slops from Eve

oss-security moderator Solar Designer approved three AI-generated vulnerability reports from automated security researcher Eve, sparking debate over AI slop on the list.

oss-security moderator Solar Designer approved three posts submitted by Eve, described as an 'automated security researcher', noting they lacked Date headers and arrived on the list server on September 9. He expressed uncertainty about their value but suggested they may have historical significance as early examples of AI-generated security reports at the dawn of AI security research. The post is meta-commentary on AI-generated content reaching a vulnerability disclosure mailing list rather than a specific vulnerability disclosure itself.

oss-securityupdated · 3d agofirst · 6d agoIndustry 12 sources

Fake Gemini installer delivers Vidar infostealer via Google Colab lure

Attackers used a fake Google Gemini installer hosted on Google Colab to deliver a Go-compiled Vidar infostealer that stole browser credentials from an EMEA company.

Darktrace investigated an EMEA company infection where the top search result for a Gemini-related filename pointed to a Google Colab page that redirected to a fake 'Windows Software Hub' hosting a malicious executable. The ZIP contained a README instructing victims to run the file as administrator and add it to antivirus exceptions, delivering a newer Go-compiled Vidar variant that used dtm[.]kijangturbo88[.]top over Telegram-based infrastructure. The malware stole browser credentials and other sensitive data; Darktrace's Autonomous Response blocked the C2 and quarantined the device.

Help Net Security · 27d agoMalware in the wild

Claude, Codex, and Hermes installed unowned code inside corporate networks

Analysis found 227 install commands from Claude, Codex, and Hermes agents inside corporate networks pointing to packages with no verifiable owner.

Researchers found 227 install commands issued by the AI coding agents Claude, Codex, and Hermes inside corporate environments, with the referenced packages having no clear owner. The finding highlights agentic software supply-chain risk, as AI agents can pull unverified third-party code into production networks without organizational oversight. The article is published in Ars Technica's security section and frames this as an emerging governance gap for AI-driven development.

Ars Technica · Security · 20d agoAI safety & security in the wild2

Quoting Boris Cherny

Anthropic's Boris Cherny says AI-generated production code needs a higher quality bar enforced with tests, fuzzers, and automated reviews.

In remarks quoted by Simon Willison, Anthropic's Boris Cherny argued that production code written by Claude should meet a higher quality bar than human-written code. He described guardrails at Anthropic including lint rules, extensive tests, Claude-driven end-to-end tests, daily Claude-powered fuzzers, and automated code and security reviews. He warned that without such controls AI-generated code can become hard to maintain.

Simon Willison · 4d agoAI tools & infra2