ZeroHour

Search: “Sality operators”

40 stories

Smart search ranks by meaning as well as keywords (one row per story, last 45 days).

Global sinkhole operation ends Sality botnet’s 23-year run

Law enforcement, CrowdStrike, and Shadowserver sinkholed the 23-year-old Sality P2P botnet, cutting 15,000+ infected machines from its operator.

Sality, active since 2003 as a file-infecting virus with two P2P networks (versions 3 and 4), distributed credential thieves, spam, proxies, and DDoS payloads, and most recently delivered the EggJagger clipboard hijacker that swapped cryptocurrency wallet addresses for at least $150,000 in operator profit. A coordinated sinkhole operation replaced the botnet's super-peer list with defender-controlled sinkholes, and investigators in the US, Bulgaria, Hungary, and Romania seized payload domains. The Shadowserver Foundation is coordinating ISP and CERT notifications to infected device owners.

Help Net Security · 14d agoMalware1

Local gradient neural operator

Researchers propose LGNO, a lightweight interpretable neural operator using learnable local stencils, matching global-operator accuracy on PDE benchmarks with fewer parameters.

LGNO builds on nonlinear gradient discretization priors and uses multilayer perceptron convolutional layers to learn translation-invariant local kernels resembling discrete stencils. A zero consistent stencil factorization separates coefficient learning from field reconstruction, and network folding shares equivalent components to cut parameter counts for symmetric problems. Evaluations on linear and nonlinear, static and dynamic, and low- and high-dimensional PDE benchmarks show maintained accuracy, parameter efficiency, and rollout stability, with applicability to diffusion, flow, and quantum problems.

arXiv cs.AI / cs.LG / cs.CL · 9d agoAI research1

International Operation Disrupts Sality P2P Botnet

US-led international operation with Europol, CrowdStrike, and Shadowserver sinkholed the 20-year-old Sality P2P botnet, once exceeding one million infected machines.

On August 31, 2026, authorities from the US, Bulgaria, Hungary, and Romania, supported by Europol, CrowdStrike, and the Shadowserver Foundation, disrupted the Sality P2P botnet by sinkholing communications and seizing domains. Sality has operated for over 20 years, at its peak controlling more than one million infected machines used for credential theft, spam, proxy services, crypto-theft, and DDoS attacks, with over 11 million unique IP addresses linked to its infrastructure since 2017. The disruption exploited the botnet's super-peer reputation mechanism by removing legitimate peers via protocol-level manipulation and inserting sinkhole entries into emptied peer lists.

Infosecurity Magazine · 13d agoMalware

Nonmaximal sums of maximally monotone operators under Rockafellar's constraint qualification

Mathematical paper constructs counterexamples on c0 and l1 disproving Rockafellar's conjecture that sums of maximally monotone operators remain maximally monotone.

The authors build counterexamples where two maximally monotone operators satisfy the interior-domain condition yet their sum is not maximally monotone, refuting Rockafellar's sum conjecture. One counterexample is constructed on c0 and another on l1 with its usual norm. A general construction theorem computes the monotone polar of a class of graphs, gives necessary and sufficient conditions for maximal monotonicity, and shows how a positive rank-one perturbation yields a nonmaximal sum.

arXiv cs.AI / cs.LG / cs.CL · 7d agoAI research

Sality, one of the longest

US and European authorities, with CrowdStrike and Shadowserver, disrupted the 20-year-old Sality peer-to-peer botnet, severing 15,000+ infected machines from operators.

US and European authorities disrupted the Sality botnet, active since at least 2003, in an operation involving the DOJ, CrowdStrike, the Shadowserver Foundation and agencies in Bulgaria, Hungary and Romania. Researchers reverse-engineered the botnet's peer-to-peer architecture and injected false data into infected machines' 'super peer' lists, cutting more than 15,000 systems off from their operators. For the past eight years Sality primarily distributed EggJagger, malware that replaces clipboard cryptocurrency addresses and is estimated to have netted the operator at least $150,000. No arrests were announced, and CrowdStrike assesses the operator works from Russia's Bashkortostan region.

The Record · 14d agoMalware in the wild

Searching for New Physics with Reinforcement Learning

Researchers apply reinforcement learning to identify SMEFT operators explaining particle physics anomalies, reproducing and improving known CDF W-mass results.

The paper introduces a reinforcement learning method to search the large Standard Model Effective Field Theory (SMEFT) operator space for explanations of measurement anomalies. It was validated on the CDF W-mass anomaly, reproducing and improving known results, then applied to a harder multi-anomaly scenario. RL efficiently navigates complex loop-level operator correlations that bias human-driven phenomenological analysis.

arXiv cs.AI / cs.LG / cs.CL · 7d agoAI research

Dogged Russia-based botnet dismantled after 23-year run

Law enforcement, CrowdStrike and Shadowserver dismantled the 23-year-old Sality P2P botnet that infected more than 11 million devices.

Sality, a Russia-based peer-to-peer botnet active for 23 years and infecting over 11 million devices, was dismantled by law enforcement working with CrowdStrike and the Shadowserver Foundation. CrowdStrike poisoned the botnet's peer list so infected machines permanently disappeared from the operator's view, while domains were seized in a coordinated effort involving the FBI, Justice Department, Europol and authorities from Bulgaria, Hungary and Romania. The financially motivated operation enabled cryptocurrency theft, DDoS attacks and other cyberattacks, and Europol said the effort dates back to 2017; the operators were not named.

CyberScoop · 14d agoMalware

A Generalization of Amari's Bayesian Duality

Paper generalizes Amari's Bayesian duality by connecting it to a convex duality of Bayes' rule.

The authors revisit Amari's less-known work on Bayesian duality from information geometry. They connect Bayesian duality to a convex duality formulation of Bayes' rule and present a generalization of it. The paper is purely theoretical and discusses relevance for modern AI, with no experiments or model releases.

arXiv cs.AI / cs.LG / cs.CL · 8d agoAI research

Authorities Turn Sality's P2P Network Against Itself, Cutting Off New Malware Payloads

US and European authorities with CrowdStrike dismantled the two-decade-old Sality P2P botnet using peer-list manipulation, blocking payload delivery to infected hosts.

The US Department of Justice announced a coordinated takedown of the Sality peer-to-peer botnet, executed August 31, 2026 by authorities from the US, Bulgaria, Hungary, and Romania with CrowdStrike and the Shadowserver Foundation. Sality, active since 2003, infects Windows executables and delivers payloads including the EggJagger crypto clipper, which stole at least $150,000, and was used for DDoS campaigns. The operation abused the botnet's peer-list maintenance cycle to insert sinkhole nodes and isolate both super peers and NAT-hidden infections, cutting off URL and payload distribution to more than 15,000 infected machines across two P2P networks.

The Hacker News · 14d agoMalware

Cross-modal learning for SAR target recognition using optical vision foundation models

Frozen DINOv3 optical prototypes supervise SAR target recognition without EO/SAR pairs, improving classification on the heavily imbalanced UNICORNv2 dataset.

The framework aligns SAR embeddings to class-level prototypes built from a frozen DINOv3 electro-optical encoder, requiring no strict EO/SAR image pairs. At inference the SAR model operates independently without access to optical imagery. On UNICORNv2, a civilian vehicle dataset with heavy speckle and severe class imbalance, EO prototype alignment improves accuracy over frozen DINOv3, SAR-only finetuning, and unpaired distribution alignment baselines.

arXiv cs.AI / cs.LG / cs.CL · 9d agoAI research

Cops, CrowdStrike disrupt Sality botnet by poisoning the network and diverting into sinkholes

Law enforcement and CrowdStrike disrupted the 23-year-old Sality P2P botnet, isolating 15,000+ infected machines and seizing linked domains.

International law enforcement, working with CrowdStrike and the Shadowserver Foundation, executed a peer-to-peer sinkhole operation against Sality, a botnet active since 2003 that delivered malware to more than 15,000 machines worldwide. Sality's primary payload for eight years was EggJagger, a clipboard hijacker that swaps copied bitcoin and ethereum wallet addresses with attacker-controlled ones, yielding at least $150,000 in stolen cryptocurrency. The US Justice Department, FBI, and DoD Office of Inspector General's Defense Criminal Investigative Service seized Sality-linked domains, with parallel action in Bulgaria, Hungary, and Romania. The Shadowserver Foundation is coordinating with ISPs and CSIRTs to identify infections and notify victims.

The Register · Security · 14d agoMalware in the wild

Explainability Assistant: A Conversational XAI Interface for Interpreting Energy Consumption Models

Researchers release Explainability Assistant, an open-source conversational XAI tool using LLM function calling, lifting intent-parsing accuracy from 76.8% to 94%.

The paper introduces the Explainability Assistant, an open-source conversational XAI system for interpreting energy consumption forecasting models such as genetic-programming symbolic regressors. It uses LLM function calling instead of rigid custom grammars, achieving 94% intent-parsing accuracy versus 76.8% for prior work TalkToModel, and adapts to different ML problem types without task-specific fine-tuning. Comparative evaluation with energy domain specialists against a traditional XAI dashboard showed improved usability, with all experts preferring the conversational interface.

arXiv cs.AI / cs.LG / cs.CL · 6d agoAI research1

Risky Bulletin: Russia tells data centers to deploy drone defenses

Russia ordered data center operators to deploy drone strike defenses under a Putin decree allowing temporary state takeover of unprotected critical infrastructure.

The Russian government instructed data center operators to deploy protections against drone strikes under a presidential decree signed by Putin that allows temporary state administration of critical infrastructure operators failing to defend against Ukrainian hacks and drone strikes. Although data centers are not formally critical infrastructure in Russia, the decree applies to them because other sectors depend heavily on cloud services; Russia has more than 180 data centers, over 80% in the European region within range of Ukrainian strikes. The digest also reports a Dropbox breach affecting nearly 5,000 accounts via the Lenovo ID integration, spyware attacks on at least 14 Serbians using NoviSpy or Pegasus, and a password recovery attack targeting hundreds of thousands of X accounts tied to the new X Money service. Other items include a 14-hour compromise of Coder's Cloudflare infrastructure delivering malicious Terraform modules, donor data breaches at Davayte and You Are Not Alone via the Stripe/WooCommerce integration, a $2.5M Aquifer crypto heist, and a TVING breach exposing data of almost 40 million accounts.

Risky Business News · 12d agoPolicy & legal

Can your coding style predict whether your code is vulnerable?

University of Massachusetts Dartmouth researchers present VulStyle, a stylometry-based vulnerability detector that also exposes benchmark reliability problems.

VulStyle combines stylometric features with syntax-tree structure and source tokens, pre-trained on about 4.9 million functions across seven programming languages and fine-tuned on five vulnerability detection datasets. It beat token-only detectors on some benchmarks but its F1 drops sharply on DiverseVul, which the authors link to noisy labels inflating reported performance across popular datasets. The authors argue style-aware detection should be harder to evade but did not test this empirically, and they note that uniform LLM-generated code may strip away the individual developer style the model depends on.

Help Net Security · 23d agoResearch1

America's Driver's License Breach Is a National Security Disaster

Dark web service Nexus sells 153 million US/Canadian driver's licenses linked to a breach of identity verifier IDScan.

Krebs on Security revealed a dark web service, Nexus, selling access to 153 million driver's licenses and 3 million travel documents from US and Canadian citizens, roughly 63 percent of all US licenses. Circumstantial evidence links the data to identity verification firm IDScan, which confirmed it is investigating a breach, and the FBI is probing the incident. Licenses belonging to senior US officials, including Pete Hegseth, an FBI assistant director, and Krebs's own contacts were verified as genuine. The exfiltration appears ongoing, with the database growing by nearly 400,000 licenses in a single day, and the data carries significant national security value for foreign intelligence services.

Hacker News · security · 1d agoData breachHN 26↑ · 4 comments3· 1 read

Measuring benchmark optimization in speech recognition

Hugging Face examines how much speech recognition systems overfit benchmarks and how to measure benchmark optimization in ASR.

A Hugging Face post on measuring benchmark optimization in automatic speech recognition, analyzing how model improvements on benchmarks reflect genuine capability gains versus overfitting. It is evaluation methodology research with no direct security impact.

Hugging Face Blog · 26d agoAI research

Srsly Risky Biz: America's Drivers Licence Breach is a National Security Disaster

Dark web service Nexus sold 153 million US and Canadian driver's licenses, linked to identity verification firm IDScan under FBI investigation.

Krebs On Security reported that a dark web service called Nexus sold access to 153 million US and Canadian driver's licenses, claiming over a year of continuous exfiltration from a major identity verification company, with roughly 400,000 new licences added in a single day. Krebs verified the data as genuine and linked the incident via circumstantial evidence to identity verification firm IDScan, whose licences of senior US officials including Secretary of War Pete Hegseth appeared in the database; the FBI is investigating and IDScan has confirmed a breach inquiry. The article argues the data has national security implications, citing how Chinese APT espionage (Anthem, Equifax, Marriott, OPM) and Bellingcat investigations exploited leaked databases. Class action suits are being prepared, and the piece calls for stricter oversight of identity verification firms.

Risky Business News · 6d agoData breach in the wild

ENCP: Episode-Normalized Conformal Prediction for Vision-and-Language Navigation

ENCP calibrates conformal prediction per navigation episode, giving step-level coverage guarantees for vision-language navigation agents despite within-episode dependence.

Episode-Normalized Conformal Prediction (ENCP) rescales a nonconformity score by a VLN policy's residual confidence and calibrates one maximum score per episode, preserving step-level coverage of at least 1−α despite dependence among steps within an episode. Across four VLN policies and three nonconformity scores on R2R and REVERIE, ENCP meets all reported empirical step-coverage targets in seen-to-unseen evaluation. The model-agnostic uncertainty estimates can signal when an agent should defer to a stronger predictor or human assistance.

arXiv cs.AI / cs.LG / cs.CL · 1d agoAI research

Modified ScreenConnect Clients Used in Worm-Like Campaign

Huntress warns of worm-like attacks using modified ScreenConnect clients to spread VBScript payloads; ConnectWise issued an advisory.

Campaigns starting in late August use social engineering, including Quick Assist abuse, to install rogue ScreenConnect clients that spawn wscript.exe and deploy four VBScript files for reconnaissance, staging, and PowerShell execution. The attackers persist via User Run Keys, attempt UAC bypass, install UltraViewer, and propagate the VBScript chain to other connected ScreenConnect endpoints. ConnectWise published an advisory about a file transfer behavior issue affecting cloud and on-premises ScreenConnect, with a CVE identifier and fix expected within a week; it recommends disabling file transfer meanwhile.

SecurityWeek · 9d agoExploit / PoC in the wild

Access Control as Verified Parse Constraints

Researchers verify a class of EverParse validators that correctly enforce access-control policies, deploying a machine-checked enforcement gate on seL4.

The paper targets enforcement-code bugs in commercial security gateways by proving that forward-only, backtrack-free EverParse validators are verified recognizers for a bounded finite-state class that includes access-control decision functions with fixed-offset fields and bounded disjunction. Encoding a bounded policy language into a fixed-size byte buffer allows an SMT solver to verify the enforcement code once, covering all byte values, policies, requests, and sessions. Editing rule content over a fixed endpoint set requires no new proof, while adding endpoints reruns the toolchain. A deployment on the seL4 microkernel ensures every request passes through the gate and unverified components cannot corrupt the enforcement chain.

arXiv cs.CR · 5d agoResearch1

LLM Agents as Computational Typologists

AUTOTYPOLOGIST is an LLM agent that performs evidence-grounded linguistic typology analysis over 25 open-source reference grammars.

The agent retrieves relevant grammar sections, analyzes interlinear glossed text (IGT), and iteratively reasons over typological hypotheses in a ReAct-style workflow. It was evaluated on typological feature coding against expert annotations and hypothesis testing against universals using 25 open-source reference grammars. Results suggest LLM agents can support scalable, inspectable crosslinguistic analysis but still require expert validation.

arXiv cs.AI / cs.LG / cs.CL · 9d agoAI research1

Canonical Color as a Lens into Concept Decodability in Vision Encoders and VLMs

Probing study shows vision encoders make canonical color linearly decodable from grayscale images and tie it to object identity.

Researchers use canonical color as a controlled testbed for measuring conceptual (not just visible) information in vision encoder representations. A dataset of objects with canonical colors was built, and probes on both color and grayscale images show canonical color remains decodable even when color is removed from the input, linked to predicted object identity. Extending to full VLMs, they find post-training has a surprisingly large effect on color decodability in the vision encoder.

arXiv cs.AI / cs.LG / cs.CL · 8d agoAI research1

HyQuant: Hybrid-Precision Quantization for LLM Attention

HyQuant keeps most LLM attention states low-bit while preserving vertical-line tokens and local windows in high precision, maintaining near-lossless accuracy.

HyQuant is a hybrid-precision quantization framework for LLM attention that quantizes most attention states to low bits while keeping accuracy-critical vertical-line tokens and local-window states in full precision, selected via lightweight attention-pattern signals. In the prefill stage it uses a hybrid-precision attention operator, and in the decode stage it applies the same principle to KV-cache compression with fused dequantization and attention computation. Across diverse tasks, models, and datasets it maintains nearly lossless accuracy; code is available on GitHub.

Hugging Face daily papers · 20d agoAI tools & infra1

Det-LIME: Detector-Aware, Multi-Instance Local Interpretable Model-Agnostic Explanations for Automated Marine Mammal Detection

Det-LIME extends LIME to multi-instance object detection explanations, improving attribution for harbor seal aerial surveys.

Det-LIME adapts LIME to object detection by combining per-detection weighting, a proximity kernel emphasizing box-adjacent regions, and IoU-based matching to track instances across perturbations. It was evaluated on aerial drone imagery for harbor seal detection plus a seabird case study, and compared against vanilla LIME, Stabilized LIME, Deterministic LIME, and gradient-based attribution. Using Attribution Ratio and Max Saliency Hit Rate metrics, it consistently improved multi-instance attribution and produced box-aligned explanations useful for debugging and data augmentation.

arXiv cs.AI / cs.LG / cs.CL · 1d agoAI research

Safe Meta-Reinforcement Learning via Information Space Reachability

Safe meta-RL framework reasons about safety in information space, learning a safety value function used for safety filtering and constrained policy optimization.

The paper proposes safe meta-RL that reasons about safety in information space, capturing both physical state and the agent's belief over the underlying task. A safety value function measures the probability of avoiding unsafe regions indefinitely and satisfies a self-consistency condition and Bellman equation, making it learnable via meta-RL. The resulting algorithm uses the learned function for safety filtering and constrained policy optimization, with effectiveness demonstrated on meta-RL benchmarks.

arXiv cs.AI / cs.LG / cs.CL · 2d agoAI research

Unifying Conformal Language Tasks with In-Context Ensembles

Researchers propose Conformal Relevance, which builds conformal score functions via in-context example curation and ensembling to improve conciseness across seven NLP tasks.

The paper targets NLP tasks like summarization and extractive QA that reduce to retrieving content under coverage and conciseness constraints. Conformal Relevance replaces hand-engineered LLM scoring prompts with curated in-context examples and ensembles, maintaining coverage guarantees while improving conciseness with minimal manual input. The authors demonstrate the framework on seven NLP tasks and contribute theory, including a complementarity condition for when ensembling improves worst-case sentence scores and a saturation bound on ensemble gains.

Hugging Face daily papers · 15d agoAI research1

ZDI-26-569: Linux Kernel Net Scheduler True Link Equalizer Race Condition Local Privilege Escalation Vulnerability

ZDI publishes ZDI-26-569, a CVSS 7.5 race condition local privilege escalation in the Linux kernel net scheduler true link equalizer.

The Zero Day Initiative disclosed a race condition in the Linux kernel's net scheduler true link equalizer component enabling local privilege escalation. Exploitation requires the attacker to first run high-privileged code on the target system. The advisory carries a CVSS rating of 7.5; no CVE id is listed in the disclosure text.

ZDI Published Advisories · Aug 13, 2026Advisory

SAS: Simple Attention Sparsification via End-to-End Optimization of Context Ranking

SAS trains attention sparsification end-to-end with the language modeling loss, beating sparse attention baselines especially under tight context budgets.

Simple Attention Sparsification (SAS) injects the selector's continuous scores into attention logits in log form inside the softmax, letting gradients from the language modeling loss directly update the ranking of context units. The method uses normalized softmax gates calibrated against the current block and a memory-efficient Triton kernel integrated into FlashAttention-style computation. Across reasoning, long-context, and agentic tasks, SAS consistently outperforms trainable sparse attention baselines across budgets, with the largest gains under tight attention budgets.

Hugging Face daily papersupdated · 5d agofirst · 6d agoAI research 2 sources1

Beneath the Surface of Chains-of-Thought: A Mechanistic Interpretation of Reasoning Operations in LLMs

Study shows LLM reasoning operations like planning and deduction are geometrically separable in hidden states, with separability peaking in middle layers.

Researchers investigate whether functional reasoning operations — problem formulation, goal decomposition, deduction — have corresponding geometric structure in LLM hidden representations. They find operations are separable in held-out representations with separability peaking in middle layers, ruling out lexical and positional confounds; token-wise operation alignment becomes more distributed across layers, and identical surface tokens are represented differently depending on their surrounding chunk. Attention-masking interventions show chunk-onset operation-aligned representations depend on preceding reasoning context; code is released on GitHub (naver-ai/beneath-cot).

Hugging Face daily papers · 13d agoAI research1

Adversarial Testing of Automated Program Repair Agents for Security Vulnerabilities

SWEADV benchmark shows adversarial issue descriptions make LLM program-repair agents write insecure fixes in 51.7% of cases, evading most detection tools.

Researchers built SWEADV, a benchmark of 750 adversarial issue descriptions derived from 150 SWE-bench Verified repair tasks, covering command execution, deserialization, path traversal, denial of service, and weak hashing attack types. Tested on mini_swe agents backed by GPT-5-Mini, MiniMax-M2.5, and DeepSeek-R, adversarial descriptions induced malicious behavior with successful repair in 51.7% of cases. Detection was weak: LLM-as-judge pre-repair screening reached only 62.3% accuracy, and post-repair detection via static analysis and LLM-as-judge achieved just 39.4% and 55.4%.

arXiv cs.CR · 2d agoAI safety & security2

Likelihood-free inference with nuisance parameters through normalizing flows

Researchers decompose normalizing flows to derive near-pivotal statistics for likelihood-free inference with nuisance parameters, recovering the t-test and beating Welch limits.

A new paper decomposes neural-network normalizing flows to uncover pivotal statistics in the presence of nuisance parameters using only a sample generator from the distribution of interest. The statistic is near-pivotal in the sense of minimum average KL-divergence of its p-values and can incorporate prior knowledge of group invariances such as translation and scale. Experiments show it recovers the one-sample t-test almost exactly, outperforms the Welch test on worst-case size over a constrained variance-ratio range, and delivers higher power and much faster runtime than profile likelihood-ratio techniques on small-to-moderate samples.

arXiv cs.AI / cs.LG / cs.CL · 7d agoAI research

Data Scarcity and Model Sparsity: Mixtures-of-Experts Overfit More to Repeated Data

Study finds Mixture-of-Experts models overfit faster than dense Transformers under repeated training data, with degradation tied to total parameter sparsity.

Across models from 80M to 1B active parameters (8.5B total), MoE architectures degrade more rapidly than dense models when training data is repeated, with the effect increasing with sparsity as dictated by total parameters. Dense 80M models tolerate 8x repetition with minimal loss while MoEs suffer at 4x and underperform dense models beyond 32x. Masking-based regularization such as dropout mitigates overfitting, letting MoEs beat dense models even at over 64x repetition, though no method matches all-unique training data. Routing stabilizes early and expert specialization correlates with overfitting to repeated data.

arXiv cs.AI / cs.LG / cs.CL · 6d agoAI research

Understanding the Impact of Model Pruning on Long-Tail Forgetting and Explanation Reliability in Medical Imaging

Systematic study finds model pruning causes frequency-dependent long-tail forgetting in medical imaging and that gradient-informed methods best preserve explanations.

Across two long-tailed medical imaging datasets, two CNN architectures, four pruning methods, and sparsity up to 95%, the study measures predictive performance, explanation stability, and faithfulness. Rare classes degrade earlier and more severely than frequent ones, while explanation reliability depends mainly on the pruning strategy, with gradient-informed methods degrading least. Mechanistic analysis ties explanation collapse to loss of class-discriminative gradients rather than vanishing feature activations, recommending class- and explanation-aware evaluation of compression.

arXiv cs.AI / cs.LG / cs.CL · 9d agoAI research

[AINews] AEF-1 standard emerges for Third Party Evaluators, as Xai, OpenAI, and Anthropic all cosign

xAI, OpenAI, and Anthropic cosign the AEF-1 third-party evaluation standard while Dario Amodei proposes embedded evaluators for safety verification.

The AI Evaluator Forum published AEF-1, a baseline standard for independent third-party AI evaluations covering access, conflicts of interest, funding relationships, recusal, and transparency, cosigned by xAI, OpenAI, and Anthropic. Dario Amodei wrote a rare personal blogpost proposing embedded evaluators such as METR with desks, badges, company laptops, and internal-risk-team-level access to verify safety commitments, plus democratic and global coordination frameworks. The roundup also covers the pacing debate: Bilal Chughtai left Google DeepMind arguing progress may outrun alignment, while critics including Aidan Gomez and Cohere push back against slowdowns and lab gatekeeping. Additional items include Cline Desktop's launch with open-weight model support.

Latent Space · 1d agoAI safety & security

EvoSafeHarness: Evolving Model- and Domain-Specific Harnesses for Securing Agents

EvoSafeHarness auto-synthesizes per-model, per-domain safety harnesses, cutting prompt-injection attack success on AgentDojo to 0.0% at 82.8% utility.

EvoSafeHarness is an optimization framework that synthesizes deployable safety harnesses for frozen LLM agents in a target domain, jointly searching natural-language policies and executable code logic guided by model behavior, domain specifications, and adversarial review. On DecodingTrust-Agent it reduces average attack success rate from 45.6% to 10.0% at a 3.3-point utility cost, and on AgentDojo reaches 82.8% utility at 0.0% ASR, twice CaMeL's utility at that operating point. It keeps mean ASR below 20% under adaptive PAIR attacks and transfers unchanged to unseen AgentDyn suites. The analysis finds domain semantics determine required safety relations while model and runtime behavior determine enforcement points.

FST Pay: Deterministic Safety-Gated Architecture for Youth Digital Payments

FST Pay proposes a deterministic safety-gated architecture for teen digital payments, pairing invariant authorization checks with decoupled post-settlement AI explanations.

Researchers propose FST Pay, a formal architecture for adolescent digital payments on rails like UPI that applies six deterministic invariant checks (spending limits, guardian co-sign policies, amount thresholds, merchant category codes, temporal intervals, hardware integrity) to classify transactions as ALLOW, REVIEW, or BLOCK. High-risk transactions trigger an asynchronous guardian co-sign workflow. Generative AI is restricted to post-settlement natural-language insights and holds no mutation privileges over the ledger, avoiding non-determinism on the real-time authorization path.

arXiv cs.CR · 6d agoResearch

PrivEscalate: Measuring and Augmenting the Threat of LLM-Automated Linux Privilege Escalation

Researchers release PrivEscalate, a 531-scenario benchmark showing LLM agents' Linux privilege-escalation success varies by vulnerability class, plus PrivEscAgent, a domain-specialized agent that boosts success.

The paper introduces PrivEscalate, an open-source benchmark of 531 Dockerized Linux privilege-escalation scenarios spanning 14 sub-categories, plus 329 parameterized variants measuring sensitivity to environmental distractors. Evaluating six LLMs across three agent architectures shows capability is heterogeneous across vulnerability classes, sensitive to perturbation, and architecture-dependent. The authors also present PrivEscAgent, a wrapper adding deterministic enumeration, category matching, and step planning that outperforms prior privesc-agent baselines without modifying the underlying LLM. The benchmark is released to support LLM agent evaluation, defensive tool validation, and red-team training.

arXiv cs.CR · 8d agoResearch

Deep-Fake CAPTCHA: Mitigating Next-Generation Social Engineering Attacks

Researchers propose DF-CAPTCHA, a challenge-response defense that verifies callers in voice and video to defeat real-time deepfake impersonation in social engineering.

The DF-CAPTCHA framework prompts call participants with simple challenge-response tasks that are easy for humans but hard for real-time deepfake systems to convincingly generate. Responses are verified on four criteria: realism, identity consistency, task completion, and response time. User studies and experiments with real-time deepfake models across audio and video modalities show substantially improved detection over passive artifact-based methods.

arXiv cs.CR · 6d agoResearch1

🔬“We have foundation models for language, not for physics” — Anima Anandkumar, Bren Professor of Computing

Caltech professor Anima Anandkumar discusses Neural Operators and FourCastNet for physics modeling, arguing inductive biases beat pure token scaling.

Anima Anandkumar, Bren Professor at Caltech and co-founder of Accelerated Understanding, describes Fourier Neural Operators that learn in frequency and spherical-harmonic domains to model weather, fusion, and fluid or heat flow. Her team built FourCastNet 3, a global weather model competitive with physics-based simulations that runs on consumer-grade GPUs. She also introduced TorchLean, a framework for writing PyTorch-style networks inside the Lean proof assistant for formal verification, and was appointed to the United Nations Scientific Advisory Board. She argues physical domains resist scaling due to tiny datasets and context lengths in the hundreds of billions, so progress comes from built-in structure and physical priors.

Latent Space · 21d agoAI research1

Google Open-Sources Mantis: A Modular Skills Toolkit That Lets Coding Agents Find, Reproduce and Patch Vulnerabilities

Google open-sourced Mantis, an Apache-2.0 modular skills toolkit that lets AI coding agents find, reproduce, and patch vulnerabilities with sandboxed verification.

Google released Mantis on GitHub under Apache 2.0 as a stack-agnostic set of slash-command skills that chain through the full vulnerability lifecycle: mining version history, building threat models, filtering findings, reproducing bugs in gVisor or network-disabled VMs, assembling exploit chains, patching, and scoring residual risk from 1 to 10. It runs with Gemini CLI, Antigravity CLI, the Google ADK, or comparable agent frameworks, and a supervisor skill (/mantis-meta-agent) can drive the whole loop. Google says the design targets the sub-7 percent true-positive rate of naive AI code scanning, and that its hierarchical summary tree cuts token overhead by over 85 percent. The toolkit is deployable for local and internal evaluation but not yet recommended for production.

MarkTechPost · 7d agoAI tools & infra