ZeroHour

Search: “DoFun”

34 stories

Smart search ranks by meaning as well as keywords (one row per story, last 45 days).

The invisible passenger in your car

Kaspersky discovered Android adware and proxy-botnet malware delivered inside legitimate DoFun car head-unit software.

Kaspersky researchers uncovered Android malware distributed through legitimate software for DoFun automotive head units. The malware displays ads on infotainment screens and enrolls infected devices into a proxy botnet for resale of residential-style proxy traffic. Delivery via trusted vendor software means victims receive it through a legitimate update channel, raising infection likelihood.

Kaspersky Securelist · 26d agoMalware in the wild

Android Car Malware Spreads Through Built

Kaspersky found MoYu Group malware infecting DoFun Android car head units via firmware updaters, enabling ad fraud and proxy botnet operations.

Kaspersky discovered in June 2026 the first documented malware specifically infecting Android-based car head units, spread through the built-in updater (TWCore) of DoFun head unit firmware via a dropper dubbed JarService. The multi-stage implant supports nine commands enabling unwanted ads, ad fraud, and additional module downloads, and installs the zhima reverse proxy module. The campaign is attributed with high confidence to the MoYu Group behind the BADBOX ad fraud and residential proxy scheme; the distribution issue was fixed after responsible disclosure.

The Hacker News · 22d agoMalware in the wild

Hackers infecting Android car systems to build proxy botnet

Kaspersky reports MoYu Group-linked malware infecting DoFun Android car head units, enrolling them in a BadBox-linked proxy botnet for ad fraud and traffic routing.

Kaspersky discovered malware on Android-based head units made by Chinese automotive supplier DoFun, the first documented case of a car head unit being infected through an attack purpose-built for such devices. Attackers abused TWCore, a legitimate DoFun system application that handles updates and can install new apps, to silently push a malicious app called JarService that displays ads, generates fraudulent ad clicks and downloads additional malware. One malware module turns infected head units into reverse proxies so other users' internet traffic can be routed through the car's connection. Kaspersky attributes the campaign with high confidence to MoYu Group, linked to the BadBox operation, which previously infected over 70,000 Android devices and resurfaced as BadBox 2.0 after German authorities disrupted the original botnet in December 2024.

The Record · 23d agoMalware in the wild

Molecular Déjà Vu: Digit-Level Retrieval of Published Values in Frontier Language Models

Audit of 22 frontier models finds widespread verbatim retrieval of published molecular property values, with higher reasoning increasing recall of memorized numbers.

An arXiv audit tests 22 frontier LLMs across 12 molecular regression benchmarks for verbatim retrieval of published values. More than 50% of the LLMs show verbatim retrieval on five datasets, and identical experiments are flagged 89% more often at a high reasoning level than at the lowest one. Suppressing retrieval moves model prediction errors closer together in relative terms, suggesting predictive capability is not determined solely by memorized values.

arXiv cs.AI / cs.LG / cs.CL · 12d agoAI research1

DianShi-RxnDB: A Large-Scale, Fine-Grained Organic Reaction Data Platform Built via a Fully Automated Pipeline for Researchers and AI Agents

Researchers release DianShi-RxnDB, a database of roughly 24 million organic reaction instances extracted automatically from USPTO and EPO patents since 1976.

DianShi-RxnDB is built by a fully automated pipeline integrating patent text, images, and reaction schemes, yielding about 24 million reaction instances, of which 14.8 million (61.7%) pass automated qualification checks. Manual evaluation of 1,300 sampled instances showed 92.95% field-level accuracy, and comparisons with Pistachio found advantages in deduplicated record counts and granularity. The platform offers a web research workbench and a Model Context Protocol (MCP) service enabling AI agents to perform composable structured retrieval.

Hugging Face daily papers · 11d agoAI research

Android car head units infected with proxy botnet malware through built-in software updaters

Kaspersky found malware delivered via car head unit updaters, attributed to the MoYu Group's BADBOX operation, recruiting devices into a proxy botnet.

Kaspersky discovered malware delivered through the built-in TWCore system updater in Android-based car head units running DoFun infotainment firmware, turning devices into ad-fraud tools and nodes in a proxy botnet. The three-stage infection chain (JarService dropper, loader, and final payload supporting nine commands) installs the zhima reverse-proxy module, which Nokia's Deepfield team independently found on TV set-top boxes. Kaspersky attributes the operation with high confidence to the MoYu Group, linked to the BADBOX supply-chain botnet first identified by HUMAN Security in 2023. DoFun closed the gap after Kaspersky's responsible disclosure.

Help Net Security · 23d agoMalware in the wild

Why AI food looks like that

Experts explain why AI-generated food images look unappetizing, citing diffusion model limitations, weak structural reasoning, and stylized training data.

The Verge examines why AI-generated food imagery from restaurants and brands often appears grotesque, citing researchers from Oxford, Naples, Zurich, and London. Diffusion models recover coarse structure before fine texture, so structural errors like extra fingers or donut shrimp get baked in early. Researchers note the models are weak at thin, continuous, terminating structures such as noodles, and reproduce the glossy conventions of professional food photography without understanding the objects. Odd internet imagery and memes in training data further skew outputs toward strange textures and clustered holes.

The Verge · AI · 12d agoAI research

Malware Hijacks Android Car Head Units

Kaspersky reports first known malware infecting Android car head units via firmware updaters, repurposing vehicles as BADBOX proxy nodes for ad fraud.

Kaspersky documented the first known malware infection of Android-based car head units, delivered through the built-in TWCore firmware updater on DoFun devices via an MQTT-driven installation flag. A multi-stage chain installs the JarService dropper and a loader that pulls a clicker and reverse proxy module ('zhima') used for ad fraud and proxy botnet infrastructure. The malware supports nine commands, including clipboard changes, HTTP requests, and JavaScript loading, checking in with C2 every 90 minutes. Kaspersky attributes the campaign with high confidence to MoYu Group, linked to the BADBOX botnet.

Security Affairs · 25d agoMalware in the wild

Evaluating Time-Series Foundation Models and Multimodal Dietary Context for CGM Forecasting

Study finds zero-shot time-series foundation models underperform on CGM forecasting; fine-tuned Chronos-Bolt cuts RMSE up to 18.4% and dietary context adds signal.

The paper evaluates time-series foundation models for continuous glucose monitoring forecasting across eight public datasets covering Type 1 diabetes, Type 2 diabetes, and non-diabetes populations. Under a unified protocol, zero-shot foundation models did not consistently outperform baselines like Elastic Net and PatchTST, but lightweight fine-tuning did, with fine-tuned Chronos-Bolt reducing RMSE by 6.5%-18.4% in the T1D cohort and 8.6%-18.2% in the non-diabetes/T2D cohort. A residual-based fusion framework adding dietary context from CGMacros reduced overall RMSE by about 3% and postprandial RMSE by about 15% versus CGM-only baselines.

arXiv cs.AI / cs.LG / cs.CL · 6d agoAI research

Jalapeño’s first results show industry-leading speed and efficiency in AI inference

OpenAI's Jalapeño custom inference chip shows faster, more power-efficient inference with higher throughput and lower latency.

OpenAI announced first results for Jalapeño, its custom AI inference chip. The company claims industry-leading speed and efficiency, delivering higher throughput and lower latency for modern models compared with existing options. The chip targets more power-efficient serving of large models at scale.

OpenAI News · 22d agoAI industry

Anthropic Says Seven China-Based AI Labs Ran Industrial-Scale Claude Distillation Attacks

Anthropic disrupted industrial-scale unauthorized Claude distillation by seven China-based AI labs, including Alibaba, DeepSeek, Moonshot, and Z.ai.

Anthropic identified and disrupted six illicit distillation campaigns since February 2026 run by seven China-based labs: Alibaba, Moonshot, DeepSeek, Z.ai (Zhipu), MiniMax, Xiaomi, and SenseTime. The largest, GTG-16005, involved 151 million exchanges targeting Claude Opus 4.6/4.7 chain-of-thought transcripts, peaking at roughly 3 million exchanges per day from more than 3,500 fraudulent accounts. Labs used proxy/relay services with fictitious identities, fake or stolen credit cards, harvested API keys, and purchased conversation transcripts from third-party resellers. Anthropic is countering by banning reseller accounts, summarizing internal reasoning before responding, and introducing preserved thinking in Fable 5.1, which encrypts reasoning and prevents context edits before it.

One Symptom, Three Levers: A Critical Review of On-Policy Self-Distillation

A review paper frames on-policy self-distillation collapse as governed by three levers: token weighting, privileged information, and guidance decay.

The paper critically reviews On-Policy Self-Distillation (OPSD), where a language model trains on its own generations scored token-by-token by a teacher conditioned on privileged information such as reference solutions or environment feedback. It identifies collapse, the progressive narrowing of producible reasoning paths, as the dominant failure mode and analyzes it through three levers: signal weighting, the nature of privileged information, and teacher dynamics. The review is restricted to mathematical reasoning, reports no new experiments, and offers a shared vocabulary separating settled findings from disputed ones.

Hugging Face daily papers · 22d agoAI research

Analysis of Smoke Loader in New Tsunami Campaign

Fake Japanese Meteorological Agency tsunami warning emails delivered Smoke Loader and AzoRult malware to steal credentials from targets in Japan.

A fake tsunami warning email impersonating Japan's Meteorological Agency asked recipients to click a link on a registered fake agency domain, delivering the commodity loader Smoke Loader to targets in Japan. Smoke Loader, active since 2011, is modular, and its payloads have included banking trojans, ransomware, cryptominers, password stealers, and PoS malware; the campaign later also deployed AzoRult. New samples add junk-jump obfuscation, encrypted network traffic and payload files, a unique machine ID used for tracking and encryption, and PROPagate injection into explorer.exe, with persistence via a Startup folder shortcut and RC4-encrypted C2 communication.

Palo Alto Unit 42 · Aug 17, 2026Malware in the wild

Banking Trojans: Ursnif Global Distribution Networks Identified

Unit 42 maps banking-trojan distribution networks: spam botnets push Shiotob downloaders and Ursnif, KINS, Tinba at Japan and European targets via compromised web servers.

Unit 42 identified the distribution networks behind banking trojan attacks against Japan, Italy, Spain, Poland, Australia, and Germany. A spam botnet delivered 75 unique Shiotob (Bebloh/URLZone) variants across 7 million spam emails, with Shiotob acting mainly as a downloader that installs Ursnif and the Pushdo spam bot from C2 commands. Over 200 malicious files were hosted on 74 compromised, mostly European small-business web servers between April 2015 and January 2017, with localized invoice and photo-themed email lures per target country.

Palo Alto Unit 42 · Aug 17, 2026Malware in the wild

Meme Coin Factories: Uncovering Large-Scale Manipulations on pump.fun

Large-scale pump.fun study of 15 million meme coins identifies five manipulation classes including wash trading and a Market-Manipulation-as-a-Service ecosystem.

Researchers analyzed all 15 million coins launched on pump.fun over the last two years plus large random samples of transaction data, identifying five manipulation classes: wash trading, creator address obfuscation, coordinated sells, copycat coins, and social media manipulation. Strategic actors bypass the platform interface and implement strategies in a highly automated, low-latency way by interacting directly with the blockchain. The study also uncovers Market-Manipulation-as-a-Service (MMaaS) third-party tools that let non-technical users run these manipulations, and proposes mitigations for traders, pump.fun, and regulators.

arXiv cs.CR · 7d agoResearch

DoppelCart fraud network uses 119,000 fake shops to steal credit cards

DoppelCart, the largest documented fake-shop network, runs 119,000 domains impersonating 44,182 brands to steal payment card details via WebSocket-connected checkout pages.

German cybersecurity startup Nebty discovered DoppelCart, a network of more than 119,000 fake e-commerce domains, mostly in the .SHOP TLD, that harvest payment card details through fraudulent checkout pages. Over 105,000 shops remain active, impersonating 44,182 brands with discounts of up to 65%, and 96% of confirmed shops share identical build files resolving to 27 commerce backends. Checkout code exfiltrates card numbers, expiration dates, CVVs, cardholder names, contact details, and even bank one-time codes to attacker C2 over WebSockets in real time, potentially bypassing bank security controls. The network surpasses BogusBazaar, the previously largest documented fake-shop cluster with 75,000 sites and an estimated 850,000 fraudulent transactions.

BleepingComputer · 8d agoPhishing & fraud

Fly Language Model (FLM) Wires the Full Fruit Fly Connectome Into a Frozen 1.2B LLM, and Its Own Controls Show the Wiring Does Not Help

Researchers wire the full fruit fly connectome (166,700 nodes) into a frozen LiquidAI LFM2.5-1.2B LLM, but controls show no fly-specific benefit.

The Fly Language Model (FLM) couples the complete MaleCNS v1.0 fruit fly connectome (166,700 nodes, 25,582,938 edges) to a frozen LiquidAI LFM2.5-1.2B-Instruct backbone, training only a 278,528-parameter readout (~0.0238% of backbone parameters). The fly readout improved NLL by 0.0222 nats/token (perplexity 3.98 to 3.90) on 32 SmolTalk dialogues, but a direct-input control without the graph beat it in all three seeds. Relabeling node identities removes the gain and the recurrence contracts state differences by 0.6 per token, so the connectome adds no long-range memory. The MIT-licensed code runs locally on Python 3.12, but study artifacts remain private, limiting independent reproducibility.

MarkTechPost · 4d agoAI research1

Anthropic details distillation campaigns from Alibaba, Moonshot AI, and DeepSeek

Anthropic reports nearly 200 million Claude exchanges tied to distillation campaigns by Alibaba, Moonshot AI, and DeepSeek.

A new Anthropic report describes five distillation campaigns totaling nearly 200 million exchanges that extracted chain-of-thought traces from Claude to train competing models, targeting agentic tool use, coding, and reasoning capabilities. The largest campaign, attributed to Alibaba, accounted for 151 million exchanges between May and July 2026 across 3,500 accounts, peaking near three million exchanges per day, allegedly to produce training material for the Qwen model family. A Moonshot AI campaign routed roughly 300,000 requests over ten days through 5,000 accounts, primarily targeting Opus, including one task analyzing CCTV footage that appeared connected to the Chinese military. Attackers used prompt techniques, such as framing queries as katakana-only Japanese translation requests, to make Claude reveal its internal thinking traces.

TechCrunch · AIupdated · 10h agofirst · 6d agoAI safety & security 18 sources2

27.5KB language-agnostic WebGPU syntax highlighter

A developer released gpu-lexer, a 27.5KB language-agnostic syntax highlighter that uses a tiny WebGPU model to label code tokens in the browser.

gpu-lexer splits source into words, whitespace, and symbols, then a small WebGPU model uses local and whole-file context to assign nine token classes, working on languages never seen in training. On held-out files, 12.57% of token labels differ from Shiki, though this measures agreement with Shiki rather than objective correctness. In benchmarks against Shiki 4.4.3, Prism.js, Highlight.js, Sugar High, and Starry Night, it highlighted 10 concatenated copies of three.min.js (5.56M characters) about 10x faster on an Apple M4 Pro in Chrome 152. The author frames it as an experiment, not a grammar-equivalent highlighter.

Show HN: LLM Attention Visualization

A developer released a browser-based tool that visualizes which past tokens influence each LLM output token using aggregated, value-weighted attention scores.

A Show HN project presents a React application built on Transformers.js that renders per-token attention influence by aggregating attention weights scaled by value-vector magnitudes across all attention heads and layers. To expose internal tensors, the author instrumented the ONNX computation graph, hosted a modified model on Hugging Face, and pre-generated prompts to avoid long model downloads in the browser. Demos with a 600-million-parameter model show how verbatim copying draws heavily on source tokens and how single outputs blend information from multiple phrases.

.blend URL Viewer

Simon Willison demos a .blend URL viewer built with GPT-6 Astra in Codex and ChatGPT Images 2.5 generating Blender models.

Simon Willison used ChatGPT Images 2.5 to generate a Faberge egg concept image themed after the TV show Pluribus, then had Codex running GPT-6 Astra (high) execute a Blender local skill to build a 3D model from it. He published the result as a .blend URL viewer tool and continues experimenting with agentic Blender workflows. The post is a hands-on demo of AI-driven creative tooling rather than a security or release announcement.

Simon Willison · 6d agoAI tools & infra

Right Tool, Right Job: Native-Language Evaluation, Tokenizer Sensitivity, and Methodological Findings from a French-Only BabyLM

French BabyLM entry METRON-FR (125M GPT-2, 92.47M words) shows tokenizer artifacts dominate child-scale zero-shot evaluation; proposes standard diagnostics.

METRON-FR is a 125M-parameter GPT-2 pretrained on 92.47M French words, submitted to the BabyLM 2026 Strict track, scoring 85.97% on the native Quebec-French QFrBLiMP benchmark and 62.80% on the BabyLM-weighted leaderboard. A cross-lingual GLUE protocol combining French task-data translation with rank-16 LoRA shows relational tasks gain while world-knowledge tasks regress. Bilingual Lexicon Induction reaches p@1 of 68.84%, 18x above chance, and ablations show single-token zero-shot scoring is dominated by tokenizer and template artifacts at child scale.

arXiv cs.AI / cs.LG / cs.CL · 1d agoAI research

[AINews] Hot Chips: OpenAI’s Jalapeño, Cerebras CS-5, Groq 3 LPX, Apple M6

OpenAI unveiled Jalapeno custom inference chip claiming 1.5-1.9x better perf-per-watt than NVIDIA GB200/GB300, deploying in-house by year-end.

At the 37th Hot Chips conference, OpenAI published first benchmark details for its custom Jalapeno inference chip, claiming 1.5-1.9x more work per watt, 1.7-3.6x lower end-to-end latency, and 2.1-4.1x higher interactive-workload performance versus NVIDIA GB200/GB300, with the 700W-rated part staying at or below 550W in tests. Deployment into OpenAI's own infrastructure begins by year-end, with Gen 2 deep in development and Gen 3 underway. OpenAI also said GPT-Astra and Codex helped write low-level kernels, reportedly 1.5-1.8x faster than human-expert code for selected attention and MoE blocks. Cerebras CS-5, Groq 3 LPX and Apple M6 were also featured at the conference.

Latent Space · 20d agoAI industry

How a researcher uses Codex and ChatGPT to search for new antimicrobial molecules

OpenAI profiles César de la Fuente's lab using ChatGPT and Codex alongside deep-learning models to accelerate antimicrobial molecule discovery.

OpenAI published a case study on bioengineer César de la Fuente's lab, which uses ChatGPT and Codex for hypothesis brainstorming, code writing, dataset processing, and bridging knowledge gaps across biology, chemistry, and computer science. The lab's deep-learning models scan genome and protein databases for antimicrobial peptide candidates, potentially cutting initial searches from years to hours. Bacterial antimicrobial resistance was associated with about five million deaths in 2021, a toll projected to roughly double by 2050.

OpenAI News · 6d agoAI industry1

More than 100,000 fake stores are out to steal your card details

Researchers uncovered DoppelCart, a network of roughly 119,000 cloned fake shops that harvest card details and one-time bank codes during checkout.

Researchers at German firm Nebty identified 118,787 .shop domains tied to cloned online stores, representing 2.72% of the TLD population examined and described as the largest publicly documented fake-shop network by domain count. The shops mimic more than 44,000 brands, advertise discounts up to 65%, and 96% of confirmed shops reportedly share identical build files using just 27 ecommerce backends. Fraudulent checkout pages send card numbers, CVVs, billing data and bank one-time confirmation codes to attacker-controlled servers in real time over WebSockets, allowing criminals to complete payments while victims are still checking out.

Malwarebytes Labs · 7d agoPhishing & fraud

CausalArena: Benchmarking Causal Discovery in the Foundation Model Era

Researchers introduce CausalArena, a unified benchmark revealing that causal discovery rankings shift substantially across structural causal model families and protocols.

The paper presents CausalArena, a unified and evolvable benchmark for causal discovery combining synthetic structural causal models, semantically grounded operational SCMs, formula-grounded scientific SCMs, and public real-world datasets. Experiments across classical, neural, and pretrained causal discovery foundation models show large ranking shifts between benchmark regimes. The authors identify pretraining-evaluation overlap and benchmark diversity as central evaluation challenges.

arXiv cs.AI / cs.LG / cs.CL · 6d agoAI research

ZGCM-1: A Fully Open and Extremely Efficient Foundation Model for Math and Agentic Search

ZGCM-1 is a fully open 7B foundation model with 256K context that stays competitive with frontier models on math reasoning and agentic search.

ZGCM-1 is a fully open 7B dense foundation model trained from scratch using an efficiency-focused recipe: interleaved gated sliding-window and full attention, a stable FP8 Muon optimizer, and MDP-based mid-training with context scaling across 16K, 64K, and 256K. On mathematical reasoning and agentic search suites it remains competitive with much larger frontier models such as Qwen3-235B-A22B and GLM-5.1. The recipe yields a ~4.2x improvement in 16K pre-training time-to-loss, and all weights, checkpoints, training code, data recipes, and W&B logs are open-sourced.

Hugging Face daily papers · 6d agoModel release

Former OpenAI researcher builds an AI model that judges options instead of writing text

TypeSafe AI launches Jev, a judgment-only model built by ex-OpenAI staff that classifies inputs with 70-500 ms latency instead of generating text.

Startup TypeSafe AI, co-founded by former OpenAI researcher and InstructGPT co-author Diogo Almeida, introduced Jev, a model that scores developer-defined answer options with probabilities rather than generating free-form text. The company claims 70-500 ms responses, parallel multi-question evaluation, and $0.042 per million input tokens with free outputs, targeting request routing, sales intent scoring, and assistant guardrail checks. Benchmarks are self-built and not independently verified, the 'no hallucination' guarantee only covers output structure, and access is currently via waitlist.

The Decoder · 8h agoAI industry

Aggah Campaign: Bit.ly, BlogSpot, and Pastebin Used for C2 in Large Scale Campaign

Aggah campaign abuses Bit.ly, BlogSpot, and Pastebin as multi-hop C2 to deliver RevengeRAT across the Middle East, US, Europe, and Asia.

Unit 42 details the Aggah campaign, which began with spearphishing emails in March 2019 spoofing a large financial institution and targeting education, media/marketing, and government organizations in the Middle East, later expanding to the US, Europe, and Asia. Delivery documents use Template Injection to load a remote OLE file whose macro runs mshta against a Bit.ly link redirecting to a BlogSpot post, which then uses Pastebin pastes to download RevengeRAT configured with a duckdns[.]org C2 domain. The embedded script also deletes Microsoft Defender signatures and kills Defender and Office processes, and modifies registry keys to enable macros. High-level TTPs resemble the Gorgon Group, but Unit 42 could not confirm attribution.

Palo Alto Unit 42 · Aug 17, 2026Threat actor

Lightweight Vision Transformer Compression for On-Device Plant Disease Detection in Resource-Constrained Agricultural Field Conditions

A unified ViT compression pipeline (H-BAC pruning, quantization, distillation) cuts plant-disease models 54.5x to 6.01 MB while keeping 95.13% accuracy.

Researchers combined Hessian-Balanced Adaptive Block Pruning (H-BAC), guided by second-order sensitivity estimation, with quantization and attention-based knowledge distillation to compress Vision Transformers for on-device chilli plant disease detection in India. On a 3-class cross-village, cross-device out-of-distribution dataset, the integrated pipeline reduced model size from 327.42 MB to 6.01 MB (54.5x) at 95.13 +/- 2.32% accuracy, matching the 95.13% FP32 baseline. Ablations also show a directly trained 6.01 MB INT8 student reaches 94.87% accuracy, indicating where pruning and distillation add limited value.

arXiv cs.AI / cs.LG / cs.CL · 12d agoAI research

XHToken/Spark-X2.5-4B-GGUF — new model trending #30 on Hugging Face

XHToken released GGUF weights of Spark-X2.5-4B, a compact model with 1M-token context and 200+ language support, under Apache 2.0.

The Hugging Face repository provides BF16 GGUF conversions of Spark-X2.5-4B, a compact general-purpose language model for conversation, writing, translation, reasoning, coding, tool use, and agentic workflows. The model uses a hybrid attention architecture, supports a native context length up to 1M tokens, and covers more than 200 languages. Local inference is supported through Ollama and LM Studio via an XHToken llama.cpp fork, with a --think=false flag to disable thinking mode for faster responses. Released under Apache License 2.0; it was trending #30 on Hugging Face at publication.

Hugging Face trending models · 19d agoModel release

Tracking Elirks Variants in Japan: Similarities to Previous Attacks

Unit 42 links new Elirks backdoor variants attacking Japanese organizations to 2012 Taiwan attacks, delivered via spear-phishing PDFs exploiting Adobe Flash CVE-2011-0611.

Unit 42 analyzed new Elirks backdoor variants found in an attack on a Japanese business, noting strong similarities to 2012 attacks on Taiwanese ministries. The backdoor retrieves its C2 address from attacker-created accounts on Japanese blog and SNS services. Recent deliveries used an airline e-ticket lure named "E-TKT" with a PDF exploiting Adobe Flash CVE-2011-0611. Shared infrastructure and tactics with the Scarlet Mimic campaign suggest possible ongoing cyber espionage across East Asia.

Palo Alto Unit 42 · Aug 17, 2026Threat actor in the wildCVE-2011-0611

Augustinian BabyLM: What Ostensive Definition Can and Cannot Teach a Small Language Model

Study shows visually grounded token embeddings in a small masked LM persist through training and improve object-property knowledge, but escape standard BabyLM benchmarks.

The paper implements ostensive definition for a small DeBERTa masked language model trained on 10M words, seeding visually grounded tokens with embeddings derived from labeled image regions before training. Visual initialization leaves a persistent, seed-replicated advantage on object-property knowledge (COMPS) and a corpus-tailored Visual-Property Swap benchmark covering color, material, size, and shape, but has no effect on most BabyLM grammar benchmarks. Synthetic grounding of previously unseeded words causally transfers the advantage to exactly those words.

arXiv cs.AI / cs.LG / cs.CL · 6d agoAI research

ICE Wants to Know Everyone Who Bought a Certain Green Beanie From REI in the Last 2 Years

DHS subpoenaed REI for all Minneapolis-area customers who bought a specific green beanie since 2024, part of an investigation into 39 ICE protest defendants.

Court filings allege Homeland Security Investigations agents subpoenaed REI in March for transaction records of all persons in the greater Minneapolis–St. Paul area who purchased a specific dark green beanie since 2024. The subpoena was one of 92 sent in a federal case against 39 people, including journalists, who attended an ICE protest at a church. Companies responded differently: T-Mobile handed over six months of a defendant's call and text logs, Google refused a request for YouTube viewers, Reddit withdrew after a First Amendment objection, and Meta pushed back on at least one summons. The 1509 customs summonses require no judicial oversight, and the total number issued under the Trump administration is unknown.

WIRED · Security · 12d agoPolicy & legal