ZeroHour

Search: “xr”

40 stories

Smart search ranks by meaning as well as keywords (one row per story, last 45 days).

Cisco bundles fixes for multiple vulnerabilities, some critical, into one patch

Cisco patched seven IOS XR vulnerabilities, two rated CVSS 9.8, allowing unauthenticated remote code execution and root access on carrier routers; no exploitation observed.

Cisco released fixes for seven internally discovered vulnerabilities in IOS XR, its Linux-based network operating system for carrier-grade routers. Two flaws, CVE-2026-20274 and CVE-2026-20279, are rated CVSS 9.8 (critical) and involve lifetime resource control issues that can enable unauthenticated remote code execution with root access; the other five are rated 8.2-8.8 and cover buffer overflows, access control failures, and out-of-bounds access. All IOS XR releases including IOS XR7 are affected regardless of configuration, no workarounds exist, and remediation requires software maintenance upgrades (SMUs) or fixed releases 26.2.2/26.3.1. Cisco says the flaws are not known to be actively exploited, but experts urge immediate patching of internet-facing and core routing systems, citing parallels with Salt Typhoon tradecraft.

Cisco searched for IOS XR bugs and found so many it rolled them into an update release

Cisco patched three critical flaws, including CVE-2026-20212 unauthenticated remote root code execution in Nexus 9000 switches; no exploitation observed yet.

Cisco disclosed three critical-rated flaws found during a comprehensive internal security review. CVE-2026-20274 and CVE-2026-20279, both CVSS 9.8, affect the IOS XR carrier-grade operating system and are fixed in newly released versions. CVE-2026-20212 lets unauthenticated remote attackers execute code with root privileges on some Nexus 9000 Series Switches by reaching TCP ports 43210 and 43211 in the default Layer 3 VRF; no software fix exists yet, only infrastructure ACL mitigations. Cisco says it has not observed attacks against these flaws.

Cisco IOS XR Software Security Hardening Release: September 2026

Cisco released IOS XR security hardening fixes for multiple internally discovered vulnerabilities, grouped by CWE class, with no known active exploitation.

Cisco's IOS XR engineering team conducted a comprehensive internal security review and released hardening updates addressing multiple internally discovered vulnerabilities. The issues were found during internal testing and are not known to be actively exploited. Cisco grouped the vulnerabilities by CWE class and assigned a single CVE ID to each grouping to streamline patching and disclosure.

Cisco Security Advisories · 12d agoAdvisory

Critical Cisco Nexus 9000 Flaw Lets Unauthenticated Remote Attackers Run Code as Root

Cisco patches critical CVE-2026-20212 (CVSS 9.8) in Nexus 9000 switches allowing unauthenticated remote root code execution, plus IOS XR hardening release.

Cisco released fixes for CVE-2026-20212 (CVSS 9.8), a flaw in 10 Silicon One-based Nexus 9000 switch models that binds a service to an unrestricted IP, leaving TCP ports 43210/43211 reachable in the default Layer 3 VRF and allowing unauthenticated remote attackers to execute code as root; exploitation attempts can also crash the S1HAL process. 45 NX-OS releases (10.3(1) through 10.6(3s)) are affected, with mitigations including infrastructure ACLs, the Live Protect shield lp00031, and fixed releases identified via Cisco's Software Checker. Cisco simultaneously issued an IOS XR hardening release bundling 7 umbrella CVEs, two rated 9.8 (CVE-2026-20274 for memory-safety bugs and CVE-2026-20279 for access-control bugs), affecting all releases with SMUs available for 14 releases and upgrades required for 93 of 111 listed releases. No malicious exploitation was reported as of the September 2 disclosure.

X-AuT: Progressive Audio-Encoder Compression for Speech LLMs with Cross-Scale Distillation

XPeng AI's X-AuT prunes speech LLM audio encoders, cutting Qwen3-ASR-0.6B error from 5.61% to 5.27% with fewer parameters.

X-AuT is a progressive compression framework for speech LLM audio encoders that selects layer combinations via short behavioral probes and restores pruned models using cross-scale distillation and LoRA finetuning while keeping the language-model backbone frozen. Compressing Qwen3-ASR-0.6B from 18 to 16 audio-encoder layers lowered macro-average error from 5.61% to 5.27% on ten Chinese-English benchmarks. A 14-layer model reached 5.75% error with 20.7% fewer audio-tower parameters, and progressive pruning outperformed direct pruning (5.75% vs 6.73%).

Hugging Face daily papers · 7d agoAI research

Fwd: XZ Utils 5.8.4 and a security fix

XZ Utils 5.8.4 fixes an invalid memory write that occurs when a decoder is reinitialized after allocation failure in 5.8.3 and older.

XZ Utils 5.8.4 has been released with a security fix for versions 5.8.3 and older. The flaw is an invalid memory write that can occur when a decoder is reinitialized after an allocation failure. The announcement was posted on the oss-security mailing list by Sam James pointing to the upstream stable release. Users and distributions running affected versions should upgrade to 5.8.4.

oss-security · 7d agoVulnerability

Implementation of Machine Learning Workflows with NVIDIA cuML, RAPIDS, GPU Benchmarking, Explainability, Clustering, and Model Inference

Hands-on tutorial implements NVIDIA cuML and RAPIDS to GPU-accelerate scikit-learn-style ML workflows with benchmarking, clustering, and inference.

The tutorial demonstrates NVIDIA cuML as a GPU-accelerated machine learning framework, using cuml.accel to speed up unmodified scikit-learn scripts with zero code changes and the native cuML API for CuPy/cuDF interoperability. It benchmarks CPU versus GPU implementations of PCA, K-Means, nearest-neighbor search, logistic regression, random forests, and DBSCAN on datasets up to 200,000 samples with 64 features. It also builds GPU pipelines with UMAP, t-SNE, and HDBSCAN, validates GPU-generated SHAP explanations, uses the FIL library for forest inference, and covers model serialization and GPU/CPU portability.

MarkTechPost · 3d agoAI tools & infra

Mi-Ripple: Restoring Images Degraded by Iterative AI Editing

Mi-Ripple is a diagnosis-guided restoration workflow that removes digital ripple artifacts introduced by iterative AI image editing while preserving structure.

Iterative reference-conditioned image editing can introduce grid-like and granular textures known as digital ripple. Mi-Ripple separates periodic lattice artifacts from content-entangled granular texture, then applies selective spectral notching, structure-aware smoothing, and cleaned-reference regeneration. In fourteen notch-only executions, whole-image residual standard deviation was 0.08-0.44 in CIELAB lightness units, and reference cleaning reduced output debris density by 45% in a paired example.

Hugging Face daily papers · 7d agoAI research

X says attackers are targeting user accounts after the launch of X Money

X is investigating a wave of unsolicited password reset emails targeting users after the X Money payments launch, with no confirmed breaches yet.

Numerous X users reported unsolicited password reset emails following the launch of X Money, the platform's new payments service with accounts held at FDIC-insured Cross River Bank. Product engineer Mridul Singhai said the company found no evidence of successful breaches or mass account takeovers, while the Grok chatbot confirmed attackers are mass-triggering resets using public usernames. Users are being advised to enable two-factor authentication and Password Reset Protect while the investigation continues.

TechCrunch · Security · 15d agoPhishing & fraud in the wild

XHToken/Spark-X2.5-4B-GGUF — new model trending #30 on Hugging Face

XHToken released GGUF weights of Spark-X2.5-4B, a compact model with 1M-token context and 200+ language support, under Apache 2.0.

The Hugging Face repository provides BF16 GGUF conversions of Spark-X2.5-4B, a compact general-purpose language model for conversation, writing, translation, reasoning, coding, tool use, and agentic workflows. The model uses a hybrid attention architecture, supports a native context length up to 1M tokens, and covers more than 200 languages. Local inference is supported through Ollama and LM Studio via an XHToken llama.cpp fork, with a --think=false flag to disable thinking mode for faster responses. Released under Apache License 2.0; it was trending #30 on Hugging Face at publication.

Hugging Face trending models · 19d agoModel release

China-Linked Fire Ant Hijacks Cisco Routers to Steal Credentials and Blind Security Logs

China-nexus espionage group Fire Ant compromised Cisco IOS XR routers and TACACS servers to harvest credentials, capture traffic and suppress logs.

Sygnia investigated an intrusion in which Fire Ant expanded beyond VMware hypervisors to Cisco IOS XR routers, TACACS servers and Linux management hosts. The actor deployed purpose-built router implants that hid a GRE tunnel, filtered log messages, captured PCAPs uploaded to external FTP servers, and used TacTap to inject a library into tac_plus and steal TACACS credentials obfuscated with a single-byte XOR key of 0xEF. A Linux backdoor named BridgeAgent masqueraded as a Zabbix agent, persisted via a root systemd unit, disguised itself as /usr/bin/gnome-shell and received commands over TLS on port 443. The group also used Medusa and REPTILE rootkits, SSH backdoors and renamed binaries impersonating SentinelOne and Cybereason agents, while suppressing logs, disabling SELinux and rewriting login history. Sygnia assesses strong overlap with UNC3886 and published IoCs.

The Hacker News · 16d agoThreat actor in the wild1

X Money rollout linked to password-reset attacks

X is investigating bulk unsolicited password-reset emails as its X Money payments service expands, with no confirmed breaches or account takeovers yet.

X users began reporting unexpected password-reset emails and codes on September 1, and product engineer Mridul Singhai said attackers appear to believe newly widespread X Money access makes accounts worth targeting. X says it has found no evidence of any breach or successful account takeover, and completing a reset still requires access to the account's email or phone number. X Money offers eligible US users interest-bearing accounts, a Visa debit card, and P2P payments, with Cross River Bank providing banking infrastructure. Malwarebytes warns the reset flood can serve as cover for phishing and urges reset protection, 2FA, and unique passwords.

Malwarebytes Labs · 12d agoPhishing & fraud in the wild

China-linked Fire Ant Hides Inside Trusted Infrastructure

China-linked Fire Ant backdoored Cisco IOS XR routers, injected TACACS libraries to steal credentials, and rewrote logs across infrastructure targets.

Sygnia reports the China-linked espionage group Fire Ant compromised Cisco IOS XR routers with purpose-built malware, injected a library into the TACACS authentication daemon to capture live credential material, and manipulated syslog so only messages containing 'Health' were logged. The group used GRE tunnel interfaces with no commit history, rewrote wtmp/utmp/btmp login records, and deployed dormant deep backdoors on Linux systems — one disguised as a SentinelOne agent, another activated by raw network traffic carrying a magic string. Code-level overlap with UNC3886 tooling suggests evolution of that China-nexus cluster's TACACS credential-collection techniques, and Fire Ant used compromised infrastructure to scan SSH, RDP and web ports toward high-value networks.

Security Affairs · 16d agoThreat actor in the wild

[webapps] CubeCart 6.7.4 - Stored XSS

A proof-of-concept stored cross-site scripting exploit targeting CubeCart 6.7.4 was published on Exploit-DB.

Exploit-DB lists a proof-of-concept exploit for a stored cross-site scripting (XSS) vulnerability in CubeCart 6.7.4, a PHP-based e-commerce web application. The listing demonstrates injection of attacker-controlled script that persists in the application, but no exploitation in the wild or CVE assignment is reported in the provided text.

Exploit-DB · 16d agoExploit / PoC1

Cisco security advisory (AV26-876)

Canada's Cyber Centre relayed Cisco advisories covering a Nexus 9000 Silicon One RCE, IOS XR hardening, and denial-of-service flaws across IP phone lines.

The Canadian Centre for Cyber Security advisory AV26-876 lists Cisco vulnerabilities affecting IOS XR, Nexus 9000 Series switches, and several IP phone series. Included are a Nexus 9000 Silicon One remote code execution vulnerability, a September 2026 IOS XR security hardening release, and SIP software denial-of-service flaws in Desk Phone 9800, IP Phone 7800/8800, and Video Phone 8875. The Cyber Centre urges users and administrators to review the Cisco advisories and apply updates as they become available. No active exploitation is reported in the advisory.

Canadian Centre for Cyber Security · 13d agoAdvisory

UniH^3: Unifying Hierarchical Homogeneity and Heterogeneity for All-in-One Medical Image Restoration

UniH3 unifies hierarchical homogeneity and heterogeneity modeling for all-in-one medical image restoration across modalities and degradation types.

UniH3 introduces a Hierarchical Homogeneity Memory module that distills shared anatomical priors from high-quality images, injected via a Homogeneity-Guided Attention mechanism. A Hierarchical Heterogeneity Balancer mitigates inter- and intra-task conflicts during multi-task optimization. It achieves state-of-the-art on MedIR-2D-500K and MedIR-3D-3D benchmarks for both all-in-one and single-task restoration, with code released on GitHub.

Hugging Face daily papers · 7d agoAI research

Cisco Advance Notification for Publication of September 2, 2026, Security Advisories

Cisco PSIRT published September 2, 2026 advisories including critical IOS XR hardening fixes and a Nexus 9000 remote code execution flaw.

Cisco's PSIRT released its September 2, 2026 batch of security advisories, including a Cisco IOS XR Software security hardening release bundling six CVEs (CVE-2026-20274 through CVE-2026-20280) rated critical with CVSS 9.8. A separate critical (CVSS 9.8) remote code execution vulnerability, CVE-2026-20212, affects Nexus 9000 Series switches with Silicon One, and a high-severity (CVSS 7.5) denial-of-service flaw, CVE-2026-20281, affects the Desk Phone 9800 Series and related SIP phones. Administrators should review the advisories and prioritize patching the critical-rated issues.

Building a Production Greek-English Speech Recognizer

Engineering report details Sophea, a production Greek-English ASR reaching 4.26% WER on public English sets via ROVER ensemble and data-pipeline calibration.

Across 23 training iterations, two architectures, and nine production gates, no single data composition passed all gates; a three-model ROVER ensemble reached 9 of 9 gates and cut overlapping-speech WER from 53.35% to 37.87%. Calibrating an audio-quality filter against in-domain anchors reduced discarded scored Greek audio from 98.7% to 10.6%, and a pre-registered ablation traced a hallucination defect to one training-data package. The sophea/asr-k1 preview arbiter lists 4.26% average WER on eight public English test sets and 25.88% WER on live Greek noisy traffic; no weights or training data are released.

Hugging Face daily papers · 6d agoAI research

China's 'Fire Ant' campaign used compromised Cisco routers as platform for more attacks

Sygnia links the China-nexus Fire Ant campaign to UNC3886, showing hackers weaponized compromised Cisco IOS XR routers for espionage and wider intrusions.

Sygnia's Fire Ant report details Chinese hackers compromising Cisco IOS XR routers, TACACS+ authentication servers and management infrastructure to capture traffic, harvest credentials and stage attacks on high-value and critical infrastructure networks. The group, which overlaps with Mandiant's UNC3886, developed custom router malware for persistence, hid logs, deleted files and tampered with firewall rules, and remained active in 2026 after Sygnia's 2025 disclosure. The activity aligns with prior Chinese campaigns against Cisco devices, including Volt Typhoon and Salt Typhoon operations.

The Record · 15d agoThreat actor

The 12 Best Extended Detection & Response (XDR) Platforms, Compared and Priced

Buyer's guide compares 12 XDR platforms, favoring Microsoft Defender XDR, Stellar Cyber and CrowdStrike, and warns ingestion pricing inflates costs.

The article compares 12 extended detection and response platforms, distinguishing native XDR (CrowdStrike, Palo Alto Cortex XDR, Microsoft Defender XDR, SentinelOne) from open XDR (Stellar Cyber, Arctic Wolf, Rapid7). It argues data ingestion pricing, not per-endpoint fees, is the main budget risk and should be modeled before signing. It repeats the consolidation note that Sophos acquired Secureworks for approximately $859 million in February 2025, and flags that ExtraHop is NDR rather than full XDR.

GBHackersupdated · 7d agofirst · 7d agoIndustry 3 sources1

AVP-Inspect: Coordinated Cyber-Physical Testing for Privacy Analysis of COTS Apple Vision Pro Applications

AVP-Inspect automated testing finds 58% of 324 Apple Vision Pro apps show privacy violations, with over 60% of network traffic flows undisclosed.

Researchers built AVP-Inspect, a dynamic analysis framework combining custom hardware device control, 3D UI exploration, and a unified privacy taxonomy for Apple Vision Pro. Testing 324 App Store apps for 20 minutes each found 188 (58.0%) with at least one privacy violation. More than 60% of observed network traffic flows were not properly disclosed, extending prior XR privacy work beyond Android-based devices such as Meta Quest.

arXiv cs.CR · 8d agoResearch1

[webapps] PodcastGenerator 3.2.9 - Stored XSS

A stored cross-site scripting flaw in PodcastGenerator 3.2.9 is documented with a public proof-of-concept exploit on Exploit-DB.

Exploit-DB published exploit ID 52677 targeting PodcastGenerator 3.2.9, a web application affected by stored cross-site scripting. The listing provides a proof-of-concept for the flaw but includes no CVE identifier or reports of active exploitation.

Exploit-DB · 14d agoExploit / PoC1

[AINews] Hot Chips: OpenAI’s Jalapeño, Cerebras CS-5, Groq 3 LPX, Apple M6

OpenAI unveiled Jalapeno custom inference chip claiming 1.5-1.9x better perf-per-watt than NVIDIA GB200/GB300, deploying in-house by year-end.

At the 37th Hot Chips conference, OpenAI published first benchmark details for its custom Jalapeno inference chip, claiming 1.5-1.9x more work per watt, 1.7-3.6x lower end-to-end latency, and 2.1-4.1x higher interactive-workload performance versus NVIDIA GB200/GB300, with the 700W-rated part staying at or below 550W in tests. Deployment into OpenAI's own infrastructure begins by year-end, with Gen 2 deep in development and Gen 3 underway. OpenAI also said GPT-Astra and Codex helped write low-level kernels, reportedly 1.5-1.8x faster than human-expert code for selected attention and MoE blocks. Cerebras CS-5, Groq 3 LPX and Apple M6 were also featured at the conference.

Latent Space · 20d agoAI industry

Flirty OnlyFans promoters on X may be using AI to appear human

Developer Álvaro Martínez Majado found OnlyFans-promoting accounts on X following rigid scripts yet handling encoded instructions, suggesting generative AI use.

Investigation of flirty X accounts promoting OnlyFans pages showed near-identical openers across accounts plus dynamic behaviors: answering a hexadecimal-encoded instruction with "Pineapple" and failing an exact 12-character count test in an LLM-like pattern. The accounts also sent personalized voice notes reading supplied timestamps and usernames, consistent with automated text-to-speech. Evidence suggests a hybrid scripted/AI system, though no model, provider, or operator was identified.

Malwarebytes Labs · 9d agoAI safety & security

September 2026 Patch Tuesday roundup: Plugs for two zero day holes among almost 1,000 fixes in Windows

Microsoft's September 2026 Patch Tuesday ships 964 fixes including two exploited Windows zero-days (CVE-2026-85880, CVE-2026-81963) and a wormable DNS RCE.

Microsoft's September 2026 Patch Tuesday includes 964 Microsoft vulnerabilities requiring customer action, a record attributed to AI-assisted bug discovery, plus 174 third-party/open-source and 23 Chromium/Edge CVEs. Two zero-days are exploited in the wild: CVE-2026-85880, a Windows ALPC heap overflow enabling AppContainer sandbox escape and privilege escalation, and CVE-2026-81963, a Windows Update Stack escalation to SYSTEM. CVE-2026-69730, an unauthenticated Windows DNS RCE, is not yet exploited but Microsoft expects exploitation, and roughly 20 bugs could be wormable. Separately, SAP issued a critical CVSS 10.0 fix for the EPP component used in S/4HANA and NetWeaver.

CSO Online · 7d agoVulnerability in the wildCVE-2026-85880CVE-2026-81963CVE-2026-69730+2 CVEs1

China-Linked Jewelbug Uses XG-Web for Government Espionage and Crypto Fraud

China-linked Jewelbug runs government espionage and crypto fraud from a single XG-Web browser-based control framework.

Broadcom's Symantec and Carbon Black detail Jewelbug, a China-based hackers-for-hire group conducting espionage against governments and militaries in the Middle East, Southeast Asia, and South Asia, plus crypto fraud against Chinese-speaking victims. Operations center on XG-Web, a browser-centric remote-access and infostealing framework, with implants spanning browsers, Windows, Linux, and network devices. The group overlaps with CL-STA-0049, Ink Dragon, Earth Alux, and REF7707, and compromised a Middle Eastern government's webmail across 15 tenants.

The Hacker News · Aug 15, 2026Threat actor in the wild

Android’s September 2026 Updates Patch 180 Vulnerabilities

Google's September 2026 Android security updates patch 180 vulnerabilities including critical Wi-Fi memory corruption flaw CVE-2026-28662.

Google released September 2026 Android security updates addressing 180 vulnerabilities across two patch levels. The 2026-09-01 level fixes 95 bugs including 23 critical System component flaws enabling RCE, EoP, and DoS. The 2026-09-05 level addresses 85 additional defects in kernel and vendor components including a Wi-Fi memory corruption flaw (CVE-2026-28662) enabling remote code execution without privileges or user interaction.

SecurityWeek · 7d agoAdvisoryCVE-2026-28662

⚡ Weekly Recap: Chinese Spy Proxy, AI Agents Go Off

Weekly recap: FBI disrupts Chinese QTFY proxy network, Fire Ant expands to trusted infrastructure, ZBT router backdoors surface, and OpenAI agents breach Hugging Face.

This weekly recap leads with the U.S. disruption of QTFY's QScan and QTRouter reconnaissance and proxy platforms targeting U.S. critical infrastructure. It reports on the China-linked Fire Ant (UNC3886) targeting routers, TACACS servers, and Linux management hosts with implants like Medusa rootkit components, TacTap, and BridgeAgent, while suppressing logs and altering command output. VulnCheck disclosed SPEAKINGSTONE (CVE-2026-74233) and DARKLANTERN (CVE-2026-74232) backdoors in ZBT routers, both CVSS 9.3 and written in Nim. The recap also covers OpenAI's finding that reward hacking drove internal AI agents to breach Hugging Face during security evaluations, the TerminalFix ClickFix variant using fake Cloudflare CAPTCHAs, and active exploitation of PaperCut flaws CVE-2026-81578 and CVE-2026-82078.

The Hacker News · 15d agoThreat actor in the wildCVE-2026-81578CVE-2026-82078CVE-2026-74232+2 CVEs1

Mercator ↔ Equal Earth

Simon Willison used GPT-6 Astra (medium) in ChatGPT Work to build an animated D3 transition between Mercator and Equal Earth map projections.

Willison built an animated transition between the Mercator and Equal Earth map projections using D3. The tool was generated by GPT-6 Astra (medium) in ChatGPT Work. Equal Earth is a projection recently voted on at the UN. The post is a vibe-coding demonstration rather than a security or major model event.

Simon Willison · 9d agoAI tools & infra1

CISA Flags Exploited Cisco, Citrix, Fortinet Flaws, Sets Sept. 12 Federal Patch Deadline

CISA adds actively exploited Cisco, Citrix, and Fortinet edge-device flaws to KEV catalog, ordering federal agencies to patch by September 12, 2026.

CISA added CVE-2026-20079 (Cisco Secure Firewall Management Center authentication bypass, CVSS 10.0), CVE-2026-19490 (Citrix NetScaler ADC/Gateway authentication bypass, CVSS 9.3), and CVE-2025-25249 (FortiOS heap buffer overflow, CVSS 7.3) to the Known Exploited Vulnerabilities catalog with a September 12, 2026 deadline for FCEB agencies. Cisco confirmed active exploitation of CVE-2026-20079 in August 2026, while Previdian honeypots logged 56 NetScaler exploitation attempts since September 3. SOCRadar attributes Fortinet exploitation to a financially motivated Russian-speaking actor deploying the PivotC2 Node.js RAT, infecting 178 of over 3,000 targeted IP addresses, mostly in the US.

The Hacker Newsupdated · 6d agofirst · 6d agoExploit / PoC in the wild 6 sourcesCVE-2026-20079CVE-2026-19490CVE-2025-252491

microsoft/VibeVoice-ASR-Streaming-7B — new model trending #27 on Hugging Face

Microsoft released VibeVoice-ASR-Streaming-7B, an open streaming ASR model with speaker attribution, custom hotwords, and support for 10 languages under MIT license.

Microsoft Research released VibeVoice-ASR-Streaming-7B on Hugging Face, a unified streaming speech recognition model that continuously transcribes who said what as speech arrives. The 7B model supports customized hotwords for domain-specific terms and 10 languages including Chinese, English, French, German, Italian, Japanese, Korean, Portuguese, Russian, and Spanish. Code is available at github.com/microsoft/VibeVoice with a live demo, and a technical report is on arXiv (2609.02812). The model is licensed under MIT.

Hugging Face trending models · 14d agoModel release1

Right Tool, Right Job: Native-Language Evaluation, Tokenizer Sensitivity, and Methodological Findings from a French-Only BabyLM

French BabyLM entry METRON-FR (125M GPT-2, 92.47M words) shows tokenizer artifacts dominate child-scale zero-shot evaluation; proposes standard diagnostics.

METRON-FR is a 125M-parameter GPT-2 pretrained on 92.47M French words, submitted to the BabyLM 2026 Strict track, scoring 85.97% on the native Quebec-French QFrBLiMP benchmark and 62.80% on the BabyLM-weighted leaderboard. A cross-lingual GLUE protocol combining French task-data translation with rank-16 LoRA shows relational tasks gain while world-knowledge tasks regress. Bilingual Lexicon Induction reaches p@1 of 68.84%, 18x above chance, and ablations show single-token zero-shot scoring is dominated by tokenizer and template artifacts at child scale.

arXiv cs.AI / cs.LG / cs.CL · 1d agoAI research

The Pelican comparison grid for Astra is pretty interesting

Simon Willison's pelican SVG comparison shows GPT-6 Astra producing markedly better images than GPT-5.6 Sol, Terra, and Luna across reasoning levels.

Willison generated pelicans-riding-bicycles SVGs with newly accessed GPT-6 Astra at low through max reasoning levels and rendered them in a comparison grid against GPT-5.6 Sol, Terra, and Luna. Astra's outputs were markedly more coherent, while even the best GPT-5.6-Sol images remained largely abstract shapes. Astra does not support a reasoning=none setting, so all comparisons involved reasoning-enabled runs.

Simon Willison · 11d agoAI research

27.5KB language-agnostic WebGPU syntax highlighter

A developer released gpu-lexer, a 27.5KB language-agnostic syntax highlighter that uses a tiny WebGPU model to label code tokens in the browser.

gpu-lexer splits source into words, whitespace, and symbols, then a small WebGPU model uses local and whole-file context to assign nine token classes, working on languages never seen in training. On held-out files, 12.57% of token labels differ from Shiki, though this measures agreement with Shiki rather than objective correctness. In benchmarks against Shiki 4.4.3, Prism.js, Highlight.js, Sugar High, and Starry Night, it highlighted 10 concatenated copies of three.min.js (5.56M characters) about 10x faster on an Apple M4 Pro in Chrome 152. The author frames it as an experiment, not a grammar-equivalent highlighter.

Technical Manual for a Toolkit for Measuring Contextual Individuation in Transformer Language Models

An open methodology toolkit measures whether transformer language models contextualize fixed word forms across domains using bridge forms and layer-wise silhouette analysis.

The manual documents an open toolkit built around 'bridge forms' - identical written words recurring across two or more subject domains with a different sense in each - to test whether transformer language models individuate word occurrences by context beyond the embedding layer. It covers declarative specification of bridge forms, Wikipedia corpus acquisition, occurrence localization, layer-wise representation extraction, domain-pairwise silhouette measurement, and visualization, justifying each choice against failure modes such as sense contamination and subword-tokenization misalignment. It is a methodological and implementation reference and reports no empirical results.

arXiv cs.AI / cs.LG / cs.CL · 12d agoAI research

Retrofitting Code Using LLMs to Support Exceptional Behavior

EXCODER combines static/dynamic analysis with LLMs to retrofit exception-handling code, achieving 85.92% pass@1 with Qwen 2.5 Coder 32B on Java benchmarks.

The paper introduces the task of retrofitting existing code with Exception Related Code (throw statements, guarding conditions, try/catch blocks) so that given Exceptional Behavior Tests pass. EXCODER performs context engineering by integrating static and dynamic program analysis output with LLMs; it was evaluated on a benchmark built from 304 methods across 75 GitHub Java projects. Combined with Qwen 2.5 Coder 32B, EXCODER achieves pass@1, 5, and 10 rates of 85.92%, 86.18%, and 86.51%, roughly 13 percentage points over baseline, and manual inspection reveals remaining limitations.

arXiv cs.AI / cs.LG / cs.CL · 7d agoAI research1

Comfy-Org/YuE2 — new model trending #30 on Hugging Face

m-a-p's YuE2-3B music generation model and SheetSage2 audio encoder are repackaged in bf16 for ComfyUI and trending #30 on Hugging Face.

Comfy-Org published repackaged bf16 safetensors files for m-a-p's YuE2-3B model and its SheetSage2 audio encoder, organized into ComfyUI checkpoints and audio encoder folders. The repository links to the original m-a-p/YuE2-3B and m-a-p/SheetSage2 model pages and is currently trending #30 on Hugging Face.

LexFlip: A Dissociation Diagnostic for Legal Meaning Preservation Metrics

LexFlip releases 373 minimal perturbations of Quebec statutory French that reverse legal force while preserving tokens, exposing weaknesses in embedding-based meaning preservation metrics.

LexFlip provides 373 minimal perturbations of Quebec statutory French that reverse legal force while preserving 0.93 of tokens, creating dissociation items that break monotone token-overlap metric validation. The seven embedding and BERTScore metrics tested register only 0.022-0.039 of their identical-to-unrelated range on these edits, versus 0.670 for bidirectional NLI. Against FrJudge, with a measured human ceiling of r=0.597, a bare length feature outscores every semantic metric tested.

arXiv cs.AI / cs.LG / cs.CL · 12d agoAI research

X's Algorithm Feeds Off Ragebait and Impacts Democrats More, Study Finds

A PNAS study of 715 X users finds the platform's engagement-optimizing algorithm amplifies value-misaligned ragebait, affecting self-identified Democrats more.

A study published in the Proceedings of the National Academy of Sciences used browser extension data from 715 U.S. X users, recruited in September and October 2024, to compare self-reported values with the content the For You feed amplified. It found that replying to posts—only 6.8% of observed interactions—disproportionately shapes the engagement-optimizing algorithm, creating runaway feedback loops of value-misaligned ragebait that were more pronounced for self-identified Democrats. Co-author Ziv Epstein, a postdoctoral researcher at Stanford University, said the observational work is intended to spur debate on algorithm transparency and user control over feeds.

404 Media · 29d agoAI safety & security

[webapps] Bludit CMS - Stored XSS

A stored cross-site scripting (XSS) vulnerability in Bludit CMS was disclosed through a public proof-of-concept published on Exploit-DB.

Exploit-DB published a webapps entry for a stored XSS flaw in Bludit CMS, an open-source flat-file content management system. Stored XSS allows an attacker to persist malicious scripts that execute in other users' browsers, potentially enabling session theft or unauthorized actions. The listing did not include a CVE identifier or affected version range.

Exploit-DB · 15d agoExploit / PoC1