Smart search ranks by meaning as well as keywords (one row per story, last 45 days).
JarvisGUI: Towards Cross-Device GUI Agents with Dynamic Task Composition
JarvisGUI benchmark tests GUI agents on cross-device workflows across Android, Windows, and Ubuntu, revealing major gaps in state transfer and long-horizon reasoning.
JarvisGUI is a dynamic benchmark that formulates GUI tasks as input-output transformations under a lightweight type system, automatically composing multi-step cross-device workflows across Android, Windows, and Ubuntu virtual environments. Evaluation shows state-of-the-art open-source GUI agents struggle with state-transfer awareness, cross-platform contextual reasoning, and long-horizon dependency management, exposing a capability gap invisible to existing single-device benchmarks.
Top 10 Best Mobile Threat Defense (MTD) Solutions in 2026
Roundup of 2026 mobile threat defense tools recommends Zimperium and Lookout for targeted-attack detection and Defender for Endpoint for Microsoft shops.
This guide ranks ten mobile threat defense solutions, recommending Zimperium and Lookout for on-device detection against targeted users such as executives and journalists, and Microsoft Defender for Endpoint mobile for organizations already licensing Microsoft 365 E5. It explains that MDM enforces configuration while MTD detects attacks, and that mobile phishing now arrives via SMS, messaging apps and QR codes rather than email. It also highlights mercenary spyware and zero-click exploits as shifting requirements for high-risk users, referencing Apple's threat-notification program and Lockdown Mode.
Can We Stop The Ads? Taxonomy and Characterization of Smartphone Splash Ads and Existing Countermeasures
Study of 108 ad-defense implementations finds only one tool blocked splash-ad navigation across ten popular apps, and it required Accessibility permission.
The paper taxonomizes smartphone splash ads — full-screen ads at app launch that trick users into trigger mechanisms such as moving the phone — and analyzes 108 documented advertising defenses for deployment barriers. Many defenses require device rooting, jailbreaking, runtime code injection, or application modification; others need extra permissions, rule maintenance, compilation, or payment. In evaluating 13 configurations of 11 tools across 10 popular apps, only one prevented ad-triggered navigation across all ten apps, requiring Accessibility permission and leaving ads visible roughly one second before dismissal. Documented harms include delayed emergency response, driver distraction, and degraded accessibility for vision-impaired users.
Hackers infecting Android car systems to build proxy botnet
Kaspersky reports MoYu Group-linked malware infecting DoFun Android car head units, enrolling them in a BadBox-linked proxy botnet for ad fraud and traffic routing.
Kaspersky discovered malware on Android-based head units made by Chinese automotive supplier DoFun, the first documented case of a car head unit being infected through an attack purpose-built for such devices. Attackers abused TWCore, a legitimate DoFun system application that handles updates and can install new apps, to silently push a malicious app called JarService that displays ads, generates fraudulent ad clicks and downloads additional malware. One malware module turns infected head units into reverse proxies so other users' internet traffic can be routed through the car's connection. Kaspersky attributes the campaign with high confidence to MoYu Group, linked to the BadBox operation, which previously infected over 70,000 Android devices and resurfaced as BadBox 2.0 after German authorities disrupted the original botnet in December 2024.
Studying Image Tokenizers as Visual Languages in Unified Multimodal Models
A controlled autoregressive testbed shows validation losses must be analyzed per task, and image tokenizer choice affects joint multimodal text modeling.
Researchers built a pure-autoregressive testbed to study image tokenizers as the 'visual language' of unified multimodal models, tracking task-specific validation losses during multimodal continual pretraining across text, image, text-to-image (T2I), and image-to-text (I2T) prediction. They found that losses exhibit distinct scaling behavior per task and rank tokenizers differently, and that I2T loss over a shared text vocabulary gives a more consistent loss–performance signal than T2I loss. Better reconstruction does not necessarily yield lower task-specific losses or stronger downstream performance, and tokenizer choice can affect text modeling under joint optimization. Case studies examine the discriminator, semantic supervision, and vocabulary size design axes.
Top 10 Best Mobile Device Management (MDM) Solutions in 2026
A 2026 MDM buyer guide ranks ten solutions, recommending Microsoft Intune for Microsoft 365 estates and Jamf for Apple-only environments.
A 2026 buyer guide evaluates ten mobile device management solutions, leading with Microsoft Intune as the default for Microsoft 365 organizations and Jamf for Apple estates. It recommends choosing the enrolment model before selecting a vendor and clarifying BYOD visibility to prevent privacy disputes. Kandji, Mosyle, Omnissa Workspace ONE, ManageEngine, Scalefusion, and Hexnode are covered as alternatives. Guidance ties MDM to Zero Trust data access policies via Apple User Enrolment and Android work profiles.
Studying Image Tokenizers as Visual Languages in Unified Multimodal Models
A controlled pure-autoregressive testbed shows task-specific validation losses rank image tokenizers differently, with I2T loss the most consistent signal.
Researchers built a controlled pure-autoregressive testbed and tracked task-specific validation losses during multimodal continual pretraining across text, image, text-to-image (T2I), and image-to-text (I2T) prediction. They find losses should be analyzed per task because they exhibit distinct scaling behavior and rank tokenizers differently, and that the loss-performance relationship depends on the predicted token space. I2T loss, computed over a shared text vocabulary, correlates consistently with both generation and visual understanding performance after supervised finetuning. Case studies revisit the discriminator, semantic supervision, and vocabulary size as tokenizer design axes.
UK Government Begins Moving 23 Million Users Away From Passwords
UK government rolls out passkeys for GOV.UK One Login, giving 23 million users phishing-resistant passwordless access to public services.
The UK government has begun deploying passkeys across GOV.UK One Login for more than 23 million users, replacing passwords and SMS one-time codes with FIDO2 cryptographic credentials. A trial saw over 300,000 people adopt passkeys, and nearly one in ten daily authentications already use them, cutting SMS verification costs by almost £600 per day. Passkeys remain optional, with password-based sign-in retained as a fallback, and the NCSC endorses the approach as phishing-resistant.
CanvasAnneal: Curriculum Reinforcement Learning for Diffusion Language Models
CanvasAnneal injects teacher reasoning traces into diffusion canvases during curriculum RL, improving diffusion LLMs on MATH500, Countdown, and Tau2.
CanvasAnneal is a curriculum-guided reinforcement learning framework for diffusion language models that addresses exploration bottlenecks in standard RL. It warm-starts exploration by injecting teacher-generated reasoning traces into the initial diffusion canvas, then gradually removes this guidance so the model generates reasoning trajectories independently. Across mathematical reasoning and tool-use benchmarks, it improves over standard diffu-GRPO on MATH500, Countdown, and Tau2 and accelerates reward improvement, though gains are task-dependent.
Apple Rolls Out Massive Security Update Fixing 273 Vulnerabilities Across Its Devices
Apple's coordinated rollout patches 273 unique vulnerabilities across iOS 27, macOS Golden Gate 27, watchOS and Safari, including remote code execution flaws.
Apple shipped one of its largest coordinated security updates on September 14, 2026, fixing 273 unique CVEs across iOS 27, iPadOS 27, macOS Golden Gate 27, watchOS 27, tvOS 27, visionOS 27, Safari 27 and Xcode 27. Highlights include CVE-2026-65414, a Bluetooth out-of-bounds write enabling remote code execution, and CVE-2026-84607, an AVEVideoEncoder race condition granting kernel privileges to sandboxed apps. macOS Golden Gate 27 covers the broadest set with 210 CVEs, and Apple states none of the flaws were exploited in the wild.
Bring your spreadsheet data to life with Sheets canvas
Google added Sheets canvas to Workspace, letting users generate interactive dashboards, trackers, and charts from spreadsheet data via prompts.
Google introduced Sheets canvas, a Workspace feature that turns spreadsheet data into interactive dashboards, custom study trackers, and seating charts from a simple prompt. The announcement is a productivity feature launch with no described security impact or risk. No metrics, model details, or pricing are provided in the notice.
Free streaming boxes may be routing criminal traffic through your home
Researchers found SuperBox streaming boxes and the CyberFlix TV app enroll home connections into the Popanet residential proxy network, routing criminal traffic.
Researchers found that SuperBox devices and the CyberFlix TV app, distributed through SuperBox's custom app store, contain Popanet proxy functionality that registers the device with servers controlled by the proxy operator, enrolling household connections into residential proxy networks. The reported configuration weakens Android safeguards with exposed ADB access, root-level privileges without authentication, and removal of app-install protections. Plume's research warns these proxy networks can also function as malware-delivery platforms, and the FBI notes foreign entities use residential proxies to conceal activity such as credential stuffing and account abuse. Malwarebytes advises disconnecting and replacing affected SuperBox/CyberFlix devices rather than factory-resetting them.
Omni Interaction Agent Technical Report
Researchers release Gander, an end-to-end omni interaction model with full-duplex streaming across video, speech, and text plus agentic capabilities.
Gander is an end-to-end model unifying omni perception, realtime interaction, and agentic capabilities in a single framework, accepting continuously streaming video, speech, and text. It uses a Cerebellum-Brain architecture where the Cerebellum handles realtime conversation and the Brain handles reasoning and agentic tasks, built on a streaming Thinker-Talker design with chunk-level token streams. Internal human evaluations report spoken dialogue on par with SOTA open source models and competitive omni interaction; the models, code, and data are released publicly.
Schools are catching on to Big Tech’s playbook
A new book warns AI firms are repeating Big Tech's education playbook, as New York City and Los Angeles restrict classroom AI use.
NYT education reporter Natasha Singer's book 'Coding Kids' documents how Apple, Microsoft and Google embedded proprietary curricula and Chromebooks in US schools over 15 years, building product loyalty and market position. Google's Chromebook and Classroom dominance positioned it to promote generative AI in classrooms. New York City banned AI in elementary and middle schools and Los Angeles imposed broader restrictions including high schoolers, as parents and teachers push back against screens and AI in classrooms.
Infostealers Are Hijacking Claude Sessions and Draining Subscriptions
Infostealers hijacked active Claude sessions to bypass 2FA, drain paid usage and run unauthorized charges; Anthropic is revoking sessions and refunding.
Anthropic confirmed that infostealer malware on user machines stole active Claude login sessions, letting attackers bypass passwords, MFA, and SSO and drain paid usage and run unauthorized charges. Affected families identified on Windows include Vidar, LummaC2, StealC, RedLine, and Acreed, plus Atomic Stealer on a small number of Macs; the malware typically arrived via unofficial downloads or malicious apps. Anthropic signed users out, removed saved payment cards, revoked affected sessions, and is refunding unauthorized charges, noting phones and tablets were not involved.
5 new ways to level up your learning with Search
Google promotes five Search study tools for students preparing for classes and standardized tests.
Google published a back-to-school post highlighting five ways to use Search tools when studying for classes and standardized tests. The post is promotional product guidance rather than a launch, research result, or security announcement.
Android 0-day Vulnerability on Google Pixel Devices Actively Exploited in Attacks
Google patched CVE-2026-58704, an actively exploited Android zero-day allowing proximal privilege escalation via the Pixel cellular modem, urging the 2026-09-05 patch.
Google confirmed CVE-2026-58704, a high-severity elevation-of-privilege flaw in the Pixel cellular modem, is being exploited in limited, targeted attacks and shipped emergency fixes in the September 2026 Pixel Update Bulletin. The low-complexity bug requires no user interaction and enables proximal/adjacent privilege escalation with no additional execution privileges, phrasing Google has historically used for spyware-vendor and state-aligned zero-days. The Pixel bulletin patches 110 flaws including 12 critical RCEs, while the broader September Android update addressed roughly 180 vulnerabilities, including Wi-Fi memory-corruption bug CVE-2026-28662.
Attention Quantization for Tabular Foundation Models
FP8 quantization of attention queries, keys, and values speeds tabular foundation model inference up to 1.7x with no accuracy loss.
The paper develops an FP8 quantization strategy targeting attention calculations (queries, keys, values) in tabular foundation models, arguing attention matters more than weight or KV cache quantization given their differing size and serving patterns versus LLMs. Aligning quantization error between test rows and training rows proves crucial, since misalignment causes drastic accuracy drops. A Triton kernel using explicit FP8 matrix multiplication achieves up to 1.7x speedup over regular 16-bit kernels, with no relevant accuracy loss on TabPFN-v3 and TabICLv2 across TabArena and BeyondArena benchmarks.
Anthropic locks out Claude users after infostealers hijack login sessions
Anthropic invalidated Claude sessions compromised by infostealers such as Vidar, Lumma and Atomic Stealer, which steal browser cookies to bypass 2FA.
Anthropic began locking users out of Claude accounts after infostealer malware stole browser session cookies, letting attackers replay logins and bypass two-factor authentication. Identified malware includes Vidar, Lumma (LummaC2), StealC, RedLine and Acreed on Windows and Atomic Stealer (AMOS) on a small number of Macs. Anthropic signed affected users out, removed saved payment methods and refunded unauthorized charges. One affected user traced the infection to a pirated game downloaded from a Russian underground forum; the company advised victims to remove the malware before resetting passwords and re-enabling 2FA.
AMIE, our research medical AI system, demonstrates real-time clinical video consultation capabilities in a first-of-its-kind study.
Google's AMIE research medical AI system demonstrates real-time clinical video consultations in a first-of-its-kind simulated study.
Google introduced AMIE, its research medical AI system, demonstrating real-time clinical video consultation capabilities in a first-of-its-kind study. The evaluation was conducted in simulated settings, extending the AMIE diagnostic dialogue research line to multimodal video consultations. AMIE remains a research system rather than a deployed clinical product.
Introducing ChatGPT for Teens: Built for learning, backed by protections
OpenAI launched ChatGPT for Teens, a teen-focused version with built-in protections, healthy-use features, and expanded parental controls.
OpenAI introduced ChatGPT for Teens, a version of its assistant positioned for learning and critical thinking with stronger built-in protections. The offering includes healthy-use features and additional controls for parents. It is a consumer product launch aimed at safer teen usage rather than a model or research release.
The 12 Best Unified Endpoint Management (UEM) Solutions, Compared and Priced
A buyer's guide compares 12 unified endpoint management platforms, recommending Microsoft Intune, Jamf and Omnissa for common scenarios.
The article compares 12 UEM solutions including Microsoft Intune, Omnissa Workspace ONE, Jamf, ManageEngine, Ivanti, SOTI and 42Gears, with pricing models and platform coverage. It flags ownership changes such as Workspace ONE becoming Omnissa, BlackBerry divesting Cylance to Arctic Wolf, and Citrix's status under Cloud Software Group. Guidance centers on checking existing Microsoft 365 licensing before purchasing, per-user versus per-device pricing, and combining platforms like Intune and Jamf for Apple estates.
UC Berkeley Researchers Release CUA-Lite, an Open Platform Unifying Sandboxes, Data, Evaluation and RL for Computer-Use Agents
UC Berkeley's CUA-Lite is an open platform unifying computer-use agent sandboxes, datasets, evaluation and RL; Lite.OSWorld cuts OSWorld memory 4.1 GB to 0.9 GB.
UC Berkeley researchers released CUA-Lite, an open platform placing agents, environments, traces, and training for computer-use agents behind one action space, one LiteSample schema, and one command across desktop, browser, and mobile. Lite.OSWorld reproduces the OSWorld task suite and evaluators in plain Docker containers (0.9 GB RAM vs 4.1 GB, cold start 23.8s, ~4.6× more parallel instances), with scores matching the QEMU/KVM VM across 13 models. The platform claims 30k+ verifiable tasks, 15+ benchmarks, 10+ agents, and 20+ datasets on Hugging Face including Aguvis, OpenCUA, and ScaleCUA. A documented SFT run lifts Qwen3-VL-2B-Instruct mean episode return from 0.138 to 0.237 on the 332-task lite.osworld split.
Top 10 Best Unified Endpoint Management (UEM) Solutions in 2026
A 2026 buyer's guide ranks UEM platforms, recommending Intune for Microsoft 365 shops, Jamf for Apple estates, and SOTI for rugged devices.
The guide ranks ten unified endpoint management platforms for 2026, recommending Microsoft Intune for Microsoft 365 organizations, Jamf for Apple-heavy estates, and SOTI for rugged, kiosk, and industrial devices. It notes VMware Workspace ONE now operates as Omnissa after Broadcom divested the End-User Computing division, and that BlackBerry sold Cylance to Arctic Wolf in February 2025 while retaining BlackBerry UEM. The article provides a coverage checklist spanning Windows, macOS, iOS, Android, Linux, kiosks, legacy on-prem Windows, and wearables/IoT.
Does Your Agent's Memory Survive a Model Upgrade? A Controlled Study of Memory Portability
A controlled study finds agent memory portability varies sharply: fixed-schema knowledge graphs survive model swaps while compressed notes degrade.
The study compares preserving an agent's history as raw long context, RAG chunks, compressed natural-language notes, or fixed-schema knowledge graphs across model upgrades, using 48 synthetic histories and two open-weight sub-10B-parameter models. Fixed-schema KG accuracy changed by only +0.0004 ± 0.0020 after a writer swap, while compressed NOTES shifted asymmetrically by +9.91 or -13.28 percentage points depending on migration direction. Mixed 50/50 embedding migrations captured only 4.96 of an 11.90-point RAG re-embedding gain; 80% of the NOTES deficit came from information lost at construction, and 81% of the RAG deficit from retrieval failures. Store-only repair of NOTES failed to reach 90% recovery in all 48 cases, while retaining raw histories enabled recovery in 34 of 48 for one direction.
FLAT: Resampling Image and Text into 1D Flexible-Length Aligned Transmodal Tokens for Retrieval and Generation
FLAT jointly trains a multimodal encoder with text-to-image and image-to-text decoders, producing flexible-length tokens that hit 83.1 GenEval on T2I after fine-tuning.
FLAT (Flexible-Length Aligned Transmodal representations) is a pre-training framework that jointly optimizes a shared multimodal encoder with T2I and I2T decoders, combining contrastive alignment with bidirectional cross-modal generative objectives. It maps visual and textual inputs into a unified continuous 1D sequence space and uses nested dropout over prefix-K tokens for dynamic output lengths. A single pre-training stage supports cross-modal retrieval and generation (71.1 GenEval), with task-specific fine-tuning reaching 83.1 GenEval on T2I, 40.5 BLEU-4 and 138.6 CIDEr on MS-COCO captioning, and strong Recall@5 on MS-COCO and Flickr30K.
How Hugging Face Inference Endpoints, Jobs, and Buckets Power Search on Papers with Code
Hugging Face explains how Inference Endpoints, Jobs, and Buckets power semantic search on Papers with Code.
Hugging Face describes the infrastructure behind search on Papers with Code, built on its Inference Endpoints, Jobs, and Buckets services. The post is a product-focused engineering walkthrough with no security impact.
Quoting Laurie Voss
Laurie Voss argues AI collapses code-writing and review costs, leaving product discovery and precise definition as the core of software engineering.
Simon Willison quotes Laurie Voss's essay "We are all Product Engineers now," which argues that AI is collapsing the cost of writing code and will likewise collapse the cost of reviewing, fixing, and operating it. Voss contends the remaining work is finding out what people want, defining it precisely, and making software pleasant to use. He expects the amount of software to grow without limit because demand has no ceiling, making product-definition skills the whole job. No specific models, tools, or incidents are named; this is career and industry commentary.
Verifiable by Construction: Claim-Level Evaluation of Verbatim Citation in Clinical Question Answering
Evaluation of twelve LLMs on 222 clinical questions shows verbatim quotes rarely substantiate claims; claude-opus-5 fully substantiates only 37.1%.
The authors build a standardized harness over four clinical practice guidelines and evaluate twelve LLMs on 222 synthetic clinical questions, measuring citation attachment, verbatim quote production, and claim substantiation. Most models attach verbatim quotes to over 90% of claims from prompting alone, though lightweight models like claude-haiku-4.5 struggle. Quotes frequently fail to substantiate claims: claude-opus-5 quotes 98.0% of claims but fully substantiates only 37.1%, exposing a capability gap for verifiable clinical QA.
Codex bundles LibreOffice
OpenAI's Codex desktop app bundles 1.7GB of runtimes including full Python, Node.js, Poppler, git, and LibreOffice binaries.
Blogger Simon Willison found that the OpenAI Codex desktop app (since rebranded to ChatGPT) keeps about 1.7GB in a ~/.cache/codex-runtimes/codex-primary-runtime folder, including full Python and Node.js installations plus native binaries for Poppler, git, and the LibreOffice office suite. Bundled skills in the plugins directory instruct Codex on how to find and use these binaries. The observation highlights the heavyweight local runtime stack shipped with agentic coding tools.
Cartesian – AI 3D Modeling for Design
Formas launches Cartesian, an AI 3D modeling tool for architecture and product design with natural-language editing and CAD exports.
Cartesian by Formas is an AI-powered 3D modeling tool aimed at architecture and product design that converts photos, rough plans, sketches, and scans into precise geometry. Users can edit models conversationally while explicitly preserved elements stay unchanged. It creates real solids and NURBS geometry and exports to AutoCAD DWG, Rhino 3DM, and SketchUp SKP, with BIM IFC support planned, removing the need for a separate desktop CAD license.
Learning never stops: How AI makes learning continuous
OpenAI report describes how students and educators use ChatGPT to extend learning continuously beyond the classroom.
OpenAI published a report examining how students and educators use ChatGPT to make learning more continuous. The report describes support that extends beyond the classroom, positioning ChatGPT as an ongoing learning companion. The release is part of OpenAI's education-focused communications rather than a technical or safety research paper.
Multi-Vector (Late Interaction) Embedding Models with Sentence Transformers
Hugging Face details building and using multi-vector late-interaction embedding models with Sentence Transformers for retrieval workloads.
Hugging Face published a guide on multi-vector, late-interaction embedding models (ColBERT-style) supported through Sentence Transformers. The post covers how practitioners can build and use these models for retrieval and RAG pipelines. It is a developer tooling and technique write-up, not a security advisory.
Technical Manual for a Toolkit for Measuring Contextual Individuation in Transformer Language Models
An open methodology toolkit measures whether transformer language models contextualize fixed word forms across domains using bridge forms and layer-wise silhouette analysis.
The manual documents an open toolkit built around 'bridge forms' - identical written words recurring across two or more subject domains with a different sense in each - to test whether transformer language models individuate word occurrences by context beyond the embedding layer. It covers declarative specification of bridge forms, Wikipedia corpus acquisition, occurrence localization, layer-wise representation extraction, domain-pairwise silhouette measurement, and visualization, justifying each choice against failure modes such as sense contamination and subword-tokenization misalignment. It is a methodological and implementation reference and reports no empirical results.
Build Low-Latency Multilingual Voice Agents: Open Weights & Full Deployment Control with NVIDIA Magpie TTS
NVIDIA's Magpie TTS open-weight multilingual speech model enables low-latency voice agents with full deployment control.
Hugging Face's blog highlights NVIDIA Magpie TTS, an open-weights multilingual text-to-speech model designed for building low-latency voice agents. The open licensing gives developers full deployment control, allowing self-hosted multilingual speech for agentic applications. The post walks through building voice agents with the model.