ZeroHour

Search: “earbuds”

22 stories

Smart search ranks by meaning as well as keywords (one row per story, last 45 days).

Skullcandy Dime 3 Bluetooth Flaw Lets Nearby Attackers Hijack Audio and Microphone

CERT/CC disclosed VU#859658: Skullcandy Dime 3 earbuds on firmware 1.0.0.28 accept unauthenticated Bluetooth pairing, letting nearby attackers hijack audio and microphone.

Skullcandy Dime 3 wireless earbuds (model S2DCW, firmware 1.0.0.28) accept Bluetooth Classic BR/EDR pairing requests from unknown devices without the owner activating pairing mode, a flaw linked to CVE-2025-20701 in Airoha Bluetooth audio SDK implementations and tracked as VU#859658 by CERT/CC. Attackers within Bluetooth range who know the device address can bond via the NoInputNoOutput configuration, establish A2DP or HFP/HSP connections, disrupt the owner's active audio session, and potentially capture live microphone audio. Firmware 1.0.0.30 addresses the issue, but Dime 3 earbuds do not support firmware updates through the Skullcandy mobile app, leaving affected users without a known upgrade path.

GBHackersupdated · 5d agofirst · 6d agoVulnerability 3 sourcesCVE-2025-207011

Apple Doesn’t Want You to Worry About the New Apple Watch's Listening Features

Apple Watch Series 12 and Ultra 4 add opt-in audio intelligence features that process microphone audio on-device via a new Secure Exclave.

The Apple Watch Series 12 and Ultra 4 ship with four opt-in audio intelligence features: Sound Recognition, Shazam music identification, Siri Recap conversation summaries, and Live Rewind 15-second transcription. Audio is held and processed in an isolated Secure Exclave buffer on the new S11 chips, with on-device speech recognition on iPhone producing a distilled transcript that foundation models in Private Cloud Compute then summarize. Apple says no raw audio is stored or accessible to the operating system, apps, the user, or Apple, and untransferred audio is automatically deleted.

WIRED · Securityupdated · 5d agofirst · 6d agoAI industry 7 sources1

VU#859658: Skullcandy Dime 3 wireless earbuds contain an unauthenticated Bluetooth pairing vulnerability

Skullcandy Dime 3 earbuds (CVE-2025-20701) accept Bluetooth pairings without owner consent, letting in-range attackers hijack audio or capture microphone; no firmware update path exists.

CERT/CC's VU#859658 describes CVE-2025-20701 in the Airoha Bluetooth audio SDK, present in Skullcandy Dime 3 (Model S2DCW) firmware 1.0.0.28. A direct Bluetooth Classic pairing request with no PIN or physical confirmation completes via NoInputNoOutput, adding the attacker's device as trusted. Attackers in radio range can hijack the A2DP audio session, access the Hands-Free/Headset profile, and capture live microphone audio. Firmware 1.0.0.30 contains the effective patch, but Skullcandy says the Dime 3 does not support app-based firmware updates, leaving existing units unpatchable.

OpenAI buys smartphone camera maker Glass Imaging for $300 million, report says

OpenAI acquired smartphone camera startup Glass Imaging for over $300 million, deepening rumored hardware ambitions after its $6.5 billion io deal.

OpenAI bought Glass Imaging, a Los Altos startup founded in 2019 by former Apple engineers Ziv Attar and Tom Bishop who led Apple's Portrait Mode team, per a Wall Street Journal report. Glass Imaging had raised about $30 million and uses neural networks that learn individual camera systems to improve image quality at capture time. The acquisition fuels speculation about OpenAI hardware plans including smartphones, earbuds, and AI companion devices, following its $6.5 billion purchase of Jony Ive's io startup in 2025.

TechCrunch · AI · 1d agoAI industry

Inaudible sounds used to fingerprint browsers catch AliExpress red-handed

AliExpress was caught using inaudible audio signals to fingerprint visitors' browsers, an outdated but still-active cross-session tracking technique.

Ars Technica reports that AliExpress fingerprinted site visitors by sending inaudible sounds to their browsers, exposing a legacy audio-based fingerprinting technique in commercial use. The method allows persistent tracking despite the technique being widely considered outdated. The finding highlights that legacy tracking methods remain deployed on major e-commerce properties, raising privacy concerns for defenders and privacy teams.

Ars Technica · Security · 22d agoResearch

Comfy-Org/YuE2 — new model trending #30 on Hugging Face

m-a-p's YuE2-3B music generation model and SheetSage2 audio encoder are repackaged in bf16 for ComfyUI and trending #30 on Hugging Face.

Comfy-Org published repackaged bf16 safetensors files for m-a-p's YuE2-3B model and its SheetSage2 audio encoder, organized into ComfyUI checkpoints and audio encoder folders. The repository links to the original m-a-p/YuE2-3B and m-a-p/SheetSage2 model pages and is currently trending #30 on Hugging Face.

Injected and Leaked: Actively Inducing Side-Channel Leakage Using Electromagnetic Injection and Hardware Nonlinearity

Researchers introduce InjectEave, using electromagnetic injection and hardware nonlinearity to induce side-channel leakage and eavesdrop on headphone audio from 30 meters.

An arXiv paper shows electromagnetic injection can actively amplify side-channel leakage: nonlinear hardware such as amplifiers, ADCs, and power converters modulates secret electrical signals onto an injected EM carrier, upconverting low-frequency secrets into measurable EM emissions. By tuning injection frequency and amplitude, an adversary can shape the effective spectrum and entropy of the resulting leakage. The InjectEave attack demonstrated eavesdropping on wired and wireless headphone audio from up to 30 meters and in through-wall scenarios using accessible RF equipment, plus leakage of smart home device power consumption and analog sensor inputs. Case studies show closed-loop eavesdropping and manipulation of landline phone conversations, and the paper discusses mitigations.

arXiv cs.CR · 12d agoResearch

A Vinyl Bar in Shibuya is a startup from a former Spotify leader for making music apps

Former Spotify innovation head raises $5.5M pre-seed for A Vinyl Bar in Shibuya, a startup building playful music-creation apps with selective generative AI features.

A Vinyl Bar in Shibuya, founded by former Spotify head of innovation Máuhan M Zonoozy, raised a $5.5M pre-seed round from Mantis VC, SV Angel, Boxgroup, Quiet Capital and others. The startup ships small music-play apps including Speed Surfer, Usersound, Stacks, Drops, Sampler, and the iOS mixer app bop, plus a new prompt-based sound creation feature. Zonoozy says the company deliberately avoids infusing AI into every product, arguing human taste and participation become more valuable as AI-generated content grows abundant.

TechCrunch · AI · 2d agoAI industry 2 sources

Researchers Show How Meta's 'Pervert Glasses' Are Used to Harass Women

University of Sydney researchers detail how pickup artists use Meta Ray-Ban smart glasses to covertly film and harass women, then post the videos on Instagram.

Researchers Joanne Gray, Milica Stilinovic, Marcus Carter, and Ben Egliston analyzed 350 Instagram videos posted between September 2023 and March 2026 showing unsolicited approaches to women filmed with smart glasses. They found a clear correlation between covert filming and harassment severity, arguing ambient capture creates 'borderline' harassment that evades platform moderation mechanisms. Instagram head Adam Mosseri said the platform would remove harassing pickup-line content, though similar videos remain widespread a month later. Meta's safeguards, such as the recording light, were previously criticized as insufficient, and users have modded glasses to disable the light.

404 Media · Aug 12, 2026AI safety & security

WearableQA: A Benchmark for Health Reasoning over Real-World Wearable Data

WearableQA benchmark tests LLM health reasoning over longitudinal wearable data; the best of 14 evaluated LLMs reaches 72.9% accuracy.

WearableQA comprises 4,084 ten-option multiple-choice questions built from wearable time series, blood biomarkers, and demographics of 200 real users with up to 500 days of daily measurements. It defines 16 question types along two axes: data versus health reasoning, and single- versus cross-signal reasoning. Evaluation of 14 proprietary and open-source LLMs shows performance from 19.6% to 72.9% against a 10% chance baseline, with most models below 60%.

arXiv cs.AI / cs.LG / cs.CL · 12d agoAI research

WearableQA: A Benchmark for Health Reasoning over Real-World Wearable Data

WearableQA benchmark introduces 4,084 questions over real longitudinal wearable data, showing 14 LLMs score 19.6-72.9% on health reasoning, far from solved.

WearableQA is a benchmark of 4,084 10-option multiple-choice questions built from wearable time series, blood biomarkers, and demographics of 200 real users with up to 500 days of daily measurements. It defines 16 question types along two axes: data versus health reasoning, and single- versus cross-signal reasoning, using a dual-grounding framework combining literature and population-validated patterns. Evaluations of 14 proprietary and open-source LLMs show accuracy ranging from 19.6% to 72.9% against a 10% chance baseline, with most models below 60%.

Hugging Face daily papers · 12d agoAI research

LG smart TVs caught logging audio with screen off and snooping on local devices

Gamers Nexus found LG smart TVs record microphone audio in standby, scan home networks, and feed LG Ad Solutions ad targeting.

A 135-minute Gamers Nexus investigation with Level1Techs and independent researchers found retail LG OLED TVs running webOS sweep local networks, gather device names and Wi-Fi metadata, and run Automated Content Recognition. Tests showed the TVs capture clean microphone audio while appearing powered down and store it offline, uploading once reconnected. The team also found RCE vulnerabilities in webOS now moving through responsible disclosure; LG claims 216 million smart TV sales, and its ad unit claims access to 363 million addressable devices in the US.

Roland is getting into generative AI music with Melody Flip

Roland launched Melody Flip, a DAW plugin that generates genre-themed MIDI loops rather than complete songs like Suno or Udio.

Roland's Melody Flip is a DAW plug-in offering around 250 genre-based 'Palettes' for generating melodies, chord progressions, basslines, and drum patterns, either from scratch or derived from a reference track. Users can control genre, note density, BPM, and key but cannot use text prompts, and outputs are simple loops with General MIDI-style tones intended for MIDI export into a DAW. The launch follows well-received hardware releases like the SH-4d, Gaia 2, and TR-1000, though music-community sentiment toward generative AI may limit goodwill.

The Verge · AI · 12d agoAI industry

BreezeBlue/Breeze-TTS-2 — new model trending #19 on Hugging Face

BreezeBlue open-weights Breeze TTS 2, a bilingual text-to-speech model it ranks #1 among open-weight models on the Artificial Analysis TTS leaderboard.

BreezeBlue released open weights and Apache 2.0-licensed PyTorch inference code for Breeze TTS 2 on 2026-08-25. The text-to-speech model supports English and Chinese, voice cloning, reference-free voice design, voice direction, and inline vocal events like (laugh) and (sigh). Reported performance includes #1 open-weight ranking on the Artificial Analysis Elo leaderboard, under 40 ms time-to-first-audio, a 0.32 real-time factor on an NVIDIA H100, and about 7.7 GiB GPU memory for eager inference.

Hugging Face trending models · 22d agoModel release

X-AuT: Progressive Audio-Encoder Compression for Speech LLMs with Cross-Scale Distillation

XPeng AI's X-AuT prunes speech LLM audio encoders, cutting Qwen3-ASR-0.6B error from 5.61% to 5.27% with fewer parameters.

X-AuT is a progressive compression framework for speech LLM audio encoders that selects layer combinations via short behavioral probes and restores pruned models using cross-scale distillation and LoRA finetuning while keeping the language-model backbone frozen. Compressing Qwen3-ASR-0.6B from 18 to 16 audio-encoder layers lowered macro-average error from 5.61% to 5.27% on ten Chinese-English benchmarks. A 14-layer model reached 5.75% error with 20.7% fewer audio-tower parameters, and progressive pruning outperformed direct pruning (5.75% vs 6.73%).

Hugging Face daily papers · 6d agoAI research

Elevenlabs makes Music v2.5 available via app and API with free and pro tier options

ElevenLabs releases Music v2.5 via app and API, claiming fuller, more natural songs preferred over v2 in blind listening tests.

ElevenLabs launched Music v2.5 for ElevenMusic, reporting that listeners preferred it in a blind test across 47,885 comparison pairs, especially for R&B, Soul, Hip-Hop, Rock, and orchestral tracks. The free tier offers five lossless downloads per day and Pro includes 400 per month, with attribution required on free and commercial use restricted by industry. Tracks based on other artists' songs are blocked, and v2 remains available alongside the API. ElevenLabs says existing Music models were trained on 'licensed stems and music,' distinguishing it from competitor Suno, which faces lawsuits for training on copyrighted content; a Universal Music Group licensing deal covers only future separate products.

The Decoder · 3d agoAI industry

m-a-p/YuE2-3B — new model trending #30 on Hugging Face

M-A-P released YuE2-3B, an open music generation model that outperforms Suno v5 on WildSongBench and runs locally on a 24GB GPU.

The M-A-P (multimodal-art-projection) team released YuE2-3B, an open-weights music generation model that turns lyrics and a style prompt into full songs with vocals and accompaniment. It uses an AR-NAR Mixture-of-Transformers backbone with symbolic planning and flow matching through a VAE, and supports editable scores (melody and chords, including ABC notation) plus agentic editing workflows. On 192 WildSongBench prompts it reports a SongBench average of 6.9632 (best-of-8) versus 6.8721 for Suno v5, claimed as state of the art among evaluated open and proprietary models. It runs 48 kHz stereo inference locally on a single 24GB NVIDIA GPU without quantization, with companion releases including YuE2-Vae, MERT-v2 encoders, the WildSongBench dataset, and SheetSage2.

Hugging Face trending models · 7d agoModel release1

microsoft/VibeVoice-ASR-Streaming-7B — new model trending #27 on Hugging Face

Microsoft released VibeVoice-ASR-Streaming-7B, an open streaming ASR model with speaker attribution, custom hotwords, and support for 10 languages under MIT license.

Microsoft Research released VibeVoice-ASR-Streaming-7B on Hugging Face, a unified streaming speech recognition model that continuously transcribes who said what as speech arrives. The 7B model supports customized hotwords for domain-specific terms and 10 languages including Chinese, English, French, German, Italian, Japanese, Korean, Portuguese, Russian, and Spanish. Code is available at github.com/microsoft/VibeVoice with a live demo, and a technical report is on arXiv (2609.02812). The model is licensed under MIT.

Hugging Face trending models · 14d agoModel release1

ZDI-26-589: BlueZ A2DP Stack-based Buffer Overflow Remote Code Execution Vulnerability

ZDI details a network-adjacent stack buffer overflow in BlueZ's A2DP stack (CVE-2026-19774, CVSS 7.1) allowing remote code execution after pairing a malicious Bluetooth device.

The Zero Day Initiative published advisory ZDI-26-589 for a stack-based buffer overflow in BlueZ, the Linux Bluetooth protocol stack. A network-adjacent attacker who can pair a malicious Bluetooth device with the target can execute arbitrary code on the affected installation. The flaw carries a CVSS 7.1 rating and is tracked as CVE-2026-19774.

Threats Making WAVs - Incident Response to a Cryptomining Attack

Guardicore researchers dissect a cryptomining attack that hid a cryptominer inside WAV files, mapping the full infection chain and response steps.

Guardicore security researchers present a full analysis of a cryptomining attack that concealed a cryptominer inside WAV audio files. The report documents the complete attack chain from detection through infection, network propagation, and malware analysis. It also includes recommendations for optimizing incident response processes in data centers.

Akamai Blog · 8d agoMalware in the wild

StepAudio 3 Gen Technical Report

StepAudio 3 Gen unifies TTS, voice design, music, and sound effects via discrete autoregressive modeling over RVQ tokens.

StepAudio 3 Gen is a general-purpose audio generation model covering zero-shot TTS, voice design, vocal generation, sound effects, music, vibe speech, and mixed audio in one framework. It uses discrete autoregressive modeling over residual vector quantization (RVQ) tokens rather than the diffusion Transformer paradigm, with a StepAudio Tokenizer representing audio at 12.5 Hz in a shared 16x2048 residual code space. Key design principles include interference-aware progressive pretraining, an RVQ Adaptor for multi-codebook acoustic representations, and shared discrete autoregressive modeling. The model reports state-of-the-art performance on TTS and voice design while retaining strong generation across speech, vocals, sound effects, and music.

Hugging Face daily papers · 5d agoAI research