ZeroHour

Search: “patient-data”

29 stories

Smart search ranks by meaning as well as keywords (one row per story, last 45 days).

4.1 Million Impacted by AdaptHealth Data Breach

AdaptHealth disclosed a breach affecting 4,115,802 people after a socially engineered attacker stole health and insurance data from cloud-based patient systems.

A threat actor used social engineering to hijack a user session at a third-party contractor and gained access to AdaptHealth cloud applications, including patient management and document storage systems, in early June. Names, contact and demographic information, and health and health insurance data were exfiltrated; Social Security numbers and financial information were not affected. AdaptHealth reported 4,115,802 affected individuals to HHS, whose breach portal listed the incident this week; Baylor Genetics separately reported 2,810,878 individuals affected in a related June healthcare breach.

SecurityWeek · 6d agoData breach 2 sources

McKesson confirms cyber incident after ShinyHunters claims patient-data theft

Healthcare giant McKesson confirmed a cyber incident after ShinyHunters claimed theft of hundreds of millions of patient records.

McKesson acknowledged a data breach following public claims by the threat actor group ShinyHunters that it stole hundreds of millions of records containing patient data. The company confirmed a cyber incident occurred but the full scope of the theft has not yet been independently verified. ShinyHunters is known for large-scale data theft and extortion against major organizations. The healthcare sector remains a frequent target for data-theft extortion groups.

Malwarebytes Labs · 16d agoData breach

Veradigm Confirms Patient Data Exposed in Third-Party Data Breach

Veradigm disclosed a third-party vendor breach exposing patient data including Social Security numbers via stolen vendor API credentials.

Veradigm filed an 8-K with the SEC on September 8, 2026, disclosing that attackers used credentials stolen from a third-party vendor to access a specific vendor-facing API and download patient personal data, including Social Security numbers for some individuals. No clinical or medical information was compromised, and Veradigm's internal infrastructure was not breached directly. The company activated incident response, notified law enforcement, and is offering credit monitoring to affected individuals.

Cyber Security News · 7d agoData breach in the wild

280,000 Impacted by Premier Medical Group Data Breach

New York healthcare provider Premier Medical Group is notifying 282,075 patients that personal and medical information was stolen in a June breach.

New York healthcare provider Premier Medical Group is notifying 282,075 patients whose personal and medical data was stolen in a June breach. Attackers accessed certain files on June 14 after some systems were disrupted, but PMG has not disclosed how the attack occurred or who was responsible. Exposed data includes names, contact information, dates of birth, treatment and diagnostic details, medication information, and health insurance details. PMG reported the incident to HHS, which added it to its public breach portal this week.

SecurityWeek · 10h agoData breach in the wild

NOAH: Learning the Full Patient Journey. A Longitudinal Multimodal Time-Aware Model for Representation and Forecasting

Researchers introduce NOAH, a time-aware generative transformer trained on 559 million MIMIC clinical events to forecast multimodal patient trajectories.

NOAH is a task-agnostic, time-aware generative transformer trained on over 559 million clinical events from 431,000 hospital visits by 299,000 patients across the MIMIC dataset family. It uses bidirectional time integration and a variational latent space to model the stochastic evolution of patient states, natively processing medical images, time-series signals, categorical events, and structured or unstructured clinical records. The model supports autoregressive forecasting with optional time control, zero-shot classification, and counterfactual intervention simulation, with strong probing performance across clinical outcomes, 15 ICD chapters, and 29 comorbidities.

Hugging Face daily papers · 9d agoAI research1

NOAH: Learning the Full Patient Journey. A Longitudinal Multimodal Time-Aware Model for Representation and Forecasting

Researchers introduce NOAH, a generative time-aware transformer trained on 559 million MIMIC clinical events to model and forecast patient trajectories.

NOAH is a task-agnostic, time-aware generative transformer designed to represent and forecast the full multimodal patient journey across medical images, time-series signals, categorical events, and clinical text. It was trained on over 559 million clinical events from 431,000 hospital visits covering 299,000 patients in the MIMIC dataset family. The architecture combines bidirectional time integration with a variational latent space to capture continuous patient state evolution and clinical stochasticity. NOAH supports autoregressive forecasting with time control, zero-shot classification, and counterfactual intervention simulation, with evaluations on 15 ICD chapters, 29 comorbidities, and time-to-event prediction.

arXiv cs.AI / cs.LG / cs.CL · 8d agoAI research2

Performance of Clinical AI System and Physicians and Frontier Language Models in primary care diagnostics

Clinical AI system Doctorina achieved 82.0% primary-care diagnostic concordance versus 57.0% for physicians across 150 synthetic consultations.

The study compared Doctorina, eight physicians, and four standalone frontier language models on 150 synthetic Polish-language primary-care consultations. Doctorina achieved 82.0% Top-1 diagnostic concordance versus 57.0% for physicians (25.0-point difference, 95% CI 17.7-32.7) and 97.3% versus 85.0% primary-or-reference-differential concordance. Normalized workup and treatment scores were 89.4 versus 66.9 and 83.7 versus 61.2. Kimi K3 ranked next on diagnosis, while Claude Opus 5 led the closely spaced management estimates among Opus, Doctorina and Kimi.

arXiv cs.AI / cs.LG / cs.CL · 8d agoAI research

Anchoring Clinical Events in Time: UID-Preserving Multimodal Reconstruction and Source-Grounded Adjudication

UID-preserving multimodal framework plus GAVEL LLM judge improves clinical timeline reconstruction, boosting event recovery 43% over prior matching.

The paper introduces a UID-preserving framework linking each narrative clinical event to its source span through text-only estimation, structured-evidence retrieval, timestamped source-row grounding, and joint revision. GAVEL, an LLM judge, compares UID-aligned timelines against narrative and structured records. Across six open-weight models and 40 mixed-critical-care summaries, GLM 5.2 multimodal revision improved temporal agreement without reducing event recovery and performed competitively with clinician annotations, while DeepSeek V3.2 did not benefit from multimodality. The pipeline achieves 43% increased event recovery with occurrence-level provenance.

arXiv cs.AI / cs.LG / cs.CL · 5d agoAI research1

Veradigm warns of patient data breach after ransomware gang claims attack

Healthcare vendor Veradigm disclosed a patient data breach via a third-party vendor's credentials, which the Gentlemen ransomware gang claims involved 3.5 million records.

Veradigm, formerly Allscripts, told the SEC that an attacker used compromised credentials from a third-party vendor to access a customer-service API and copy patient data, including personal details and Social Security numbers, without touching clinical data or the broader network. The Gentlemen ransomware group listed Veradigm on its leak site claiming 3.5 million patient records and threatened to publish the data by September 11 unless ransom negotiations start. The gang, active since mid-2025, runs double extortion across Windows, Linux, NAS, BSD and ESXi, lists 800+ victims in 86 countries, and has been linked to a SystemBC proxy botnet and the GentleKiller EDR killer. Veradigm is notifying affected individuals, offering credit monitoring, and says it does not expect a material business impact.

BleepingComputer · 7d agoData breach

Emergency Department Revisit Quality Review Screening: Exploring Human Decision-Making and Artificial Intelligence Support

A retrospective study found GPT-4 over-flagged emergency department revisit cases while an LLM knowledge-graph screener achieved 83-100% positive predictive value.

In an exploratory retrospective study of 99 emergency department diagnosis pairs from a multihospital health system, clinicians and GPT-4 independently judged whether revisit pairs warranted further assessment. GPT-4 responses correlated poorly with clinicians, flagging 94% of pairs for follow-up, 4.4-13.3 times more than clinicians, though prompt engineering was minimal. An algorithm leveraging an LLM-populated knowledge graph (KGA) achieved 83-100% positive predictive value against at least one clinician rater, suggesting LLM-based screening could broaden revisit quality review without substantially increasing reviewer workload.

arXiv cs.AI / cs.LG / cs.CL · 7d agoAI research1

Healthcare organizations can now connect EHR and additional industry data to ChatGPT

OpenAI announced ChatGPT integrations allowing healthcare organizations to connect EHR and industry data so clinicians can access patient context and research securely.

OpenAI said healthcare organizations can now connect electronic health records and other industry data sources to ChatGPT. The feature is aimed at letting clinicians securely access patient context and medical research within the assistant. The announcement was published on OpenAI's news site without disclosure of a new model release.

OpenAI News · 15d agoAI industry

LongAgent: History-Guided Agentic Search for Longitudinal Outcome Prediction

LongAgent autonomously searches variable sets and temporal windows to predict longitudinal medical outcomes, beating the strongest non-agent baseline on synthetic data.

The paper proposes LongAgent, an agent-based method that searches over combinations of variable sets, temporal windows and aggregation functions for outcome prediction on heterogeneous medical longitudinal data. It uses a history memory of previous searches and numerical evidence to guide exploration. On synthetic data it achieves mean RMSE 1.7376, improving over the best non-agent baseline by 0.0151 (95% CI [0.0045, 0.0260]; p=0.0273), and performs comparably to the best baseline on a real clinical dataset.

arXiv cs.AI / cs.LG / cs.CL · 2d agoAI research

WearableQA: A Benchmark for Health Reasoning over Real-World Wearable Data

WearableQA benchmark introduces 4,084 questions over real longitudinal wearable data, showing 14 LLMs score 19.6-72.9% on health reasoning, far from solved.

WearableQA is a benchmark of 4,084 10-option multiple-choice questions built from wearable time series, blood biomarkers, and demographics of 200 real users with up to 500 days of daily measurements. It defines 16 question types along two axes: data versus health reasoning, and single- versus cross-signal reasoning, using a dual-grounding framework combining literature and population-validated patterns. Evaluations of 14 proprietary and open-source LLMs show accuracy ranging from 19.6% to 72.9% against a 10% chance baseline, with most models below 60%.

Hugging Face daily papers · 13d agoAI research

WearableQA: A Benchmark for Health Reasoning over Real-World Wearable Data

WearableQA benchmark tests LLM health reasoning over longitudinal wearable data; the best of 14 evaluated LLMs reaches 72.9% accuracy.

WearableQA comprises 4,084 ten-option multiple-choice questions built from wearable time series, blood biomarkers, and demographics of 200 real users with up to 500 days of daily measurements. It defines 16 question types along two axes: data versus health reasoning, and single- versus cross-signal reasoning. Evaluation of 14 proprietary and open-source LLMs shows performance from 19.6% to 72.9% against a 10% chance baseline, with most models below 60%.

arXiv cs.AI / cs.LG / cs.CL · 12d agoAI research

Hackers claim millions of patient records stolen during data breach at healthcare giant McKesson

Hackers claim theft of millions of patient records from US healthcare distributor McKesson, which confirmed a hack and expects service degradation.

McKesson, which distributes medicines and medical devices to hospitals and healthcare practices across the US, said it was hacked. Threat actors claim millions of patient records were stolen in the breach. The company said it expects intermittent service degradation as it responds to the incident.

TechCrunch · Security · 16d agoData breach in the wild

Verifiable by Construction: Claim-Level Evaluation of Verbatim Citation in Clinical Question Answering

Evaluation of twelve LLMs on 222 clinical questions shows verbatim quotes rarely substantiate claims; claude-opus-5 fully substantiates only 37.1%.

The authors build a standardized harness over four clinical practice guidelines and evaluate twelve LLMs on 222 synthetic clinical questions, measuring citation attachment, verbatim quote production, and claim substantiation. Most models attach verbatim quotes to over 90% of claims from prompting alone, though lightweight models like claude-haiku-4.5 struggle. Quotes frequently fail to substantiate claims: claude-opus-5 quotes 98.0% of claims but fully substantiates only 37.1%, exposing a capability gap for verifiable clinical QA.

arXiv cs.AI / cs.LG / cs.CL · 2d agoAI research

OmniMed-FL: A Robust Multimodal Federated Learning Framework for Clinical Diagnosis

OmniMed-FL benchmarks multimodal federated learning for chest radiograph diagnosis across 3-20 clients, with FedProx leading under severe non-IID skew.

OmniMed-FL studies multimodal federated learning combining chest radiographs and clinical notes for five-class condition classification under HIPAA/GDDR-compliant decentralized training. It benchmarks eight fusion strategies, imputation rules, and federated baselines under Dirichlet non-IID partitioning across 3-20 hospital clients. With 5 clients and severe skew (alpha=0.1), FedProx scored 0.737 macro-F1 versus 0.662 for FedAvg and 0.297 for local-only training. Multimodal fusion beat unimodal inputs (0.956 vs 0.934 text, 0.664 images) on the synthetic corpus.

arXiv cs.AI / cs.LG / cs.CL · 7d agoAI research1

Dental contractor set up secret account with access to 4,000 patient records then left the company

A dental practice left an unknown vendor admin account with access to 4,000 patient records active for over three years, creating HIPAA risk.

Chris Kirksey, CEO of Direction, found three admin accounts on a dental practice's patient database, including one belonging to a scheduling company dropped in 2021 that retained access to 4,000 patient records for at least three years. A contractor had created the account without telling anyone and then left, so nobody knew to remove it, creating HIPAA compliance risk. Kirksey removed the accounts, adopted mandatory vendor access shutdown and twice-yearly reviews, and later found similar orphaned-account issues at six other healthcare practices.

The Register · Security · 6d agoIndustry

The 2026 PNPL Competition: Word Classification and Efficient Cross-Subject Generalisation in LibriBrain100

The 2026 PNPL competition releases LibriBrain100, a MEG speech dataset with 32 extra subjects, targeting word classification and cross-subject BCI generalization.

The 2025 PNPL competition on non-invasive speech decoding from MEG achieved F1-macro scores of 95.6% for speech detection and 73.6% for phoneme classification, built on LibriBrain's ~50 hours of single-subject data. The 2026 edition extends this with LibriBrain100, adding 32 subjects (~40 minutes each) plus ~80 hours of within-subject data. Two tracks target within-subject word classification at scale and cross-subject generalization with subject-specific fine-tuning shrinking from ~40 to ~20 to ~10 minutes, aiming at clinically feasible non-invasive BCIs for people with profound paralysis.

Hugging Face daily papers · 14d agoAI research

Healthcare cyberattacks hit pacemakers and millions of patient records

McKesson confirms a data breach as ShinyHunters demands $55.2M amid healthcare cyberattacks affecting millions of patient records and pacemaker devices.

McKesson, one of the largest US healthcare distribution companies, has admitted a data breach. The ShinyHunters group is demanding a $55.2M extortion payment. The Register reports broader healthcare cyberattacks tied to the story, affecting millions of patient records and involving medical devices such as pacemakers. Details on the exact number of compromised records have not been confirmed.

The Register · Security · 15d agoData breach in the wild

Electronic health record company CareCloud says 3.7 million people affected by breach

Healthcare software firm CareCloud reported a breach affecting 3,756,469 people after a hacker spent eight hours in an AWS EHR environment.

CareCloud disclosed to HHS that a hacker accessed one of its AWS environments from March 10 to March 16 and exfiltrated data within an eight-hour window in its electronic health record environment. Stolen data includes Social Security numbers, ID numbers, credit and debit card details, medical information, and insurance data. The company serves more than 45,000 providers and reported the incident to the SEC on March 24. No hacking group has claimed responsibility.

The Record · 28d agoData breach in the wild

Electronic health record company says customer data stolen in breach

Veradigm disclosed that attackers used stolen vendor credentials via an API to steal patient data including Social Security numbers, as the Gentlemen ransomware gang claims 3.5 million patients' records.

Electronic health records company Veradigm filed an 8-K with the SEC stating that an unauthorized party obtained credentials from a vendor's environment and used them to access a Veradigm API, downloading patients' personal data including Social Security numbers; no clinical or medical data was involved. The Gentlemen ransomware gang added Veradigm to its leak site, claiming theft of 3.5 million patients' health records. Access was limited to the specific API interface, with no operational disruption. Veradigm was previously hit by SamSam ransomware in 2019 and disclosed a December 2024 breach affecting 2,672,036 people.

The Record · 7d agoData breach in the wild

Target leakage, not model class, explains reported accuracy in survey-based cardiovascular screening: a leakage-tiered audit of glass-box and tabular foundation models

Audits of 10 classifiers on BRFSS show target leakage, not model class, drives the reported 0.89 AUROC in survey-based cardiovascular screening.

The study benchmarks ten model classes, including glass-box and tabular foundation models, for prevalent myocardial infarction on 442,067 respondents of the 2022 BRFSS across five feature tiers of decreasing leakage risk. Removing two post-diagnostic features costs every model 0.049-0.051 AUROC and collapses performance into a 0.0045-wide band, and the explainable boosting machine matches all alternatives within 0.005 while scoring roughly 104x faster than the strongest foundation model. Frozen models transport within 0.002 AUROC to 2023 data; the authors conclude evaluation practice and feature sets, not model capacity, are the binding constraint.

arXiv cs.AI / cs.LG / cs.CL · 6d agoAI research1

AMIE, our research medical AI system, demonstrates real-time clinical video consultation capabilities in a first-of-its-kind study.

Google's AMIE research medical AI system demonstrates real-time clinical video consultations in a first-of-its-kind simulated study.

Google introduced AMIE, its research medical AI system, demonstrating real-time clinical video consultation capabilities in a first-of-its-kind study. The evaluation was conducted in simulated settings, extending the AMIE diagnostic dialogue research line to multimodal video consultations. AMIE remains a research system rather than a deployed clinical product.

Google · AI · Aug 11, 2026AI research

Personal Data of Approximately 220,000 Domestic and International Gangnam Unni Users Leaked

Healing Paper's Gangnam Unni beauty platform leaked personal data of roughly 220,000 domestic and international users via abnormal API access on September 4.

Healing Paper, operator of the beauty medical platform Gangnam Unni, announced on September 7 that personal information of approximately 220,000 domestic and international users was leaked. The incident involved abnormal access on September 4 to the integration API used to view consultation records. The company disclosed the event through a public notice.

DataBreaches.net · 9d agoData breach

18 ways to check whether data can be trusted for AI

ETSI published TR 104 180 defining 18 data quality metrics, including fairness and privacy, to assess whether datasets are fit for AI.

ETSI's technical report TR 104 180 defines 18 metrics across four groups - intrinsic quality, usability and lineage, fairness, and privacy - each with calculation formulas, plus an open-source tool that scores datasets. Testing on an aircraft engine sensor dataset and a US census dataset revealed a roughly threefold gender gap in high earners (about 31% of men versus 11% of women) and two privacy failures: re-identification via age, race, sex, and country, and sensitive fields stored in plaintext. The working group included Sejong University, EGM, TTA, Daejeon University, and CNIT.

Help Net Security · 9d agoAI policy

Canada’s Hospital for Sick Children attacked by cybercriminals again as employee data stolen

Canada's Hospital for Sick Children reported a data theft incident exposing current and former employee data via a third-party application.

The Hospital for Sick Children (SickKids), Canada's largest pediatric health center, disclosed that hackers stole personal information of current and former employees, job applicants, and SickKids Foundation staff, likely through a third-party software application. The attack briefly took down the hospital's careers website but did not involve clinical systems or patient information. Affected individuals were notified and offered two years of credit monitoring. The hospital was previously hit by ransomware in 2022.

The Record · 26d agoData breach in the wild

DynSHAP: Towards Explainable Dynamic Survival Analysis

DynSHAP extends SHAP explainability to dynamic survival analysis, treating time-feature pairs as Shapley players for longitudinal clinical predictions.

DynSHAP adapts marginal SHAP estimators to dynamic survival analysis by treating time-feature pairs as players in the Shapley game, handling longitudinal irregular inputs and functional survival outputs. Temporal DynSHAP learns linear feature dependencies over time and addresses them with conditional sampling. On synthetic data with ground-truth attributions it recovers temporally dependent features more accurately than marginal estimators, and it produces faithful attributions on two real-world clinical datasets across two DSA architectures.

arXiv cs.AI / cs.LG / cs.CL · 5d agoAI research