ZeroHour

Search: “Nine”

16 stories in the last 3d

ProgramDistill: From Interactive Web Apps to Verifiable Reference-Guided SWE Tasks

ProgramDistill is a benchmark evaluating coding agents on reconstructing web app features from reference applications, testing nine frontier agents.

ProgramDistill evaluates coding agents on features discovered through interaction with fully functional reference applications, factorizing apps into features with replayable behaviors verified via gold patches. Its mine-craft-patch pipeline discovered 1,975 replay-verified behaviors across 26 applications and built 4,063 tasks without human intervention. On cumulative full-application reconstruction workflows, GPT-6 Astra achieved 49.2% and Claude Opus 5 28.8% success. In partial reconstruction, success drops from 100% to 64.0% and from 96% to 32% as restoration depth increases from 1 to 8.

Hugging Face daily papers · 1d agoAI research

Confidence Comes from Experience: Experiential Confidence Estimation from Reasoning to Agents

XConf estimates LLM confidence from accumulated past episodes, improving calibration and discrimination across nine benchmarks at one-tenth self-consistency cost.

Researchers propose XConf, an experiential confidence estimator that augments current inference with a stored record of the model's graded past episodes, including reflections, stated confidence, outcomes, and lessons. A Recall stage retrieves episodes from similar tasks with similar stated confidence and reads off historical success rates, while a Reflect stage prompts the model to name recurring failure modes and restate confidence. Across nine benchmarks in reasoning, coding, multimodal QA, and interactive agents, and four models from three families, XConf beats or matches ten-sample self-consistency in AUROC on 23 of 24 comparisons with much lower ECE, at a tenth of the generation cost. For selective prediction, abstaining on the 10% least-confident episodes raises delivered success rate by up to 8.7 points on agent tasks.

Hugging Face daily papers · 2d agoAI research

The AI security question leaders should be asking instead

Gremlin security officer Frederic Bull argues AI has eroded the attacker-defender skill asymmetry while least-privilege controls remain essential for securing AI agents.

In a Help Net Security interview, Gremlin Security Officer Frederic Bull says AI has narrowed the expertise gap between attackers and defenders, enabling faster exploit discovery even by less-skilled actors. His team processed roughly nine times more vulnerabilities in the past year with unchanged staffing using LLM-based tooling, cutting time-to-remediate by about 5%. He argues least privilege, session-based RBAC via OIDC/OBO, and human-in-the-loop oversight remain the bedrock defenses for AI agents, and that hiring should favor engineers able to catch confidently wrong AI output.

Help Net Security · 4h agoIndustry

Reporting Practice Matters: The Impact of Reference Choice on Chest X-ray Report Evaluation

Study shows radiology reporting-style variations in reference reports can flip rankings of chest X-ray report generation models; releases MIMIC-CXR-Ext-ReRef dataset.

The paper quantifies how variations in radiologists' reporting practices distort evaluation of radiology report generation (RRG) models, introducing a radiologist-informed taxonomy and the ReRef method for rewriting reference reports while preserving clinical meaning. On MIMIC-CXR with RadCliQ-v1, condensing normal-findings discussion caused Libra to drop from first to second while CheXOne rose from third to first among nine models. The authors release MIMIC-CXR-Ext-ReRef, a radiologist-validated dataset of 120 original/alternative reference pairs, arguing metrics conflate clinical correctness with stylistic conformity.

arXiv cs.AI / cs.LG / cs.CL · 16h agoAI research

Detecting Logic Vulnerabilities Across the Contract and Device Layers of Blockchain-Enabled IoT With Multi-Agent Heterogeneous Graph Attention

MA-HGAT framework detects logic vulnerabilities across smart contract and IoT device firmware layers using multi-agent heterogeneous graph attention.

Researchers extend MA-HGAT into a cross-layer multi-agent heterogeneous graph attention framework that models smart contracts, firmware artifacts, device fleets, and transaction streams for blockchain-enabled IoT security. A four-role, nine-relation schema supports graph-, link-, and node-level detection tasks, while a gateway-cloud partition enables lightweight edge inference on resource-constrained devices.

arXiv cs.CR · 1d agoResearch

MSPs say nearly half their customers rely on them for CISO services

Sophos survey finds MSPs act as CISOs for an average 46% of customers, mostly with partial compliance offerings and fragmented manual reporting.

A Sophos survey found MSPs estimate that 46% of their customers rely on them to act as CISOs, and most providers deliver only four to six of seven measured compliance services. More than half say they manage customers' full compliance programs, while compliance requirements influence roughly half of customers' security purchases. Nearly nine in ten providers use software, but most juggle multiple tools that cannot feed a central reporting platform, forcing staff to combine data manually. Providers estimated a unified platform would cut time spent on posture, compliance and reporting by about half; Sophos sells CISO Advantage via its Sophos Fusion system for this work.

Help Net Security · 1d agoIndustry

The AI data center boom is colliding with cities scarred by big industry

Philadelphia activists rally against AI data center construction amid energy and pollution concerns, joining a wave of U.S. city moratoriums.

Residents of Philadelphia's Grays Ferry, home to a former oil refinery, launched the 'No Data Centers in Philly' campaign over pollution, noise, and resource concerns. BloombergNEF projects U.S. data centers will consume more natural gas than Germany and Japan combined by 2035. New York Governor Kathy Hochul signed an executive order pausing permits for large data center projects, and moratoriums have passed in Denver, Indianapolis, Asheville, Charlotte, and Reno.

TechCrunch · AI · 1d agoAI industry

1Password's AI patching benchmark is misleading

Trail of Bits reanalysis says 1Password's 26% AI clean-fix rate is misleading; 86% of eligible patches blocked exploits.

Trail of Bits critiques 1Password's FLAWED AI patching benchmark, arguing its 26% clean-fix headline mixes trials where agents were instructed to apply wrong fixes (22% of data) with trials that prohibited compiling or testing (36%). Restricting to reasonable conditions, 2,634 of 3,067 patches (86%) blocked the supplied exploit. Trail of Bits also reports 12.5% of 2,265 developer first fixes failed in its own 2024-2026 assessments, and released post-patch-validation and review-walkthrough agent skills.

Lobsters · security · 1d agoResearch1

What’s next for CISA’s CDM program that gives cybersecurity tools to federal agencies

CISA officials outline future plans for the CDM program, emphasizing speed, automation, unified data, and data-driven federal risk management.

Speaking at an Elastic Federal Cyber Defense Breakfast, CISA officials described next steps for the Continuous Diagnostics and Mitigation (CDM) program that supplies cybersecurity tools to federal agencies. Acting deputy program manager Richard Grabowski named velocity, unification, and data-driven risk management as core goals, including a three-year roadmap to expand SIEM-as-a-Service. Federal CISO Mike Duffy urged aggregating demand across agencies, buying outcomes rather than products, and designing acquisition for continuous improvement. CISA's Matt House tied the program's evolution to post-SolarWinds needs for a government-wide common operating picture.

CyberScoop · 1d agoPolicy & legal

Building a Linux GPU Driver for the M4 Mac Mini in One Month

Two developers built a fully OpenGL ES 3.0 compliant Linux GPU driver for the M4 Mac Mini in one month via clean-room reverse engineering.

Niklas and the author reverse engineered Apple's AGX GPU firmware ABI and user-space components in about a month, a process that normally takes years, producing an OpenGL ES 3.0 conformant driver fast enough to run Minecraft at 200fps on an M4 Mac Mini. The work was done transparently using hypervisor traces without examining Apple binaries, following clean-room practices, and included a custom shader compiler, command stream builder, and a full Linux kernel driver for the firmware ABI. The A18 Pro firmware ABI proved significantly more complex than the M1's, with 1.5x as many structs and twice as many pointers. All experiments and provenance evidence were published in public agx-re repositories.

US data centers could consume more natural gas than Germany and Japan combined by 2035

BloombergNEF projects US data centers will consume about 18 billion cubic feet of natural gas daily by 2035, exceeding Germany and Japan combined.

A new BloombergNEF report forecasts US data centers will consume roughly 18 billion cubic feet of natural gas per day by 2035, nearly double the estimate from nine months ago. On-site gas plants planned by Meta, Microsoft, Google, and Amazon would use 2.9-3.4 billion cubic feet per day, while grid-connected data centers drive an additional 15 billion cubic feet per day of power-sector gas demand. The added demand would generate about 1 million metric tons of extra greenhouse gas pollution daily, roughly 12% of current US emissions.

TechCrunch · AI · 1d agoAI industry

Coding Agents Have Converged: Why the SWE-bench Leaderboard Can No Longer Order Its Top Entries, and What to Measure Instead

Audit of 254 SWE-bench submissions finds top coding-agent entries statistically inseparable, so small leaderboard gaps no longer establish rank.

The paper audits 254 SWE-bench submissions across four splits without running models. On Verified, the top two entries each resolve 396 of 500 instances, and exact paired McNemar tests separate none of the 29 adjacent top-thirty pairs at alpha=0.05. Within-model scaffold score ranges reach 29.8 percentage points, versus an 8.8-point spread among the top thirty. The authors release a five-step audit protocol and recommend reporting comparison-set-specific resolution and model-scaffold provenance.

arXiv cs.AI / cs.LG / cs.CL · 1d agoAI research

AI for everyone in every language

Google says its AI now spans 300+ languages reaching 7 billion people, unveiling Gemini 3.5 Transcribe, Live Translate, and TranslateGemma models.

Google announced its technologies now support more than 300 languages spoken by 7 billion people, 86% of the global population. Gemini 3.5 Live Translate powers real-time spoken translation across 70 languages and 2,000+ language pairs, while Gemini 3.5 Transcribe is its most precise speech-to-text model. Its Universal Speech Model was trained on 12 million hours of audio using cross-lingual transfer learning, and TranslateGemma is a family of lightweight open translation models covering 55 languages that run on-device. Open-data partnerships include WAXAL covering 27 Sub-Saharan African languages and Project Vaani with 30,000+ hours of speech across 109 languages.

Google · AI · 1d agoAI industry

Swiss court sentences 52-year-old Ukrainian ransomware dev to nearly 13 years in the cooler

Zurich court sentences Ukrainian ransomware developer to 12 years, 9 months for LockerGoga, MegaCortex and Nefilim attacks including Stadler Rail.

Zurich District Court sentenced a 52-year-old Ukrainian to 12 years and 9 months for developing LockerGoga, MegaCortex, and Nefilim ransomware, plus a 10-year ban from Switzerland; the verdict can be appealed. The operations hit over 1,800 victims across 71 countries with losses of several hundred million Swiss francs, including Stadler Rail (2020, $6 million Nefilim demand), Meier Tobler, and Crealogix. Alleged mastermind Volodymyr Tymoshchuk, indicted in the US and tied to at least 250 companies including Norsk Hydro, remains at large with an $11 million FBI bounty.

The Register · Security · 1d agoPolicy & legal

Microsoft AI Code of Conduct Sets Cyberattack Boundaries, Chain of Command, Safety Constraints

Microsoft AI's draft Humanist AI Code of Conduct blocks MAI models from producing exploit code and constrains autonomous agent behavior.

The draft code sets 'Absolute Constraints' preventing MAI models from generating working exploit code, attack tooling, or intrusion guidance, while permitting authorized defensive work such as vulnerability discovery and malware analysis. A 'Chain of Command' rule means tool outputs, file contents, and webpages carry no authority over model behavior, countering injected instructions. Microsoft opened a six-week public consultation; a revised version will guide 2027 model development, and current MAI Models were not trained on the document.

SecurityWeek · 2d agoAI safety & security1

LLM-Based Schema-Aware Split Learning for Privacy-Preserving Mental Distress Prediction Across Heterogeneous Surveys

Schema-aware split learning uses LLaMA-3.2-3B-Instruct as shared semantic encoder to harmonize heterogeneous mental-health surveys while raw data stays local.

The paper proposes a schema-aware split learning framework where an LLM serializes heterogeneous mental health survey records into natural language and is fine-tuned via LoRA, partitioned across client and server. Clients keep raw survey responses local and run only a lightweight front-end while the resource-intensive backbone runs server-side. Using LLaMA-3.2-3B-Instruct, the framework attains an average ANLS of 0.708 with 2,000 training samples, beats federated learning in eight of nine settings, and cuts per-client computation by three orders of magnitude while generalizing to unseen datasets.

arXiv cs.AI / cs.LG / cs.CL · 2d agoAI research