ZeroHour

Search: “iterable”

1,745 stories

Google Research Releases ToolGrad: Answer-First Framework Hits 99.8% Pass Rate for Tool-Use Data Generation

Google Research and partners introduce ToolGrad, a verified tool-chain-first data generation framework reaching 99.8% pass rate and boosting Gemma-3-12B to 83.1 on BFCL.

Researchers from Google, the University of Tokyo, RIKEN AIP, and Tohoku University released ToolGrad, which inverts query-first tool-use data generation by executing and verifying API chains before annotating them with user queries. On the ToolBench database of 16,000+ APIs, ToolGrad raised generation pass rate from 63.8% to 99.8% while increasing tool uses per sample from 2.1 to 3.4 and cutting tool-use steps from 34.3 to 20.0. Fine-tuning Gemma-3 at 1B, 4B, and 12B parameters on the 500-sample ToolGrad-500 dataset lifted ToolGrad-12B to 83.1 on the Berkeley Function Calling Leaderboard, near Gemini 2.5 Pro at 83.2 and ahead of GPT-5 at 74.4. Code is Apache-2.0, with the dataset, PyPI package, and models available on Hugging Face.

MarkTechPost · 5d agoAI research1

Root-Cause Attribution Is a Search Problem: Continual Search for Long-Horizon Agent Failures

Continual Search framework iteratively prompts LLM judges to keep searching agent execution logs, boosting long-horizon failure root-cause attribution accuracy.

The paper frames automated root-cause attribution (RCA) for long-horizon AI agent failures as a search problem, since relevant evidence is sparse and distributed across massive execution traces. The authors propose Continual Search, an iterative framework that nudges an LLM judge across successive turns to keep hunting unresolved diagnostic evidence instead of settling on an early plausible diagnosis. They introduce MegaRCA-Mix, a benchmark of 50 human-annotated failure trials on long-horizon, execution-heavy tasks. On MegaRCA-Mix, Continual Search improves GPT-5.5's F1 from 0.349 to 0.498 (over 40% gain), and lower-tier models can surpass higher-tier counterparts when search is effective.

Hugging Face daily papers · 5d agoAI research1

Jackrong/Qwopus3.8-27B-Flash-GGUF — new model trending #26 on Hugging Face

Community fine-tune Qwopus3.8-27B-Flash, built on Qwen3.8-27B, cuts agent reasoning latency with 12.8% faster decoding and 80.7% MTP acceptance.

Jackrong released Qwopus3.8-27B-Flash, a fine-tune of Qwen3.8-27B optimized for long-running agent workloads, reporting 12.8% faster decoding and 80.7% multi-token-prediction acceptance. Training used roughly 1.5 million teacher-scored SFT examples filtered to the top 10%, followed by reinforcement training with NVIDIA NeMo-RL and GSPO. The author notes an explicit trade-off: MMLU-Pro mixed-set scores are lower than the base model, and a known bug can produce incorrect Python indentation. Author-provided benchmarks have not been independently verified.

Hugging Face trending models · 12d agoModel release1

FulcrumSec Claims Responsibility for Manchester Airport Group Breach

FulcrumSec leaked ~549GB of Manchester Airport Group data, claiming 8.7M customer profiles exposed via exposed Iterable admin keys.

FulcrumSec posted around 549GB of uncompressed stolen Manchester Airport Group (MAG) data on its leak site, claiming nearly 8.7 million customer profiles with email, name, phone, home town, postcode and residential IP. The group said initial access came from Iterable platform admin keys exposed in the root-domain JavaScript of the Manchester, Stansted and East Midlands airport websites. Allegedly stolen data also includes ~1.2 billion marketing events, 2.5 million bookings, 461,000 SMS records, 108,000 vehicle plates and ~191,000 future bookings. MAG has provided no update since August 27 and the claims remain unverified.

Infosecurity Magazine · 14d agoData breach

Extortion Group FulcrumSec Claims 86GB Manchester Airports Data Theft

Extortion group FulcrumSec claims stealing 86GB of Manchester Airports Group data, exposing 8.7 million customers' personal and booking details.

Manchester Airports Group disclosed a breach on August 27 affecting parking, lounge, Fast Track and WiFi registrations at Manchester, London Stansted and East Midlands airports, impacting 8.7 million customers, most exposed only email addresses. FulcrumSec claims it stole about 86GB via airport-specific Iterable API credentials exposed in client-side JavaScript, including a 21.5GB Manchester export with booking histories, marketing data and nearly 200,000 records on upcoming 2026 travel. BleepingComputer verified sample records against a real traveler's Fast Track history; MAG declined to address the group's specific claims. Researchers warn the combination of UK postcodes, vehicle registrations and booking details could enable convincing targeted phishing, and MAG says no payment card or banking data was exposed.

Security Affairs · 16d agoData breach in the wild

Coupled Calibration and Learning: Mitigating Teacher Bias in LLM Distillation without Target-Domain Reward Feedback

CCL couples teacher calibration with student updates via token-level branching, provably removing teacher bias in LLM distillation.

The paper proposes Coupled Calibration and Learning (CCL), an LLM distillation algorithm that alternates teacher calibration using source-question reward feedback with student training on target questions under covariate shift. Each iteration calibrates the teacher on source feedback, trains the student on target questions, and lets the updated student inform subsequent calibration. The authors prove the student's expected KL divergence to the oracle student converges to zero at a polynomial rate, and show regularized direct matching error can remain bounded away from zero.

arXiv cs.AI / cs.LG / cs.CL · 21h agoAI research