35
35
42
Google Research Releases ToolGrad: Answer-First Framework Hits 99.8% Pass Rate for Tool-Use Data Generation
Google Research and partners introduce ToolGrad, a verified tool-chain-first data generation framework reaching 99.8% pass rate and boosting Gemma-3-12B to 83.1 on BFCL.
Researchers from Google, the University of Tokyo, RIKEN AIP, and Tohoku University released ToolGrad, which inverts query-first tool-use data generation by executing and verifying API chains before annotating them with user queries. On the ToolBench database of 16,000+ APIs, ToolGrad raised generation pass rate from 63.8% to 99.8% while increasing tool uses per sample from 2.1 to 3.4 and cutting tool-use steps from 34.3 to 20.0. Fine-tuning Gemma-3 at 1B, 4B, and 12B parameters on the 500-sample ToolGrad-500 dataset lifted ToolGrad-12B to 83.1 on the Berkeley Function Calling Leaderboard, near Gemini 2.5 Pro at 83.2 and ahead of GPT-5 at 74.4. Code is Apache-2.0, with the dataset, PyPI package, and models available on Hugging Face.
58
Root-Cause Attribution Is a Search Problem: Continual Search for Long-Horizon Agent Failures
Continual Search framework iteratively prompts LLM judges to keep searching agent execution logs, boosting long-horizon failure root-cause attribution accuracy.
The paper frames automated root-cause attribution (RCA) for long-horizon AI agent failures as a search problem, since relevant evidence is sparse and distributed across massive execution traces. The authors propose Continual Search, an iterative framework that nudges an LLM judge across successive turns to keep hunting unresolved diagnostic evidence instead of settling on an early plausible diagnosis. They introduce MegaRCA-Mix, a benchmark of 50 human-annotated failure trials on long-horizon, execution-heavy tasks. On MegaRCA-Mix, Continual Search improves GPT-5.5's F1 from 0.349 to 0.498 (over 40% gain), and lower-tier models can surpass higher-tier counterparts when search is effective.
42
Jackrong/Qwopus3.8-27B-Flash-GGUF — new model trending #26 on Hugging Face
Community fine-tune Qwopus3.8-27B-Flash, built on Qwen3.8-27B, cuts agent reasoning latency with 12.8% faster decoding and 80.7% MTP acceptance.
Jackrong released Qwopus3.8-27B-Flash, a fine-tune of Qwen3.8-27B optimized for long-running agent workloads, reporting 12.8% faster decoding and 80.7% multi-token-prediction acceptance. Training used roughly 1.5 million teacher-scored SFT examples filtered to the top 10%, followed by reinforcement training with NVIDIA NeMo-RL and GSPO. The author notes an explicit trade-off: MMLU-Pro mixed-set scores are lower than the base model, and a known bug can produce incorrect Python indentation. Author-provided benchmarks have not been independently verified.
35
FulcrumSec Claims Responsibility for Manchester Airport Group Breach
FulcrumSec leaked ~549GB of Manchester Airport Group data, claiming 8.7M customer profiles exposed via exposed Iterable admin keys.
FulcrumSec posted around 549GB of uncompressed stolen Manchester Airport Group (MAG) data on its leak site, claiming nearly 8.7 million customer profiles with email, name, phone, home town, postcode and residential IP. The group said initial access came from Iterable platform admin keys exposed in the root-domain JavaScript of the Manchester, Stansted and East Midlands airport websites. Allegedly stolen data also includes ~1.2 billion marketing events, 2.5 million bookings, 461,000 SMS records, 108,000 vehicle plates and ~191,000 future bookings. MAG has provided no update since August 27 and the claims remain unverified.
75
Extortion Group FulcrumSec Claims 86GB Manchester Airports Data Theft
Extortion group FulcrumSec claims stealing 86GB of Manchester Airports Group data, exposing 8.7 million customers' personal and booking details.
Manchester Airports Group disclosed a breach on August 27 affecting parking, lounge, Fast Track and WiFi registrations at Manchester, London Stansted and East Midlands airports, impacting 8.7 million customers, most exposed only email addresses. FulcrumSec claims it stole about 86GB via airport-specific Iterable API credentials exposed in client-side JavaScript, including a 21.5GB Manchester export with booking histories, marketing data and nearly 200,000 records on upcoming 2026 travel. BleepingComputer verified sample records against a real traveler's Fast Track history; MAG declined to address the group's specific claims. Researchers warn the combination of UK postcodes, vehicle registrations and booking details could enable convincing targeted phishing, and MAG says no payment card or banking data was exposed.
82
60
45
55
60
30
57
42
57
60
60
45
55
30
60
30
42
30
57
30
60
55
35
60
30
60
55
30
60
30
60
30
30
Coupled Calibration and Learning: Mitigating Teacher Bias in LLM Distillation without Target-Domain Reward Feedback
CCL couples teacher calibration with student updates via token-level branching, provably removing teacher bias in LLM distillation.
The paper proposes Coupled Calibration and Learning (CCL), an LLM distillation algorithm that alternates teacher calibration using source-question reward feedback with student training on target questions under covariate shift. Each iteration calibrates the teacher on source feedback, trains the student on target questions, and lets the updated student inform subsequent calibration. The authors prove the student's expected KL divergence to the oracle student converges to zero at a polynomial rate, and show regularized direct matching error can remain bounded away from zero.
20