ZeroHour

Search: “ModernBERT-Base”

28 stories

Smart search ranks by meaning as well as keywords (one row per story, last 45 days).

Re: Retrospective by 'gpg.fail' authors

Sam James replies to GnuPG author Werner Koch that upstream may freely adopt stricter -W compiler warnings, while distros should avoid blanket -Werror.

This is an oss-security mailing list reply by Sam James (Gentoo toolchain maintainer) to Werner Koch in the thread on the 'gpg.fail' authors' retrospective about GnuPG. James states that upstream projects like GnuPG should feel free to enable whatever -W* warning flags they need. He adds that distributions should not use -Werror indiscriminately, allowing narrow exceptions such as -Werror=format-security, and may reject warning-driven bugs that do not reflect real issues. No new CVE or vulnerability details are disclosed in this reply.

oss-securityupdated · 8h agofirst · 1d agoVulnerability 9 sources

Climate-ModernBERT: Revisiting Corpus Composition for Domain-Adaptive Continued Pretraining

Climate-ModernBERT domain-adapted encoders reach 76.3 average F1 across nine climate benchmarks, 2.8 points above vanilla ModernBERT-Base.

The authors continue pretraining ModernBERT-Base on three climate corpora - academic text, climate-filtered web data, and synthetic documents - and compare joint mixtures against parameter-space merging of specialized checkpoints. The best model achieves 76.3 average F1 across nine climate NLP benchmarks, a 2.8-point improvement over the vanilla baseline. Academic climate corpora provide the strongest adaptation signal, and parameter-space merging outperforms joint multi-source training while preserving complementary corpus information; all variants are released.

arXiv cs.AI / cs.LG / cs.CL · 9d agoAI research

τ^τ-Bench: An Environment for End-To-End, Realistic Agent Construction

New τ^τ-bench tasks coding agents with building deployable customer-service agents; best config, Claude Opus 5, passes only 23.9% of simulations.

Researchers introduce τ^τ-bench, an end-to-end benchmark where a developer agent must build a complete customer-service agent from real business records, a client with requirements, a production API, an inherited codebase, and cost/model limits, then is scored by deploying it against held-out simulated users. Across 53 tasks in four domains, the strongest configuration, Claude Opus 5 under Claude Code, passes just 23.9% of evaluation simulations versus an 82.2% expert-authored reference ceiling. Failure modes mirror those of human developers: shallow queries instead of deep record comprehension, almost no client communication, and shipping the first architecture that runs rather than experimenting.

Hugging Face daily papers · 13d agoAI research1

We got admin access to Baseten's production GitHub in 25 minutes

Strix autonomous hacking agent extracted a working GitHub token with repo admin rights from Baseten's public Harbor image; Baseten rotated it next day.

Strix, an autonomous hacking agent, scanned *.baseten.co without credentials and found a public Harbor container registry project anonymously exposing the baseten/baseten-app image. A GitHub personal access token for basetenbot, embedded in Docker build history since March 2023, still worked in July 2026 and granted admin/push rights to basetenlabs/baseten, flux-cd, and homebrew-tap plus read/write on private customer repos. Baseten, valued at $13 billion, confirmed the issue as critical and rotated the token within a day.

Show HN: Self-hosted company OS, Claude Code and Codex agents in departments

OtoDock, a self-hosted company OS that organizes Claude Code and Codex AI agents into departments, was launched on GitHub via Show HN.

OtoDock is a self-hosted 'company OS' shared on GitHub through a Show HN post, presenting Claude Code and Codex AI agents organized into department-style teams. The Hacker News feed entry shows the post reached 20 points with 5 comments; no further technical details are provided in the available text.

Xen Security Advisory 511 v3 (CVE-2026-79603) - Unconditionally do TLB flushing ahead of page scrubbing

Xen Project released XSA-511 (CVE-2026-79603) fixing missing TLB flushes before page scrubbing that can leak x86 PV guest data.

Xen Security Advisory 511 v3 publicly discloses CVE-2026-79603, a TLB handling flaw in the Xen hypervisor on x86. x86 PV guests can free memory pages while a stale TLB entry still points to them, and Xen only flushes the TLB when the page is reused, potentially exposing stale data ahead of scrubbing. The advisory changes Xen to unconditionally flush the TLB ahead of page scrubbing. The issue was published in version 3 of the advisory.

USN-8752-1: Konsole vulnerability

Ubuntu patches Konsole URL-handling flaw that could let a remote attacker execute arbitrary code under specific circumstances.

USN-8752-1 fixes a Konsole vulnerability where certain URLs are incorrectly handled under specific circumstances. A remote attacker could possibly exploit this to execute arbitrary code on a victim's system. Ubuntu released updated packages for affected releases.

Ubuntu Security Notices · 3d agoAdvisory

USN-8571-2: Apache HTTP Server regression

Ubuntu issues USN-8571-2 fixing an Apache HTTP Server regression that prevented startup when HTTP/2 proxying was enabled.

Ubuntu released USN-8571-2 to fix a regression introduced by USN-8571-1 in Apache HTTP Server. The earlier fix was incomplete due to a missing library symbol, causing Apache to fail to start when HTTP/2 proxying was enabled. The original advisory addressed CVE-2026-33007, a memory-handling flaw in mod_authn_socache allowing remote denial of service, and an HTTP response splitting vulnerability affecting multiple modules, credited to Pavel Kohout, Arkadi Vainbrand, Haruki Oyama, Merih Mengisteab, and Dawit Jeong.

Ubuntu Security Noticesupdated · 5d agofirst · 6d agoAdvisory 7 sourcesCVE-2026-330071

Claude's new system prompt really doesn't want to reproduce song lyrics

Anthropic published updated Claude consumer system prompts, including changes steering the model away from reproducing song lyrics, likely over copyright concerns.

Anthropic publishes system prompts for Claude.ai and Claude mobile apps, including historic revisions, and has reorganized them into an index with per-model pages such as the Haiku 4.5 page showing the original October 15, 2025 prompt and an updated January 18, 2026 version. The latest consumer prompt strongly discourages reproducing song lyrics, a behavioral constraint likely tied to copyright considerations. Prompts for Claude Cowork and Claude Code are not included in the published set.

Simon Willison · 14d agoAI safety & security1

BotBase for Operators: A clearer path to joining Cloudflare's directory of bots and agents

Cloudflare launched BotBase for Operators, a dashboard for bot and agent operators to manage directory submissions, track status, and declare content usage.

Cloudflare rolled out BotBase for Operators, giving bot operators a dashboard to manage their listings in Cloudflare's directory of bots and agents. The update adds submission status tracking, submission editing, and a behavior model for declaring how bots use site content. It is aimed at operators of crawlers and automated agents seeking placement in the directory.

Cloudflare Blog · 19d agoAI tools & infra

Agora: Git as Shared Memory for Collective AutoResearch

Agora records multi-agent research as an append-only Git DAG; 13 LLM workers ran nearly 12 days on a weight-transfer problem.

Agora stores every result, hypothesis, and verification as an immutable commit in a Git-stored DAG, with a derived index exposing the frontier and verification status of claims. In a nearly 12-day run, 13 language-model workers with no assigned tasks or central planner published 1,703 contributions on initializing a frozen 119.6M-parameter attention-SSM hybrid from 141 donor models. They improved the evaluator from 3.39 to 1.899 bits per byte, closing 62% of the gap to a trained GPT-2 124M, with 165 independent reproductions posted and none failing.

Hugging Face daily papers · 1d agoAI research

Fluid Notarization: Verifiable Evolution of Concurrently Edited Structured Documents

Fluid Notarization anchors delta-CRDT change graphs on blockchain, providing verifiable provenance for concurrently edited documents, demonstrated on collaborative electronic health records.

The paper introduces Fluid Notarization, a paradigm that notarizes the evolution of collaboratively edited structured documents rather than isolated snapshots. It builds on Melda, a JSON-native delta-CRDT representing changes as compact content-addressed deltas linked by causal dependencies, with blockchain notarization reduced to recording identifiers of evolution artifacts while synchronization, reconstruction, and conflict resolution remain off-chain. The architecture combines deterministic CRDT convergence with independently auditable proof-of-existence, provenance, and publication evidence, validated through a prototype based on collaboratively edited electronic health records.

arXiv cs.CR · 19h agoResearch

On Identifying Sound Conditions for Frontrunning Resistance

Researchers formally define smart-contract frontrunning resistance, showing 55% of 393 audited vulnerabilities escape state-of-the-art detection, and find two undisclosed Ethereum flaws.

The paper gives the first formal definition of frontrunning vulnerability for smart contracts, grounded in how honest users interact with contracts rather than contract code alone. In a large-scale study of 287 smart contract audits, 55% of the 393 vulnerabilities reported by leading auditors fall outside the scope of state-of-the-art dynamic detection criteria. The authors present a sound algorithm for synthesizing secure interaction conditions and apply it to real-world contracts, uncovering previously undiscovered vulnerabilities in two Ethereum contracts.

arXiv cs.CR · 6d agoResearch

[webapps] Metabase 0.61.0 - Authenticated Remote Code Execution

Exploit-DB published an authenticated remote code execution exploit targeting Metabase version 0.61.0.

A new Exploit-DB entry (ID 52680) describes an authenticated remote code execution vulnerability in Metabase 0.61.0. The listing provides minimal detail, but authenticated RCE in a widely deployed BI tool is notable for defenders running exposed instances. No CVE id or in-the-wild exploitation is mentioned in the listing.

Exploit-DB · 14d agoExploit / PoC1

China-Linked Hackers Exploit Chrome-Windows Zero-Day Chain to Deploy GRIMWEDGE

Volexity attributes spear-phishing campaign exploiting Chrome-Windows zero-day chain to Chinese actors UTA0560 and APT31 deploying GRIMWEDGE and LONGTALE.

Volexity attributes a September 1, 2026 spear-phishing campaign targeting NGOs to China-linked UTA0560, which abused a reflected XSS flaw on a US university website to trigger a three-part exploit chain (CVE-2026-85046, CVE-2026-87491, CVE-2026-85880) escaping the Chrome V8 and browser sandboxes to deploy the GRIMWEDGE JavaScript backdoor with reconnaissance, file management, and command execution capabilities. The same chain was used near-simultaneously by JungleBamboo (APT31) to deploy SUPERSTOMP, installing the LONGTALE/GemStone credential-stealing Chrome extension masquerading as Google Gemini. The Chrome flaws were patched in Chromium but not yet in stable Chrome, creating an unusual patch-gap zero-day window attackers raced to exploit.

The Hacker Newsupdated · 2d agofirst · 2d agoThreat actor in the wild 8 sourcesCVE-2026-85046CVE-2026-87491CVE-2026-858801· 1 read

Can Foundation Models Moderate Online Content? Evaluating Instruction- vs. Example-Driven Policy Operationalization

ModerationBench shows foundation models can nearly triple Bluesky's moderation F1 (0.60 vs 0.22), with instruction- and example-driven guidance performing comparably.

Researchers built ModerationBench, a new benchmark of 4,000 manually annotated in-the-wild posts from Bluesky, to test whether foundation models can reliably operationalize content moderation policies. They systematically compare instruction-driven guidance (reasoning from policy precepts) with example-driven guidance (generalizing from precedents) for Vision-Language Models. Both paradigms achieve comparable peak effectiveness, and foundation models nearly triple the F1 of Bluesky's deployed moderation system on Random Posts (0.60 vs 0.22).

arXiv cs.AI / cs.LG / cs.CL · 7d agoAI research1

[webapps] PodcastGenerator 3.2.9 - Stored XSS

A stored cross-site scripting flaw in PodcastGenerator 3.2.9 is documented with a public proof-of-concept exploit on Exploit-DB.

Exploit-DB published exploit ID 52677 targeting PodcastGenerator 3.2.9, a web application affected by stored cross-site scripting. The listing provides a proof-of-concept for the flaw but includes no CVE identifier or reports of active exploitation.

Exploit-DB · 15d agoExploit / PoC1

Hackers Disable Endpoint Protection and Deploy Sliver Across Compromised Windows Domain

The Hunter's Ledger tracked campaign UTA-2026-024 using Sliver C2, Domain Admin account creation, and Ethereum-based C2 rotation to compromise a US organization's Windows domain.

The Hunter's Ledger tracked an intrusion at one unnamed US organization as UTA-2026-024, staged from exposed server 193.233.202.17 with a Sliver beacon. Operators created a non-expiring Domain Admin account, enabled RDP with NLA disabled, dumped SAM, SYSTEM and SECURITY hives plus LSASS memory, and disabled eight endpoint protection services. A Node.js implant resolved its C2 server from an Ethereum smart contract that rotated domains five times in five months, while SYSTEM scheduled tasks with backdated dates and DNS allowlist manipulation provided persistence. The infrastructure ties to a confirmed ransomware incident, but no encryptor deployment was proven in this intrusion.

Cyber Security News · 8d agoThreat actor in the wild

DPRK APTs: Ted backdoor and curlRAT target South Korean media and automotive sectors

Rapid7 uncovered a DPRK-linked Linux toolkit using a HAProxy-embedded ted backdoor, SSH keylogger, and curlRAT against South Korean media and automotive firms.

Rapid7 Labs identified a previously undocumented framework attributed with medium confidence to DPRK actors, targeting South Korean automotive and media organizations likely since early 2025. The toolkit embeds a backdoor compiled into HAProxy 2.8.12 using its filter API, plus trojanized crond, agetty, atd, sshd, and polkitd, an SSH keylogger storing credentials under /var/lib/sshd/, and a curl-based RAT with a watchdog thread. It enables remote command execution, malicious script injection into served webpages (a watering-hole loop), credential harvesting, and long-term surveillance. Hardcoded C2s are associated with APT37 via ThreatFox, and exposed groupware portals and mail servers align with Kimsuky tradecraft; the initial access vector and any CVE remain unconfirmed.

Rapid7 Blog · 12d agoThreat actor in the wild1

USN-8670-3: curl vulnerability

Ubuntu issued USN-8670-3 updating curl for Ubuntu 26.04 LTS to fix a flaw where wrong client certificates could be used on reused connections.

Ubuntu Security Notice USN-8670-3 extends the curl fix from USN-8670-1 to Ubuntu 26.04 LTS. The flaw, discovered by Joshua Rogers, involves incorrect handling of connection reuse when client certificate settings change, potentially causing the wrong client certificate to be presented. The issue can lead to authentication mix-ups rather than remote code execution.

Ubuntu Security Notices · 8d agoAdvisory

USN-8726-1: Linux kernel vulnerabilities

Ubuntu issued kernel security update USN-8726-1 fixing an Arm TLB invalidation flaw (CVE-2025-10263) that enables local privilege escalation, plus other kernel fixes.

Ubuntu released USN-8726-1, a security update for the generic Linux kernel. It fixes CVE-2025-10263, where certain Arm processors complete broadcast TLB invalidation before related memory writes are globally observed, potentially letting local attackers bypass memory protections or escalate privileges. The update also addresses additional kernel flaws in ARM64, ARM32, RISC-V, S390 and other subsystems.

Ubuntu Security Noticesupdated · 10d agofirst · 10d agoAdvisory 6 sourcesCVE-2025-10263

Akamai Valkey Managed Database: Real-Time Memory for Enterprise AI

Akamai launched Valkey Managed Database, a low-latency in-memory data layer aimed at cutting AI inference costs and accelerating RAG.

Akamai introduced Valkey Managed Database, a managed in-memory data service based on the open-source Valkey project. The company positions it as real-time memory for enterprise AI, optimizing inference costs, accelerating retrieval-augmented generation, and powering real-time AI agents.

Akamai Blog · 29d agoAI tools & infra1

USN-8757-1: cgit vulnerability

Ubuntu USN-8757-1 fixes cgit path-handling flaw letting remote attackers read files outside repositories during HTTP cloning.

Ubuntu Security Notice USN-8757-1 addresses a cgit vulnerability in which repository paths are incorrectly handled when HTTP cloning is enabled. A remote attacker could exploit the flaw to access files outside the repository and obtain sensitive information. The notice provides no CVE identifier or exploitation details.

Ubuntu Security Notices · 2d agoAdvisory

Fortunate Recall: Ontology-Driven Memory Lifecycle Management for Persistent Coherence in LLMs

Fortunate Recall introduces ontology-based lifecycle policies for LLM memory, cutting confabulation roughly in half (e.g., 45.1% to 22.4%) versus Mem0.

Fortunate Recall (FR) is a composable policy layer that classifies personal facts into a 10+1 behavioral ontology and applies category-specific lifecycle rules including differential temporal decay, slot-key supersession, event-time validity, and retrieval routing. FR-Bank scores 76.9% on the new 516-question LifecycleBench, ahead of Mem0, A-MEM, Memory-R1, and MemoryOS (61%-70.5%), and 75.2% on LongMemEval-S. End-to-end, confabulation drops from Mem0's 45.1% to 22.4% over answered queries, with the ranking replicating on open-weight Kimi K2.5 and transferring to the independent BEAM benchmark (46.8% vs 32.9%).

arXiv cs.AI / cs.LG / cs.CL · 7d agoAI research

commit-rewriter 0.1

Simon Willison released commit-rewriter 0.1, a tool that rewrites git commit messages from the first edited commit, with a timestamped revert branch.

Simon Willison built commit-rewriter 0.1, a small web app for editing git commit messages, motivated by cleaning up Datasette security release commits that contained coding agent cruft and private issue IDs. It runs via 'uvx commit-rewriter path/to/repo' and creates a timestamped branch of the repo state before rewriting every commit from the first edited one to the most recent, allowing easy reversion.

Simon Willison · 3d agoTools2

Claude is a Contrarian

Opinion piece argues Claude habitually contradicts explicit user instructions, injecting contrarian content despite CLAUDE.md rules and user objections.

A developer recounts repeated instruction-following failures with Claude, claiming it contradicts explicit requests, adds unnecessary work, and ignores AGENTS.md and CLAUDE.md directives. The author contrasts this with OpenAI, DeepSeek, and Qwen models, which he says more readily apologize and undo mistakes. He theorizes Claude's training makes it assume the human is wrong and needs correcting. The post is personal commentary with no benchmarks or systematic evaluation.

USN-8514-2: OpenSSH vulnerability

Ubuntu backports an OpenSSH fix to older LTS releases for an scp flaw enabling setuid file placement and privilege escalation.

USN-8514-2 extends the USN-8514-1 OpenSSH fix to Ubuntu 14.04 LTS, 18.04 LTS, and 20.04 LTS. The flaw stems from incorrect file permission handling when downloading files as root using the legacy scp protocol without the preserve-mode option. An attacker could exploit this to install setuid or setgid files on the system, potentially leading to privilege escalation.

Ubuntu Security Notices · 1d agoVulnerability

Give Your Coding Agents a Memory You Own

Hugging Face introduces Funes, a tool that gives coding agents persistent, self-owned memory outside vendor clouds.

A Hugging Face blog post presents Funes, an approach for giving coding agents a persistent memory that developers own and control. The piece targets agent workflows where context must survive across sessions without ceding data to third-party services. No article body was available in the feed, so specifics beyond the title are limited.

Hugging Face Blog · 14d agoAI tools & infra1