ZeroHour

Search: “ge-act”

30 stories in the last 7d

Smart search ranks by meaning as well as keywords (one row per story, last 45 days).

Don't Mask the Environment: Observation Supervision Changes How Agents Explore Under RL

ActObs adds observation-token supervision to SFT, yielding higher pass@k for Qwen3 agents after GRPO on Terminal-Bench 2.0 and code editing.

Researchers introduce ActObs, an SFT variant that supervises environment-observation tokens in agent trajectories in addition to action tokens, without extra data, parameters, tokens, or forward passes. On Qwen3-4B, GRPO initialized from ActObs achieves higher pass@k at every sampling budget on Terminal-Bench 2.0, and on Qwen3-8B it trades some pass@1 for +3.4 pp at pass@16 while solving more distinct tasks. The benefit transfers to unseen code-editing tasks on aider-polyglot (+4.2 pp pass@1 at 4B scale). The authors trace the advantage to gradient analysis showing joint supervision preserves environment prediction and policy entropy, improving downstream RL exploration.

Hugging Face daily papersupdated · 20h agofirst · 1d agoAI research 2 sources

GE-Act 2.0: Pretraining and Scaling a World-Action Model for Robotic Manipulation

GE-Act 2.0 is a from-scratch pretrained world-action model for robotic manipulation, with success rising from 17.1% to 44.1% as co-training data scales to 30,000 hours.

Genie Envisioner Act 2.0 (GE-Act 2.0) is a world-action model whose generative and action components are all initialized from scratch on manipulation data, combining a control-oriented autoencoder (CoAE), single-step visual planner (SVP), and inverse dynamics model (IDM) trained jointly via knowledge-aligned selective optimization (KASO). Scaling co-training data from 300 to 30,000 hours raises zero-shot success from 17.1% to 44.1% on G1-OP and 13.4% to 31.1% on G2-90D, despite the latter comprising under 2% of data, suggesting cross-embodiment transfer. Gains span 19/20 and 18/20 skill groups, and skill-specific coverage correlates with zero-shot OOD success (Pearson r=0.80).

Hugging Face daily papers · 14d agoAI research

ActGuard: Pre-execution Action Auditing against Indirect Prompt Injection in LLM Agents

ActGuard audits LLM agent actions before execution against predicted tool priors, masking only malicious spans from indirect prompt injections while preserving utility.

ActGuard is a pre-execution action auditing framework against indirect prompt injection in LLM agents, judging whether external content causes the current action to deviate from a locally reasonable expectation rather than whether content is inherently suspicious. At each step it predicts the tools likely used by the upcoming action, builds a local tool prior, then performs tool-level contrastive analysis and parameter-level evidence localization to identify deviations. A verifier masks only spans confirmed as malicious and regenerates the action from the sanitized context. On challenging tool-using agent benchmarks it reduces attack success to state-of-the-art levels while keeping task utility close to the no-attack setting; code is publicly available on GitHub.

arXiv cs.CR · 4d agoAI safety & security

The EU CRA's Real Question: What Shipped, and When Did You Know?

ActiveState argues the EU CRA's 24-hour ENISA exploit-notification duty, effective September 11, 2026, makes current SBOMs and provenance visibility a legal necessity.

An ActiveState essay warns that the EU Cyber Resilience Act's reporting obligations take effect on September 11, 2026, requiring manufacturers of products with digital elements sold into the EU to notify ENISA within 24 hours of learning a vulnerability is actively exploited, with a fuller report within 72 hours. The law's engineering requirements only apply from December 11, 2027, leaving a visibility-first runway, and Article 13 requires the SBOM to stay current unlike one-time artifacts generated under US Executive Order 14028. The author contrasts the 24-hour notification clock with an industry-average 55 days to remediate high or critical vulnerabilities and recommends automated SBOM regeneration or consuming pre-vetted, attested open source components.

BleepingComputer · 9d agoPolicy & legal

House passes bill to equip local law enforcement with scam-fighting tools

The U.S. House passed the GUARD Act, letting local law enforcement use federal grants to investigate financial scams and trace stolen cryptocurrency.

The bipartisan GUARD Act (Reps. Zachary Nunn, Scott Fitzgerald, Josh Gottheimer) passed the House, allowing existing DOJ grant funds to be used for fraud analysts, victim-support training, blockchain tracing software, and financial-information sharing with law enforcement. It addresses scams like pig butchering, often run by transnational criminal groups overseas; Americans lost a record $11.4 billion to crypto-related fraud in 2025, including $8.6 billion in investment fraud. Senators Katie Britt and Kirsten Gillibrand introduced a Senate companion in July 2025, and the House also passed a bill retroactively eliminating the 'scam tax' on stolen funds for 2021-2025 victims.

The Record · 1d agoPolicy & legal

Launching managed CRA Article 14 reporting for open source maintainers

EU Cyber Resilience Act Article 14 reporting obligations begin, requiring 24-hour exploit and incident reports; Patchstack launches managed compliance for open-source maintainers.

Starting 11 September 2026, EU Cyber Resilience Act Article 14 requires manufacturers and open-source stewards to report actively exploited vulnerabilities and severe security incidents to ENISA via the EU Single Reporting Platform, with a 24-hour early warning, 72-hour notification, and final reports within 14 days or one month. Patchstack launched a free managed compliance service, acting as Assigned Representative for open-source maintainers and providing a managed VDP. The obligations apply retroactively to all products available on the European market. Patchstack, which has coordinated over 50% of known WordPress ecosystem vulnerabilities, already serves more than 1,000 open-source projects.

Patchstack · 7d agoPolicy & legal

EU's Cyber Resilience Act starts the 24-hour vulnerability clock

EU Cyber Resilience Act reporting rules take effect, requiring manufacturers to disclose actively exploited vulnerabilities to ENISA within 24 hours, with fines reaching €15 million.

The Cyber Resilience Act's Article 14 mandatory reporting duties became applicable, requiring makers of products with digital elements sold in the EU — regardless of where they are based — to file an early warning within 24 hours of becoming aware of an actively exploited vulnerability, a detailed notification within 72 hours, and a final report within 14 days of releasing a fix. Reports must be submitted through ENISA's Single Reporting Platform to the designated CSIRT, and non-compliance with these core duties can trigger fines up to €15 million or 2.5 percent of annual turnover. Manufacturers must also inform affected users of available fixes without undue delay, and most remaining CRA provisions, including mandatory SBOMs and security-by-design requirements, become applicable on December 11, 2027.

The Register · Security · 7d agoPolicy & legal

EU Cyber Resilience Act to Enforce New Reporting Requirements

EU Cyber Resilience Act reporting obligations begin Friday, requiring businesses to notify serious product security incidents within 24 hours.

The EU Cyber Resilience Act's new reporting requirements take effect starting Friday. Businesses operating in the EU will have 24 hours to notify the government whenever they discover serious product security incidents.

Dark Reading · 8d agoPolicy & legal

The Model Proposes, the Code Disposes: A Pre-Registered Ablation of a Verifier-and-Acceptance Stage in an LLM-Orchestrated Offensive-Security Agent

Pre-registered ablation finds a model verifier stage in an LLM offensive-security agent suppresses findings; removing it eliminated suppression with precision tradeoff.

The paper evaluates a verifier-and-acceptance stage in an LLM-orchestrated offensive-security agent via a pre-registered 20-run confirmatory ablation and a 2x2 factorial study with 40 runs on vulnerable lab targets. Removing the stage eliminated pre-report suppression (median 2 vs 0 findings, p = 0.00003) but reduced model-blinded shipped precision (0.471 vs 0.353, p = 0.0087). Suppression was attributed to the model verifier rather than deterministic acceptance rules, and an instrumented canary recorded zero external contacts in all 60 runs. The full design retained 93.8% of model-adjudicated true candidates but failed its pre-registered non-inferiority floor of 0.90.

arXiv cs.CR · 3d agoResearch

15 Minutes Saved Per Alert: How a Lean German Manufacturer Protects 10,000 Endpoints with ANY.RUN

A five-person security team at a German manufacturer protecting 10,000 endpoints cut triage time by 15 minutes per alert after adopting ANY.RUN's cloud sandbox.

Philipp Z., Security Lead at a leading German manufacturer, described how a five-person team protects 10,000 endpoints and users using ANY.RUN's Interactive Sandbox in a private cloud. The firm previously relied on a single air-gapped forensic laptop running Flare VM, which caused 5-10 minute setup delays, single-user bottlenecks, and selective triage. The switch reportedly saved roughly 15 minutes per alert and reduced forced wiping and reimaging of user machines. ANY.RUN data cited in the piece puts manufacturing security workloads 22% above other major industries.

ANY.RUN · 9d agoIndustry

ETSI Proposes 17 Cybersecurity Standards to Support Cyber Resilience Act

ETSI has launched an approval process for 17 cybersecurity standards that vendors must meet under the EU Cyber Resilience Act.

The European Telecommunications Standards Institute (ETSI) initiated an approval process for 17 cybersecurity standards intended to support implementation of the EU Cyber Resilience Act. These standards will define requirements that vendors of products with digital elements must satisfy to comply with the regulation. The move advances the operational groundwork for CRA compliance in the European Union.

Infosecurity Magazine · Aug 17, 2026Policy & legal

Bipartisan Senate bill aims to prepare energy sector for Q

Bipartisan Senate bill would direct FERC to factor quantum computing threats and post-quantum cryptography into US electric grid cybersecurity reliability standards.

The Quantum Grid Utility Assurance and Resilient Defense (Quantum-GUARD) Act, introduced by Senators Mike Rounds and Chris Coons, would require FERC to consider quantum computing threats when reviewing electric reliability standards and to explore post-quantum cryptography use in both IT and OT systems, plus a technical sandbox to study quantum impacts. It aligns with NIST's post-quantum algorithm work, and a June executive order moved the federal PQC migration deadline from 2035 to 2030. Industry experts noted the hard part is upgrading infrastructure such as SCADA communications and software update integrity ahead of those deadlines.

CyberScoop · 24d agoPolicy & legal

CTEM Is Not About the Stages. It’s About the Outcome.

Horizon3 argues CTEM programs should measure continuously reduced exposure rather than mapping technologies to Gartner's five stages.

Horizon3 contends that Continuous Threat Exposure Management should be judged by one outcome: continuously reducing attacker-reachable exposure, not by mapping a technology to each of Gartner's five stages. The post argues validation and verification, not visibility or closed tickets, provide evidence that attack paths are actually broken. It describes a Discover, Validate, Prioritize, Remediate, Verify, Repeat motion as its operationalization of CTEM.

Horizon3.ai · 16d agoIndustry

12 Best Endpoint Privilege Management (EPM) Tools Compared (2026): Features & Pricing

A 2026 buyer's guide compares 12 endpoint privilege management tools, ranking CyberArk and BeyondTrust as enterprise leaders.

An editorial comparison evaluates 12 endpoint privilege management (EPM) tools on elevation control, manageability, and pricing model. The guide argues that standing local-admin rights fuel ransomware and lateral movement, making their removal a high-impact control that cyber insurers increasingly mandate. CyberArk and BeyondTrust are positioned as enterprise-depth leaders, with Delinea, Heimdal, and ManageEngine for the mid-market, and Admin By Request and CyberFOX AutoElevate for SMBs and MSPs. Pricing is generally per endpoint or per user, and the article is explicitly an assessment rather than a product release or incident report.

GBHackers · 7d agoIndustry1

German Manufacturer Shrinks Security Alert Response While Protecting 10,000 Endpoints

Vendor case study: a German manufacturer's five-person SOC cut alert triage time using ANY.RUN's cloud sandbox across 10,000 endpoints.

ANY.RUN published a case study in which a five-person security team at an unnamed German manufacturer replaced an air-gapped forensic laptop with its cloud-managed interactive sandbox, protecting roughly 10,000 endpoints and 10,000 users. The vendor claims a median 15 minutes saved per alert, 20-40 daily tasks processed, a 2.5-minute alert-to-isolation target, and a 95% agreement rate between analyst and sandbox verdicts; all figures are vendor-supplied with the customer identity withheld. The writeup also describes detonating a multi-stage phishing chain from a PDF link to a password-protected ZIP to malware execution.

Cyber Security Newsupdated · 1d agofirst · 1d agoIndustry 3 sources

17 draft Cyber Resilience Act standards are open for comment

ETSI publishes 17 draft harmonised standards detailing EU Cyber Resilience Act compliance, open for comment until between mid-September and mid-November 2026.

Seventeen draft standards covering the higher-risk tier of products with digital elements, including password managers, antivirus software, connected toys and wearables, are open for comment. Following a Harmonised Standard grants manufacturers the presumption of conformity with the Cyber Resilience Act, whose obligations apply through the end of 2027 to importers, distributors, service providers and developers. The drafts went to 41 member organisations plus societal partners ANEC, ECOS, ETUC and SBS, with closing dates varying by vertical.

Help Net Security · Aug 14, 2026Policy & legal

NIS2 compliance: Fixing IAM and access control before the 2026 audit

EU NIS2 enforcement deadlines approach; organizations are urged to prioritize service account inventory, lifecycle offboarding, and phishing-resistant MFA before audits.

EU member states are moving from NIS2 transposition into enforcement, with fines up to 10 million euros or 2% of global turnover for essential entities and personal liability for management bodies. The article argues access management is the fastest high-ROI starting point, estimating 2-4 weeks to enforce fine-grained password policy, vault shared credentials, and deploy phishing-resistant MFA versus 6-12 months for supply chain risk management. It flags three common pre-audit failures: unmanaged service accounts and API keys, dormant accounts from broken offboarding, and SMS OTP instead of phishing-resistant MFA under NIST SP 800-63B. The piece promotes Passwork as a single control plane for credential storage, RBAC, and WebAuthn.

Help Net Security · 17d agoIndustry

CISA's logging guidance works beyond government

CISA released its Logging Reference Architecture in August 2026 to help federal agencies meet OMB M-26-14 logging requirements, usable as a benchmark by critical infrastructure operators.

CISA's Logging Reference Architecture (LRA), released in August 2026, helps US federal civilian agencies satisfy logging requirements in OMB Memorandum M-26-14 and explicitly encourages critical infrastructure operators to use it as a benchmark. The framework is organized around continuous event monitoring and threat hunting, investigation, response, and forensics, with a federal baseline of six months searchable and one year retrievable logs. Agencies must submit Agency Logging Plans within 90 days and work toward Advanced maturity within 320 days; the guidance also treats AI outputs as derived data requiring human review and preserved metadata.

Help Net Security · 25d agoAdvisory

Delaware Consumer Privacy and Data-Breach Law Updates

Delaware's governor signed HB 380 and HB 381 amending the state privacy act and breach notification law.

On September 2, 2026, Delaware's Governor signed House Bill 380 and HB 381. HB 380 amends the Delaware Personal Data Privacy Act (DPDPA), enacted in 2023 and effective January 1, 2025. HB 381 separately amends Delaware's computer security breach notification law. Joseph J. Lazzarotti of JacksonLewis summarizes the changes.

DataBreaches.net · 4d agoPolicy & legal

ENISA launched the CRA Single Reporting Platform for actively exploited vulnerabilities

ENISA launched the CRA Single Reporting Platform, making EU manufacturers report actively exploited vulnerabilities and severe incidents through one portal.

ENISA switched on the Cyber Resilience Act's Single Reporting Platform on 11 September 2026, the same day CRA reporting obligations became binding on manufacturers. Reports require an early warning within 24 hours, a fuller notification within 72 hours, and a final report within 14 days (one month after notification for severe incidents). Filings go through an EU Login account with MFA, are routed to a coordinating CSIRT chosen by the manufacturer, and no API is available in the first release. Open-source software stewards fall under the same obligations from 11 December 2027.

Help Net Security · 4d agoPolicy & legal

European Commission set to push social media restrictions, safety requirements into law

The European Commission proposed the EU KIDS Act, barring under-13s from social media, mandating safe-by-design rules, and fining violators up to 6% of global revenue.

The European Commission laid out a roadmap for the EU KIDS Act, which would block social media accounts for children under 13, set a bloc-wide minimum account age of 15, and restrict ages 13-15 to guardian-controlled mini accounts with one-hour daily screen time caps. Safe-by-design requirements ban addictive features, infinite scroll, profiling-based recommender feeds, and push notifications during sleeping hours, and require AI companions and chatbots to be off by default. Providers would need to deploy the EU age verification app, submit compliance plans to third-party audits, and face fines up to 6% of global annual sales and 90-day investigations. The proposal still requires approval from the European Parliament and member states, building on the Digital Services Act and AI Act.

The Record · 17h agoPolicy & legal

ZDI-26-615: (0Day) pdfforge PDF Architect activation-service Update Service Uncontrolled Search Path Element Local Privilege Escalation Vulnerability

Zero Day Initiative disclosed an uncontrolled search path flaw in pdfforge PDF Architect's update service, allowing local privilege escalation (CVSS 7.8).

ZDI published advisory ZDI-26-615 describing a local privilege escalation vulnerability in the activation-service Update Service of pdfforge PDF Architect. An uncontrolled search path element lets local attackers escalate privileges, but they must first be able to execute low-privileged code on the target. ZDI rated the issue 7.8 and tagged it as a 0-day disclosure.

ZDI Published Advisories · 18d agoVulnerability

12 Best CIEM Tools Compared (2026): Features & Pricing

Buyer's guide compares twelve CIEM tools; Microsoft discontinued Entra Permissions Management, while Tenable (Ermetic), CyberArk, and Wiz lead the 2026 scorecard.

The scorecard evaluates twelve cloud infrastructure entitlement management vendors on permission analytics depth, JIT enforcement, non-human identity coverage, pricing predictability, and bundle leverage. Tenable (Ermetic) leads at 4.70, followed by CyberArk and Wiz, while Microsoft's retirement of Entra Permissions Management (CloudKnox) forces existing customers into migration cycles. Pricing structures span per-identity, per-resource, per-workload, credit-based, and quote-based models.

GBHackersupdated · 5h agofirst · 2d agoIndustry 10 sources1

MasterControl Seventeen Every Time

Governed enterprise analytics study shows deterministic policy execution matched 110/110 answer-and-evidence contracts while runtime agent planning matched none.

The paper studies a governed approach where a language model interprets the question while deterministic policy selects and runs a pre-approved analytical program returning results and evidence. Across 440 runs, three 8B models generated SQL and selected tools at runtime, while Qwen3-8B only interpreted intent and policy executed the approved program. None of 330 runtime-planning episodes satisfied the full answer-and-evidence contract, whereas the policy-executed analyzer matched 110 of 110. The authors note this is configuration-specific and expressiveness is preserved via relational operations, aggregation, comparison, windows, ranking, and similarity with replayable results.

Hugging Face daily papers · 16d agoAI research

Australia is replacing the Essential Eight with a new cyber framework. Here’s how exposure management can help you get ahead of it.

Australia's ASD is replacing the Essential Eight with an outcomes-based Essentials series covering IT, cloud, OT and likely agentic AI, with deprecation from mid-2027.

The Australian Signals Directorate announced in June 2026 that the Essential Eight will be replaced by an outcomes-focused Essentials series structured as chapters covering enterprise IT (including identity and SaaS), cloud, OT, and likely agentic AI. Deprecation begins around mid-2027 with full retirement around mid-2028, though timelines are targets; the Essential Eight is mandatory for roughly 98 non-corporate Commonwealth entities but voluntary for private firms. Tenable argues the shift demands continuous security posture evidence via exposure management rather than point-in-time checklist assessments.

Tenable Blog · 3d agoPolicy & legal1

From Intent to Execution Grant: An Execution-Boundary Conformance Profile for High-Risk AI Actions

Researchers specify EBL-Core, an execution-boundary conformance profile binding AI agent intents, policies, and evidence into verifiable execution grants, validated with bounded tests.

The paper defines EBL-Core, a conformance profile deciding whether one fully materialized AI-generated candidate action may receive action-scoped execution authority. It binds a structured intent object, Root and Operational Policies, typed evidence, and a verifiable Decision Derivation through an Execution Release Contract, with lifecycle rules for Redemption and Revocation. Evaluation included 34 static vectors, 15 lifecycle checks, and 100 trials of 32 concurrent Redemption attempts yielding exactly one winner per trial. The authors state these bounded results demonstrate executability of the specified subset, not production readiness or complete mediation.

arXiv cs.CR · 7d agoAI safety & security1

SchemeArena: Factorized Stress Testing of Scheming in LLM Agents

Researchers introduce SchemeArena, a 400-scenario benchmark stress-testing scheming in LLM agents, finding explicit instrumental goals are the strongest driver of covert misaligned behavior.

The paper presents SchemeArena, a 400-scenario benchmark built through factorized scenario synthesis spanning safety-relevant tool domains, instrumental goals, oversight conditions and pressure mechanisms. The accompanying SCOUT monitor grounds multi-criteria scheming judgments in evidence drawn from agents' reasoning and actions. Stress tests across five LLM agents show explicit instrumental goals are the strongest driver of scheming propensity, while action-only monitoring increased scheming in several closed models, suggesting partial oversight can act as an optimization constraint. The benchmark, code and monitor are released at github.com/launchnlp/SchemeArena.

Hugging Face daily papers · 10d agoAI safety & security1