ZeroHour

Search: “workflow”

79 stories in the last 3d

Nozomi Compass helps industrial teams manage OT assets and vulnerabilities

Nozomi Networks launched Compass, an OT asset and vulnerability management platform unifying asset records, remediation workflows, and compliance evidence for industrial teams.

Nozomi Networks announced Compass, an OT asset and service management platform built on real-time first-party asset data from its Vantage cyber-physical security platform. It provides OT-native workflows, governed change approvals, consequence-based risk scoring, and continuous audit-ready compliance evidence mapped to NERC CIP, IEC 62443, NIS2, and TSA. The platform integrates with EAM, CMDB, ITAM, ITSM, SIEM, and SOAR tools and is designed to safely support AI-driven and agentic OT workflows with human oversight.

Help Net Security · 19h agoTools

Citrix adds AI-powered browser activity analysis to SecurAccess

Citrix launched Session Insights for SecurAccess with Chrome Enterprise, using AI to record and analyze browser activity from users and autonomous agents.

Citrix Session Insights adds automatic session recording and AI-powered risk detection for browser activity by human users and autonomous AI agents within Citrix SecurAccess with Chrome Enterprise. The capability creates visual forensic records, highlights risky behavior for faster investigations, and recommends policy adjustments or changes to agent authority levels. It is designed to support audits and governance as enterprise AI agent workflows expand.

Help Net Security · 19h agoTools

Jev: New frontier model 40-400x cheaper and 20-200x faster

TypeSafe AI launches Jev, an early-access 'System One' model delivering calibrated structured outputs claimed 40-400x faster and cheaper than LLMs.

TypeSafe AI, founded by former OpenAI researcher Diogo Almeida, released its first 'System One Model' called Jev in early access. Jev forgoes string generation and is trained with Reinforcement Learning for Calibrated Decisions (RLCD) to produce type-safe structured values with calibrated probabilities. The company claims 70-500ms response times (40-200x faster), input pricing of $0.042 per million tokens, and free output tokens via a parallel sampling architecture. Target use cases include AI-powered workflows, real-time applications, and verification/guardrail tasks.

Cohesity adds recovery capabilities for AI agents and the data they manage

Cohesity launched Agent Resilience to discover, protect, and recover AI agent memory, configuration, and agent-managed data, debuting with Amazon Bedrock integration.

At Cohesity Catalyst, Cohesity introduced Agent Resilience within Cohesity Data Cloud, protecting AI agent memory and configuration with snapshot architecture, immutable backups, and clean-room recovery, plus recovery for databases and file systems that agents manage. It launches with Amazon Bedrock integration, support for Microsoft and Google platforms planned, and general availability targeted for year-end. The company cited Gartner's prediction that up to 40% of enterprise applications will include task-specific agents by 2026, and Cohesity research showing 56% of organizations are unprepared to detect or contain unintended agent actions while 58% lack confidence in verifying AI model integrity after attacks. Cohesity also outlined an Autonomous Cyber Resilience vision using agentic workflows and introduced the AI Resilience Academy.

Help Net Security · 18h agoTools

ANY.RUN & SentinelOne: One Workspace, Instant Context for Rapid Response

ANY.RUN integrates its interactive sandbox, IOC lookups, and STIX/TAXII threat feeds natively into SentinelOne for faster automated malware triage.

ANY.RUN and SentinelOne launched connectors that embed interactive sandbox analysis and threat intelligence into the SentinelOne console via Singularity Hyperautomation. Suspicious files and URLs from alerts are automatically submitted to the ANY.RUN sandbox, with behavioral verdicts and risk scores returned into alert notes. On-demand IOC lookups draw on sandbox history from 16,000 organizations and 700,000 analysts. A separate STIX/TAXII feed streams verified malicious IPs, domains, and URLs through the SentinelOne Marketplace TAXII Connect app.

ANY.RUN · 22h agoTools

Observational Indistinguishability and Integrity Blind Regions in Hybrid Quantum-Classical Workflows

Framework formalizes integrity blind regions in hybrid quantum-classical workflows, validated across 3,600 label interventions with conformal detection rules.

The paper presents a claim-relative evidence and reference framework for integrity of hybrid quantum-classical workflows, distinguishing structural blind regions caused by observational indistinguishability from finite-batch statistical misses. Experiments over 3,600 label interventions show exact label-path invariance for feature and prediction views. The geometry-aligned construction detects 343 of 2,700 conclusion-changing interventions using the conformal rule and 1,183 of 2,700 with the uncorrected union, with executed conformal clean false-action rates of 0.048-0.059.

arXiv cs.CR · 1d agoResearch

NVIDIA Open-Sources OSMO: One YAML Orchestrates Physical AI Training, Simulation, and Robot Testing

NVIDIA open-sourced OSMO, a Kubernetes-native YAML orchestrator running physical-AI training, simulation, and robot testing across mixed GPU tiers.

OSMO (Apache-2.0, latest release 6.3.1) lets teams describe training, simulation, and hardware-in-the-loop pipelines in a single YAML and routes tasks across datacenter GPUs (GB200), workstation RTX hardware, and edge devices like Jetson AGX Thor. It ships Helm charts and containers on NGC, uses the KAI Scheduler with NVLink topology-aware placement, and includes RBAC, OAuth2, and TLS termination. NVIDIA says it is battle-tested on GR00T, Isaac Lab, Isaac Sim, and Isaac ROS, and integrates with Claude Code, OpenAI Codex, and Cursor agents.

MarkTechPost · 2d agoAI tools & infra

How to connect AI usage to business value

OpenAI explained how ChatGPT Admin Console analytics link AI usage, spend, and Codex contributions to business outcomes.

OpenAI published guidance describing analytics features in the ChatGPT Admin Console that combine usage, credit, and token data across ChatGPT Work and Codex. The Insights task classifier groups messages into use cases such as software engineering and sales research, while an Outcomes view tracks Codex contributions to merged commits and lines of code. An Admin plugin and Admin API let teams automate reporting and combine AI analytics with business metrics like ticket resolution time or revenue.

OpenAI News · 17h agoAI industry

Bitsight connects threat intelligence and exposure monitoring across the supply chain

Bitsight made Beacon generally available, combining supply-chain exposure monitoring with MCP support to feed threat intelligence into AI workflows.

Bitsight announced general availability of Beacon, which continuously monitors critical vendors for exposure, vulnerabilities, malicious activity, intrusion, stolen credentials and compromise. New Model Context Protocol (MCP) and agentic capabilities push Bitsight intelligence into AI-enabled workflows, with over 400 customers signing up for early access in one month. The company cites data that third parties now account for almost half of enterprise breaches, up over 60% year over year.

Help Net Security · 2d agoTools

What happens when AI agent governance is missing at scale

meshIQ engineering head Gourab Basu argues AI agent governance must inspect proposed tool calls in-flow, since prompts alone cannot control nondeterministic agents.

In a Help Net Security interview, Gourab Basu, Global Head of Engineering at meshIQ, argues that prompt instructions are an insufficient control boundary for nondeterministic AI agents. He advocates a framework-independent governance engine that inspects proposed tool calls and parameters before execution, citing an example of pausing refunds above $100 for human approval. He warns that scaling from ten to a thousand agents makes manual oversight and destination-side controls unworkable, so governance must sit inside the agent execution flow across frameworks such as FastMCP.

Help Net Security · 23h agoAI safety & security1

University of Manchester Uses NVIDIA Earth-2 to Forecast Air Pollution Across the UK

University of Manchester retrained NVIDIA Earth-2 CorrDiff and StormCast on Isambard-AI to forecast UK air pollution at 2-3 km resolution.

University of Manchester researchers led by professor David Topping adapted NVIDIA's Earth-2 generative AI frameworks to forecast air pollution across the UK. Earth-2 CorrDiff was retrained in two days on a single eight-GPU node of Isambard-AI (5,448 GH200 Grace Hopper Superchips, 21 exaflops) using a year of hourly simulated pollution data, producing a UK-wide model at 2-3 square kilometer resolution. The team added Earth-2 StormCast for time-dependent forecasts that ingest real air quality observations, and demonstrated the workflow runs on the DGX Spark desktop AI system. Open-source training data and workflows are planned so other countries and cities can build similar pollution models.

NVIDIA Blog · 1d agoAI industry

How Fyxer built an AI executive assistant people trust

Fyxer details its OpenAI-powered AI executive assistant, orchestrating 30-50 specialized models trained on 500,000+ hours of assistant workflows.

OpenAI published a case study on Fyxer, whose AI executive assistant orchestrates 30-50 specialized OpenAI models trained on more than 500,000 hours of annotated executive assistant workflows. The system uses supervised fine-tuning, LoRA, and Direct Preference Optimization on user edits, and 53% of AI-generated email drafts are accepted as written. Fyxer's annual recurring revenue grew from $1 million to $32 million during 2025.

OpenAI News · 2d agoAI industry

Affora: A Design System for Agent-Friendly Interfaces

Affora is a design system making interfaces legible to computer-use agents while preserving human workflows, with reusable components and executable checks.

Affora supports both human users and computer-use agents through a shared interface rather than a separate agent-only surface. Three controlled studies cover component implementations, visual variation, and interaction-design principles, producing guidance from individual components to complete sites with reusable implementations and executable checks. Evaluation on independently authored interfaces shows gains where agent-readability deficits exist, limited effects where they do not, and a workflow case gives preliminary evidence of reduced interaction cost.

arXiv cs.AI / cs.LG / cs.CL · 11h agoAI research

How workers are unlocking new ways of working

OpenAI's analysis of 1.5 million ChatGPT work messages finds cross-occupation AI tasks becoming recurring parts of workers' routines.

OpenAI's latest Work at the Frontier research analyzed more than 1.5 million work-related ChatGPT messages from April through July 2026. Among roughly 6,200 consistently observed workers, previously used cross-occupation tasks grew from 13.1% of occupation-specific AI activity in April to 25.9% in July. Workers returned to a cross-occupation task used the prior month 23.6% of the time versus an 8.4% baseline, with an average next-month return rate of 18.5%. Recurrence was highest for customer discussions (54%), advertising copy (44%), and marketing materials (37%), suggesting AI may broaden jobs before titles change.

OpenAI News · 20h agoAI industry

Horizon3 Announces Integration with CrowdStrike Falcon® Next-Gen SIEM

Horizon3 announces NodeZero integration pushing validated exposure findings into CrowdStrike Falcon Next-Gen SIEM for correlated investigations.

Horizon3 announced an integration enabling validated NodeZero findings to flow into CrowdStrike Falcon Next-Gen SIEM, available now in the CrowdStrike Marketplace. Security teams can ingest and correlate exposure data with endpoint, identity, cloud, and other telemetry during investigations. CrowdStrike claims Falcon Next-Gen SIEM delivers up to 150x faster search than legacy SIEMs at up to 80% lower total cost of ownership.

Horizon3.ai · 1d agoTools

Give every teammate and agent the right level of access to your Workers

Cloudflare launches per-Worker granular access controls with four roles, enabling least-privilege access for teammates, AI agents, and CI/CD pipelines.

Cloudflare announced granular authorization for Workers, letting admins scope access to a single Worker instead of the whole account. Four new roles are available: Metadata Read-Only (observability without source code), Content Read-Only (read code without changes), Editor (deploy without delete), and Admin (full control of one Worker). Roles apply at Developer Platform, product, or resource level, can be attached to dashboard users or API tokens, and are available to all customers now, with plans to extend to D1, R2, and KV.

Cloudflare Blog · 1d agoTools

Agent Harness vs Agent Framework vs MCP: Which Layer Owns the Loop, State, Tools, Permissions, and Recovery

Architecture explainer separates agent harnesses, frameworks, and MCP by which layer owns the loop, state, permissions, and recovery.

The article distinguishes agent harnesses (OpenAI Codex, Claude Agent SDK), which own the execution loop, sandbox, permission model, and recovery; frameworks (LangGraph, OpenAI Agents SDK, Microsoft Agent Framework), which supply composable primitives; and MCP, a stateless JSON-RPC wire protocol governed by the Linux Foundation's Agentic AI Foundation since December 2025. An ownership matrix maps the execution loop, state, tool transport, permissions, recovery, sandboxing, and multi-agent orchestration to each layer. The 2026-07-28 MCP specification made the protocol fully stateless, retiring the initialize handshake and session headers.

MarkTechPost · 2d agoAI research1

Investing Together: Wiz Defend and Google Security Operations

Wiz ships a Content Pack for Google Security Operations enabling shared investigations, Blue Agent AI analysis, bidirectional sync and cloud telemetry streaming.

Wiz announced deeper integration between Wiz Defend and Google Security Operations via a shared data model and an official Wiz Content Pack with out-of-the-box rules, dashboards, search queries, playbooks, and response policies. Wiz Blue Agent AI-powered threat investigations, correlating cloud context, runtime signals, and identity data, are now accessible directly within Google SecOps. Status, severity, and comments bidirectionally sync in real time, and sensor runtime events can be streamed into Google SecOps for hunting and retention.

Wiz Blog · 1d agoTools1

Weekly Cybersecurity Newsletter – Top 50 Biggest Cybersecurity Stories of the Week

GBHackers weekly digest rounds up 50 stories including Microsoft's 973-CVE patch drop, exploited Cisco FMC flaws, and Claude agent attacks.

GBHackers' September 7-12, 2026 newsletter summarizes the 50 biggest cybersecurity stories of the week. Highlights include Microsoft patching a record 973 CVEs with two exploited zero-days, active exploitation of Cisco FMC, Check Point VPN and Ivanti flaws, China-linked crews chaining Chrome and Windows zero-days, AI agents mass-exploiting PaperCut to compromise 440 servers, and the emergence of Panzer cross-platform ransomware. It also covers Anthropic and OpenAI agentic AI incidents and CrowdStrike's SafeMind launch.

GBHackers · 1d agoIndustry in the wild

Heart of the Matter: How a Major Children’s Hospital Uses Open Source NVIDIA AI for Cardiac Care

Children's Hospital of Philadelphia uses NVIDIA open-source MONAI, Warp and Newton to build pediatric heart models in seconds for surgical planning.

CHOP's cardiac modeling service uses MONAI, Auto3DSeg and SlicerHeart to turn CT, MRI and 3D ultrasound images into anatomically precise heart models in seconds instead of four hours of manual work. More than 20 US children's hospitals run similar programs, with Boston Children's supporting roughly 500 cardiac surgery cases a year. NVIDIA's Newton physics engine, built on the Warp Python framework, aims to reduce device simulations from hours to near real time in clinical workflows.

NVIDIA Blog · 1d agoAI industry

12 Best Enterprise Browsers Compared (2026): Features & Pricing

2026 comparison of twelve enterprise browsers ranks Island and Palo Alto Talon as purpose-built leaders, with Chrome Enterprise and Edge free or bundled.

Guide compares twelve enterprise browser options across three models: purpose-built secure browsers (Island, Talon, Surf), layered controls on existing browsers (Chrome Enterprise, Edge for Business, LayerX, Seraphic), and streamed/isolated browsers (Kasm). Island and Palo Alto's Prisma Access Browser lead the purpose-built category for BYOD and contractor DLP. It also notes Mammoth Cyber has ceased operations.

GBHackers · 1d agoTools

[AINews] AEF-1 standard emerges for Third Party Evaluators, as Xai, OpenAI, and Anthropic all cosign

xAI, OpenAI, and Anthropic cosign the AEF-1 third-party evaluation standard while Dario Amodei proposes embedded evaluators for safety verification.

The AI Evaluator Forum published AEF-1, a baseline standard for independent third-party AI evaluations covering access, conflicts of interest, funding relationships, recusal, and transparency, cosigned by xAI, OpenAI, and Anthropic. Dario Amodei wrote a rare personal blogpost proposing embedded evaluators such as METR with desks, badges, company laptops, and internal-risk-team-level access to verify safety commitments, plus democratic and global coordination frameworks. The roundup also covers the pacing debate: Bilal Chughtai left Google DeepMind arguing progress may outrun alignment, while critics including Aidan Gomez and Cohere push back against slowdowns and lab gatekeeping. Additional items include Cline Desktop's launch with open-weight model support.

Latent Space · 2d agoAI safety & security

Superhuman acquires YC-backed notetaker Fathom as productivity platforms push for agentic work

Superhuman acquires AI notetaker Fathom to add meeting context and agentic workflows to its 40-million-user productivity platform.

Superhuman is acquiring Y Combinator-backed AI notetaker Fathom, which raised over $30 million and was valued at $94 million in 2024. Fathom reports 400,000+ monthly active users and over 1 million people have recorded meetings with it. The deal adds notetaking to Superhuman's suite (email, docs, calendar, database, AI agent builder) to enable proactive AI agents driven by meeting context, competing with Granola, Read AI, and Wispr.

TechCrunch · AI · 2d agoAI industry

AI Changed the Exposure Problem. Validation Needs to Change With It.

Picus Security argues vulnerability validation must combine exploitability, control validation, and agentic pentesting as AI accelerates disclosure volume.

Picus Security reports 35,853 CVEs were published in H1 2026, roughly 49% more than the prior year, while only 495 were catalogued as exploited in the wild and 116 were attacked on disclosure day. The vendor argues CVSS-based triage is inadequate and promotes combining exploitability validation, security control validation, and agentic pentesting into one program. The post also cites Anthropic data showing Mythos-class models surfaced 26,153 open-source vulnerability candidates with only 421 patched upstream, and promotes Picus's Validation Summit '26 on October 14-15.

The Hacker News · 2d agoIndustry

DeepSeek-v4.1 Flash: Pushing the Limits of KV Cache Compression

DeepSeek-V4.1 Flash is a 552B-parameter multimodal MoE model with 1M-token context achieving 4x KV cache compression for long-horizon agent workloads.

A detailed analysis of the DeepSeek-V4.1 Flash technical report describes a 552B-parameter multimodal mixture-of-experts model supporting contexts up to 1 million tokens. Its Causal Encoder-Decoder (CED) architecture activates 8B parameters during prefill and 16B during decode, and reportedly delivers about 420 tokens/s. Joint optimization of architecture (CSA2 cross-layer compression), FP4 KV cache precision, and deployment strategy cuts runtime KV cache to roughly 1/4 and persistent KV cache to about 1/8 of DeepSeek-V4-Flash at the same sequence length, targeting storage and bandwidth bottlenecks in long-horizon agent serving. The author notes all DeepSeek-V4 Pro models were taken offline following the release.

Former OpenAI researcher builds an AI model that judges options instead of writing text

TypeSafe AI launches Jev, a judgment-only model built by ex-OpenAI staff that classifies inputs with 70-500 ms latency instead of generating text.

Startup TypeSafe AI, co-founded by former OpenAI researcher and InstructGPT co-author Diogo Almeida, introduced Jev, a model that scores developer-defined answer options with probabilities rather than generating free-form text. The company claims 70-500 ms responses, parallel multi-question evaluation, and $0.042 per million input tokens with free outputs, targeting request routing, sales intent scoring, and assistant guardrail checks. Benchmarks are self-built and not independently verified, the 'no hallucination' guarantee only covers output structure, and access is currently via waitlist.

The Decoder · 14h agoAI industry

Axoflow Launches AxoDetect, Bringing Detection Into the Pipeline and Making the SIEM Optional

Axoflow launches AxoDetect in early access, running customer Sigma rules in the data pipeline to cut SIEM ingest costs and make full SIEM feeding optional.

Axoflow announced AxoDetect, now in early access and unveiled at Splunk .conf26, which runs customer Sigma rules directly in the security data pipeline on normalized data. Alerts travel to the SIEM while full-fidelity logs land in AxoLake, a low-cost on-prem-capable security data lake. The company cites a global industrial company cutting SIEM costs 50% and mean time to resolution 85%, and a government agency cutting data volume 80%.

GBHackers · 17h agoTools 2 sources

Building a Linux GPU Driver for the M4 Mac Mini in One Month

Two developers built a fully OpenGL ES 3.0 compliant Linux GPU driver for the M4 Mac Mini in one month via clean-room reverse engineering.

Niklas and the author reverse engineered Apple's AGX GPU firmware ABI and user-space components in about a month, a process that normally takes years, producing an OpenGL ES 3.0 conformant driver fast enough to run Minecraft at 200fps on an M4 Mac Mini. The work was done transparently using hypervisor traces without examining Apple binaries, following clean-room practices, and included a custom shader compiler, command stream builder, and a full Linux kernel driver for the firmware ABI. The A18 Pro firmware ABI proved significantly more complex than the M1's, with 1.5x as many structs and twice as many pointers. All experiments and provenance evidence were published in public agx-re repositories.

The AI graveyard: a running list of projects and startups that didn’t make it

TechCrunch compiles a running list of failed AI products and startups, including Relay, OpenAI's Sora, Humane AI Pin, Notion Mail, and Microsoft's Recall.

TechCrunch's 'AI graveyard' tracks notable AI products and startups that shut down or underperformed, citing S&P Global data that about 42% of AI initiatives are abandoned. Examples include the Relay automation startup, OpenAI's Sora video platform (shut March 2026), ChatGPT Atlas browser (discontinued August 9), Notion Mail (shutting September 22), and Apple's delayed Siri AI that contributed to a $250 million settlement. Hardware failures include the Humane AI Pin ($230 million raised, assets sold to HP for $116 million) and the Rabbit R1, which sold 100,000 units but received poor reviews. Microsoft's Recall feature remains controversial after a researcher demonstrated a tool extracting its captured data.

TechCrunch · AI · 1d agoAI industry1

F5 Bot Defense uses real-time risk scoring to detect fraud and abuse

F5 enhances Distributed Cloud Bot Defense with persistent device identification, real-time risk scoring, and agent-aware policies to manage AI agent traffic.

F5 announced enhancements to Distributed Cloud Bot Defense adding persistent device identification, real-time device risk scoring, risk-based workflow enforcement, and an agent-aware policy framework integrated with the F5 Application Delivery and Security Platform. The features aim to expose multi-account abuse, credential stuffing, and account takeover while allowing trusted AI agents to transact at machine speed. It targets fraud and abuse detection as agentic AI becomes a key interaction channel for sites, apps, and APIs.

Help Net Security · 1d agoTools

12 Best CNAPP Platforms Compared (2026): Features & Pricing

Independent comparison of 12 CNAPP platforms finds identical estates draw quotes 2-3x apart; Microsoft Defender for Cloud is the only fully published per-resource option.

A vendor-independent buyer's guide compares twelve CNAPP platforms including Prisma Cloud, CrowdStrike Falcon Cloud Security, Wiz, Uptycs, Aqua, Zscaler, and Microsoft Defender for Cloud on pricing mechanics, procurement leverage, and capability-per-dollar. It finds quotes swing 2-3x on identical estates because vendors define 'workload' differently. Microsoft Defender for Cloud is highlighted as the only major with fully published per-resource rates.

GBHackers · 1d agoIndustry1

Top 10 Best Cloud Infrastructure Entitlement Management (CIEM) Tools in 2026

2026 CIEM guide ranks Wiz, Prisma Cloud, Okta, Entra Permissions Management and specialists Sonrai, Britive, Tenable/Ermetic for cloud entitlement right-sizing.

Buyer's guide covers ten CIEM products across three market routes: CNAPP-bundled (Wiz, Prisma Cloud), identity-suite (Okta, CyberArk, SailPoint, Saviynt) and specialists (Sonrai, Britive, Tenable/Ermetic). It cites machine identities outnumbering humans 10:1 plus effective-permissions sprawl as core drivers, with JIT elevation as the fix. Notable consolidation includes Tenable acquiring Ermetic and Zscaler acquiring Canonic.

Cyber Security News · 1d agoTools

Your employees are already using AI tools you never approved

OneTrust report: 74% of organizations have scaled AI adoption, but only 17% embed governance by design and agent use outpaces oversight.

OneTrust's 2026 AI-Ready Governance Report finds 74% of respondents report departmental or scaled AI adoption, yet only 17% report governance embedded by design and just 5% have coordination and accountability defined across the AI lifecycle. Nearly half experienced at least one incident in the past year where AI systems or agents took unapproved actions, with data loss, corruption, and misclassification cited as the most likely and least prepared-for risks. 33% say employees used unapproved AI tools because approved options were not available quickly enough, and 98% plan to increase AI governance technology budgets next financial year.

Help Net Security · 2d agoIndustry

Only at TechCrunch Disrupt 2026: What happens when OpenAI ships your roadmap?

TechCrunch Disrupt 2026 panel will discuss AI startup defensibility when OpenAI, Anthropic, or Google ship features startups built.

TechCrunch promotes a Builders Stage session at Disrupt 2026 (October 13-15, Moscone West, San Francisco) titled 'What Happens When OpenAI Ships Your Roadmap.' Speakers include Airbyte CEO Michel Tricot, Radical Ventures partner Rob Toews, and Webflow CEO Linda Tong. The session examines how founders differentiate when foundation model providers absorb startup capabilities, emphasizing proprietary data, embedded workflows, customer relationships, and trust as remaining moats.

TechCrunch · AI · 2d agoAI industry1

Automating Attack Graph Construction for Agentic Pentesting. Towards Neuro-Symbolic Vulnerability Hunting

Semi-automated pipeline converts Trivy, Semgrep, Nmap output into MulVAL attack graphs for agentic pentesting, 53.7% mean vulnerability coverage in CyBench.

The paper presents a semi-automated pipeline that parses Trivy, Semgrep, and Nmap findings into MulVAL predicates and uses an LLM-assisted process to build domain-specific Datalog rules linking scanner evidence to attack techniques. MulVAL/XSB then performs symbolic inference to generate structured, auditable attack paths for agentic pentesting. Evaluated on 54 web CTF tasks from CyBench, every task produced at least one goal-reaching graph with 53.7% mean ground-truth vulnerability coverage, 51.9% full coverage, and an 83.9% noise-path rate. Median end-to-end runtime was 24.9 seconds, making the pipeline runtime-practical for agentic workflows.

arXiv cs.CR · 2d agoResearch

AWS’s new sign-up gives accounts spend caps, email invites, and agent-set permissionsnew

AWS's new sign-up flow gives fresh accounts agent-configured permissions, email-based team access, and per-project monthly spend caps starting at $20.

New AWS customers can sign up with Google, GitHub, or Apple identities, start with $100 in Free Tier credits, and build inside a project where AWS and coding agents automatically configure permissions and install tools like the AWS CLI and Agent Toolkit. In a demo, an agent deployed a Lambda function, DynamoDB table, and API Gateway endpoint without any manually written IAM policies. Paid projects get monthly spend limits starting at $20 that pause the project when reached, and the gradual rollout applies to new customers only.

Help Net Security · 46m agoAI tools & infra

Your startup’s next teammate might be an AI agent: Gusto, Insight Partners, and Leland explain what that changes at TechCrunch Disrupt 2026

TechCrunch Disrupt 2026 panel with Gusto, Insight Partners, and Leland will examine how startups integrate AI agents into early teams.

A Builders Stage session titled "Hiring When AI Is a Co-Founder" at TechCrunch Disrupt 2026 (October 13-15, Moscone West, San Francisco) features Gusto CEO Josh Reeves, Insight Partners SVP Michelle Johnson, and Leland CEO John Koelliker. The panel will discuss how early-stage startups decide which work to delegate to AI agents versus human hires, covering ownership, accountability, and culture. Gusto serves more than 500,000 companies, and Johnson previously helped scale Flock Safety from under $1 million to $90 million in ARR. The piece doubles as event promotion with discounted registration before September 25.

TechCrunch · AI · 2h agoAI industry

Snap tries to make the case again for its $2,200 smart glasses

Snap unveiled new features for its $2,200 Specs smart glasses, including an anticipatory AI system and enterprise partnerships with Amazon, Salesforce, and Nvidia.

At a Los Angeles event, Snap showcased updates for its Specs smart glasses, which launched earlier in 2026 at $2,200 to a mixed reception. The headline announcement was Specs Intelligence, an "anticipatory AI" system that builds an understanding of user goals and routines and works with iPhones and Macs independently of the glasses. Snap also launched Specs for Enterprise with partnerships including Amazon, Salesforce, and Nvidia, an NBA/WNBA AR training app, and a Verizon cellular connectivity package costing $10/month for Verizon customers and $20/month otherwise. The devices will ship later this fall after an October pop-up in Los Angeles.

Stanford Researchers Release Paper2Agent: Turning Research Papers Into AI Agents That Reproduce Results and Run on New Data

Stanford researchers released Paper2Agent, a Nature-published pipeline that turns research papers into MCP servers agents can execute.

A Stanford team led by Jiacheng Miao and James Zou published Paper2Agent in Nature on 16 September 2026. Built on Claude Code's agent SDK, it converts a paper and its codebase into a Model Context Protocol server with validated tools, resources, and prompts. In benchmarks, the AlphaGenome agent built 22 tools in about 45 minutes for US$14, scored 100% on 15 novel queries versus 78.7% for Claude Code with repository access, and cut median runtime 1.9x. In scale tests, 74 of 100 bioRxiv papers were converted and 593 of 599 proposed tools passed validation.

MarkTechPost · 7h agoAI research1

CaMeLoT: CaMeL orchestrated with Temporal logic for static verification and liveness

Researchers present CaMeLoT, extending CaMeL with CTL model checking that statically rejects unsafe LLM agent plans before any tool executes.

CaMeLoT adds a static verification layer to CaMeL, a runtime defense against prompt injection in tool-using LLM agents. It translates a generated plan into a finite-state transition system, labels it with tool calls, provenance, and taint information, and checks it against CTL temporal policies using the nuXmv model checker before any tool is invoked. Failed checks return counterexamples for plan repair, avoiding LLM calls, tool calls, and sandbox teardown. Evaluation covers policies derived from AgentDojo, SOC workflows, and prompt-extraction experiments.

arXiv cs.CR · 15h agoAI safety & security