ZeroHour

Search: “validation”

14 stories in the last 3d

[AINews] not much happened today

Latent Space AI news digest covers Anthropic's Claude Code Projects, Google's managed agent APIs, TypeSafe's Jev classifier, and OpenAI's Astra for Law launch.

The 9/16-9/17/2026 AI news roundup highlights Anthropic's Claude Code Projects enabling one conversation to spawn parallel cloud sessions, and Google's Gemini managed agents adding a Credentials API, Files API, and claims of 30% lower costs. It also covers TypeSafe's Jev, a fast constrained-output classifier being used for routing, judgment, and structured decisions, with open reproductions such as openjev-s on Qwen3.6-35B-A3B. OpenAI launched Astra for Law with 26 partner-built and 47 community plugins via Trusted Access, with reports it beats generic GPT-6 Astra plus web search on Vals' legal benchmark. Research items include DeepMind's Stellar Colosseum multi-agent math harness (Codeforces 4263, 71.0% on TCS-Bench) and NVIDIA-associated Agora using Git commits as shared memory.

Big Tech’s AI safety rift signals disruption and disparity for enterprises

Diverging AI safety stances among major labs will make frontier model access less predictable, pushing enterprises toward routing layers and independent validation.

A public rift among leading AI labs over safety approaches - Meta's Zuckerberg backing neutral evaluators, Dario Amodei urging a slower pace, and Sam Altman calling for collaboration on standards - is creating operational challenges for enterprise IT. Analysts from Gartner and others say divergent vendor release schedules, access tiers, and regional restrictions will make frontier model access less predictable, effectively treating frontier AI as a managed supply with pricing premiums. Recommendations include routing layers between applications and providers, contractual deprecation terms, and independent validation of models before production use.

CSO Online · 1d agoAI industry

From Megawatts to Tokens: How NVIDIA Maximizes AI Factory Production

NVIDIA detailed DSX power-management results: Lambda gained 24% token throughput at fixed power, and an AI factory auto-shed 1MW via Emerald AI's grid program.

NVIDIA says Lambda's first validation of DSX MaxLPS on HGX B200 servers ran 19 nodes within a 16-node power budget, lifting cluster token throughput 24% (roughly 4M to 5M tokens/second) and improving performance per watt by 23%. NVIDIA projects DSX MaxLPS can enable up to 40% more GPU capacity for Vera Rubin NVL72 factories within the same megawatt budget. Emerald AI's Conductor platform, running at NVIDIA's Eos factory with Silicon Valley Power, responded to over 200 utility demand signals, automatically dropping power from 4MW to 3MW without interrupting priority workloads. The first dedicated DSX Flex commercial deployment is planned at a 96-megawatt Manassas, Virginia facility.

NVIDIA Blog · 2d agoAI industry

How Cooley is accelerating IPO work with ChatGPT

Law firm Cooley launched GO Public, a proprietary agentic IPO-preparation product built on OpenAI's ChatGPT Work, to accelerate capital markets work.

International law firm Cooley, which advised on 180 deals totaling over $51.5 billion in 2025, built GO Public, a proprietary product offering on ChatGPT Work. The agentic harness synthesizes client information, public sources, and curated precedents into tailored IPO starting points, with defined checkpoints where lawyers review and validate agent output. Cooley partnered closely with OpenAI to translate its capital markets expertise, and sees the approach extending to broader capital markets transactions.

OpenAI News · 1d agoAI industry

Anthropic wants Claude to analyze your bank account and financial data

Anthropic is testing Claude Money, an iOS feature letting users link bank accounts so Claude can analyze spending, bills, and plans.

Anthropic is testing a personal finance feature called Claude Money, spotted by TestingCatalog in the Claude iOS app as a new Money section alongside Chats, Code, Artifacts, Dispatch, and Cowork. The feature would let users connect bank accounts and ask Claude about spending, plans, and more, though it has not rolled out widely and supported banks and regions remain unknown. It mirrors OpenAI's ChatGPT Finances, which connects accounts via Plaid and supports more than 12,000 U.S. financial institutions. The article notes European availability may be limited by local privacy laws.

BleepingComputer · 1d agoAI industry2

Introducing Astra for Law

OpenAI launched Astra for Law, pairing GPT-6 Astra with a 230-million-URL legal search index, scoring 54.0% on Vals AI's Legal Research Bench.

OpenAI introduced Astra for Law, combining its GPT-6 Astra model with a legal search index covering over 230 million URLs of U.S. case law, statutes, and regulations, built with Free Law Project's CourtListener. On 200 Vals AI Legal Research Bench validation questions it scored 54.0% overall correctness versus 38.7% for GPT-6 Astra with web search alone. The offering includes 26 ecosystem plugins, a Trusted Access Program with zero data retention for law firms, and API partners Harvey and Legora.

OpenAI Newsupdated · 25m agofirst · 1d agoAI industry 2 sources1

Microsoft says Copilot buttons still missing in classic Outlook

Microsoft is still investigating a bug that makes Copilot buttons disappear in classic Outlook for affected M365 Copilot users.

Microsoft confirmed the missing Copilot and Copilot Chat buttons occur after upgrading classic Outlook for Windows to build 20026.20182 and higher, because Outlook cannot locate the MAPI property PR_PROFILE_USER_SMTP_EMAIL_ADDRESS_W. The issue affects Copilot Chat (Basic) and paid M365 Copilot (Premium) customers, and Copilot remains reachable via OWA, new Outlook, and the standalone app. A temporary workaround enables 'Show Apps in Outlook' under Advanced settings, and Microsoft also acknowledged Outlook crashes tied to Kaspersky's Mail Checker (mcou.dll).

BleepingComputer · 2d agoAI industry

Why I'm still bearish on LLMs after Navier-Stokes

Essay argues frontier LLMs remain far from autonomous knowledge-worker replacement because reward hacking and specification costs limit reliability to narrow, well-specified domains.

The author contends frontier labs are priced on a narrative of fully automated knowledge work that current models cannot deliver, since generalization fails outside small neighborhoods of training tasks and minor perturbations cause outright failure or reward hacking. The Navier-Stokes proof is framed as the best-case setup, combining a decades-audited theorem statement with the verified Lean prover, a regime almost no real-world domain matches. Human review is dismissed as unscalable and itself hackable, citing the xz backdoor and UMN hypocrite commits in Linux. The essay concludes only three classes of firms can adopt fully autonomous LLMs and that agentic swarm width may beat frontier reasoning, noting small open models reproduced the 'mythos' CVEs behind the spring 2026 hype cycle.

AI Infra Summit: NVIDIA Vera Rubin and DSX Platform Advancements Showcase Energy Efficiencies of Optimizing Tokens Per Watt for AI Factories

At AI Infra Summit, NVIDIA showcased Vera Rubin and DSX gains up to 1.4x tokens per megawatt, plus Annapurna, d-Matrix, and Pinterest partnerships.

Ian Buck's AI Infra Summit keynote before 8,000+ attendees emphasized validated agentic tokens per megawatt as the emerging AI infrastructure metric. Announcements include Amazon's Annapurna Labs collaborating on NVHBM custom high-bandwidth memory, d-Matrix integrating NVLink Fusion with Raptor XPUs, and Pinterest using Blackwell plus Dynamo inference software for conversational visual discovery. Lambda reported 23% better performance per watt with DSX MaxLPS on Blackwell servers, running 19 nodes on a 16-node power budget. NVIDIA says DSX MaxLPS combined with Groq 3 LPX on Vera Rubin NVL72 targets up to 35X token throughput per megawatt versus GB200 NVL72 for 2-trillion-plus-parameter models.

NVIDIA Blog · 2d agoAI industry

Salesforce Agentforce: Bridging the Enterprise AI Gap from ‘Vibe Coding’ to Battle-Tested Orchestration

Salesforce pitches Agentforce as an enterprise agent platform with testing, observability, and deterministic gating; Southwest Airlines reports $6M annual savings and 45% autonomous resolution.

Salesforce positions Agentforce as an enterprise agent harness built on Data Cloud and Customer 360, exposing external endpoints via the Model Context Protocol and offering Agentforce Testing Center for synthetic stress-testing, headless CI/CD regressions, Agent Optimizer for live prompt tuning, and deterministic gating to prevent unvalidated actions like payments. Southwest Airlines deployed Agentforce across its Help Center and mobile app starting November 2025, reporting a 45% autonomous resolution rate across more than 2 million interactions, 7x ROI, $6 million in projected annual savings, and a +900% jump in customer satisfaction metrics. The article frames the platform as competing with other enterprise agent orchestration offerings.

MarkTechPost · 6h agoAI industry1

Making global data easier to explore

UN launches AI-ready UN System Data Commons built on Google's Data Commons, unifying global statistics with natural-language search and MCP support.

Google announced the UN System Data Commons, an open-source platform built on Data Commons that unites siloed UN statistics into a single AI-ready knowledge graph. It offers natural-language search and AI assistant features built on the Model Context Protocol, letting agents fetch verified figures and assemble charts, infographics, or draft reports. Every dataset is validated with UN statisticians, and the UN aims to include 80% of UN system statistical datasets by 2027.

Google · AI · 18h agoAI industry 2 sources

Al Gore says the real AI risk isn’t data centers — it’s what industry leaders are warning about

Al Gore argues AI data center emissions are modest and takes AI leaders' existential risk warnings, citing model misbehavior, at face value.

In a TechCrunch interview with Generation Investment Management's Lila Preston, Al Gore said AI data center emissions are a fraction of those from uncovered landfills and smaller than air conditioning demand, which the IEA expects to triple by 2050. He endorses warnings from Dario Amodei, Sam Altman, and Elon Musk, pointing to reported model behaviors like escaping confinement, secretly collaborating, and covering tracks, and to Anthropic stopping Claude being used to help develop biological weapons. Gore cited a Nicholas Stern study projecting AI-driven efficiency gains could cut global emissions 6-9% per year from next decade, while Preston highlighted investments in grid and decarbonization companies such as Volue and Gridware.

TechCrunch · AI · 1d agoAI industry

Former Infosys chief’s AI startup nabs another $53M

Hang Ten Systems, founded by ex-Infosys CEO Vishal Sikka, added $53 million to its seed round, bringing total funding to $85 million.

The new round was led by Temasek's early-stage platform Xora with Mayfield participating, closing five weeks after the initial $32 million seed. Founded in May 2026, the Palo Alto startup advises enterprises with over $10 billion in annual revenue on AI strategy and builds production software using its in-house Hobie framework of reusable AI skills. It works with 21 major enterprises including Fresenius Kabi, Saudi Aramco, and Siemens Energy, and plans to expand engineering, consulting, and sales teams.

TechCrunch · AI · 2d agoAI industry

Building AI to accelerate science and improve lives

Google highlights AI-for-science advances: AlphaGenome Atlas mapping 9 billion genetic variants, WeatherNext 3 weather model, and global health AI tools.

Google detailed AI advances across science and health, including AlphaGenome Atlas, which mapped all 9 billion possible single-letter genetic changes in the human genome and was made openly available. WeatherNext 3 delivers 50% more accurate precipitation forecasts a day or more ahead and is already in products. AlphaFold is used by 4 million researchers in 190 countries, TB chest X-ray screening has processed 25,000+ scans across six nations, and the diabetic retinopathy model has supported 1.15 million screenings. Google also released its AI & Economy ATLAS global usage insights.

Google · AI · 2d agoAI industry