ZeroHour

Search: “mscms”

30 stories

Smart search ranks by meaning as well as keywords (one row per story, last 45 days).

Top 10 Best Mobile Device Management (MDM) Solutions in 2026

A 2026 MDM buyer guide ranks ten solutions, recommending Microsoft Intune for Microsoft 365 estates and Jamf for Apple-only environments.

A 2026 buyer guide evaluates ten mobile device management solutions, leading with Microsoft Intune as the default for Microsoft 365 organizations and Jamf for Apple estates. It recommends choosing the enrolment model before selecting a vendor and clarifying BYOD visibility to prevent privacy disputes. Kandji, Mosyle, Omnissa Workspace ONE, ManageEngine, Scalefusion, and Hexnode are covered as alternatives. Guidance ties MDM to Zero Trust data access policies via Apple User Enrolment and Android work profiles.

Cyber Security News · 7d agoIndustry

Causal Foundation Models

A paper introduces causal foundation models (CFMs): pretrained networks that estimate treatment effects on new datasets via in-context learning without fine-tuning.

Causal foundation models (CFMs) apply the foundation-model paradigm to causal inference, replacing bespoke per-problem estimator pipelines with networks pretrained once at scale. CFMs estimate causal quantities such as the average treatment effect on entirely new datasets through in-context learning, without model updates. The work serves as a practical introduction to the emerging area, covering background in causal inference and machine learning and including example code and Jupyter notebooks.

Hugging Face daily papers · 15d agoAI research

Large Language Models for HVAC Operations in Building Energy Systems: A Critical Review of Methods, Applications, and Deployment Readiness

Systematic review of 66 studies finds LLMs for HVAC operations are mostly research-stage, with no ready-now deployment and only four pilot-level studies.

A critical review of 66 peer-reviewed studies from 2023 to March 2026 examines LLMs for HVAC operations in building energy systems. Only four studies reach pilot-level evidence, none reports sustained operational deployment, and 63 of 66 are research-only. Conventional ML, MPC, and RL remain dominant for high-frequency control and short-horizon forecasting, and the evidence supports LLMs primarily as semantic and workflow layers rather than autonomous controllers.

arXiv cs.AI / cs.LG / cs.CL · 12d agoAI research

mySCADA myPRO Manager

CISA advisory reveals unauthenticated privileged API access and arbitrary SMS sending in mySCADA myPRO Manager <=2.1, CVSS 9.8.

CISA advisory ICSA-26-258-03 discloses two vulnerabilities in mySCADA myPRO Manager <=2.1 with aggregate CVSS v3 of 9.8. CVE-2026-73807 (CVSS 9.8) lets unauthenticated network attackers access privileged management functions via the command API, while CVE-2026-82567 exposes an unauthenticated HTTP endpoint that sends arbitrary SMS messages through a connected GSM modem. Deployments span critical manufacturing, energy, food and agriculture, transportation, and water and wastewater sectors. CISA states no known public exploitation has been reported at this time.

The 12 Best Mobile Device Management (MDM) Solutions, Compared and Priced

A comparison of 12 MDM platforms ranks Microsoft Intune as best value for Microsoft 365 estates and Jamf, Kandji, and Mosyle for Apple fleets.

The buyer's guide compares 12 mobile device management (MDM) products, naming Microsoft Intune best value since it is included in Microsoft 365 E3/E5, and Jamf, Kandji, and Mosyle as Apple specialists with day-one OS support and automated compliance remediation. Eight of the twelve publish rates; per-device pricing punishes multi-device users, while Microsoft, Omnissa, and IBM offer per-user options. Free tiers from Mosyle, Miradore, and ManageEngine support genuine small deployments.

GBHackersupdated · 6d agofirst · 6d agoIndustry 4 sources

12 Best Endpoint Privilege Management (EPM) Tools Compared (2026): Features & Pricing

A 2026 buyer's guide compares 12 endpoint privilege management tools, ranking CyberArk and BeyondTrust as enterprise leaders.

An editorial comparison evaluates 12 endpoint privilege management (EPM) tools on elevation control, manageability, and pricing model. The guide argues that standing local-admin rights fuel ransomware and lateral movement, making their removal a high-impact control that cyber insurers increasingly mandate. CyberArk and BeyondTrust are positioned as enterprise-depth leaders, with Delinea, Heimdal, and ManageEngine for the mid-market, and Admin By Request and CyberFOX AutoElevate for SMBs and MSPs. Pricing is generally per endpoint or per user, and the article is explicitly an assessment rather than a product release or incident report.

GBHackersupdated · 13h agofirst · 5d agoIndustry 14 sources

MetroLLM-Bench: Evaluating Language Models as Transit Kiosk Runtimes

MetroLLM-Bench is a 955-case benchmark testing language models as transit kiosk tool-calling runtimes across six real metro systems.

The benchmark covers 37-414-station metro systems and eleven task categories including routing, fare calculation, disruptions, accessibility, and adversarial input, with 14 deterministic and 8 semantic scoring components. Of 26 models from six vendors, a PEFT-tuned 4B Qwen 3.5 student scored 91.3 on Tier 1, exceeding GPT-5.6 (90.6/90.0), while Muse Glimmer 30B led the composite ranking. A deterministic rule-based baseline reached 84.6, and PEFT gains over base models shrank from +7.03 points at 2B to -0.91 at 27B.

Hugging Face daily papers · 8d agoAI research

How to build an exposure management program the business trusts: Lessons from Tenable’s CSO

Tenable's CSO describes an AI-driven exposure management program that consolidates tool sprawl and translates cyber risk for boards.

A Tenable blog post shares lessons from CSO Robert Huber on moving to an AI-driven exposure management program. It argues tool sprawl and data silos hinder holistic risk assessment and that exposure management unifies attack-surface data into business-level metrics for the C-suite.

Tenable Blog · 20d agoIndustry

xDailyBench: Benchmarking LLMs on Professional Consultation for Real-Life Problems

xDailyBench tests 11 frontier LLMs on 248 real-life consultation tasks; the best models score 75.6% and lag on implicit requirements.

The benchmark spans 51 scenarios across personal life, white-collar work, learning and research, and cross-domain activities, grounded in requests users actually completed or intended to complete with AI. Tasks are scored with fine-grained binary rubrics covering explicit and implicit requirements under standardized agentic settings. Across 11 frontier models, the best achieved a 75.6% task-level score, with all models performing at least 9 percentage points worse on implicit than explicit requirements.

arXiv cs.AI / cs.LG / cs.CL · 9d agoAI research

Fortunate Recall: Ontology-Driven Memory Lifecycle Management for Persistent Coherence in LLMs

Fortunate Recall introduces ontology-based lifecycle policies for LLM memory, cutting confabulation roughly in half (e.g., 45.1% to 22.4%) versus Mem0.

Fortunate Recall (FR) is a composable policy layer that classifies personal facts into a 10+1 behavioral ontology and applies category-specific lifecycle rules including differential temporal decay, slot-key supersession, event-time validity, and retrieval routing. FR-Bank scores 76.9% on the new 516-question LifecycleBench, ahead of Mem0, A-MEM, Memory-R1, and MemoryOS (61%-70.5%), and 75.2% on LongMemEval-S. End-to-end, confabulation drops from Mem0's 45.1% to 22.4% over answered queries, with the ranking replicating on open-weight Kimi K2.5 and transferring to the independent BEAM benchmark (46.8% vs 32.9%).

arXiv cs.AI / cs.LG / cs.CL · 7d agoAI research

ZGCM-1: A Fully Open and Extremely Efficient Foundation Model for Math and Agentic Search

ZGCM-1 is a fully open 7B foundation model with 256K context that stays competitive with frontier models on math reasoning and agentic search.

ZGCM-1 is a fully open 7B dense foundation model trained from scratch using an efficiency-focused recipe: interleaved gated sliding-window and full attention, a stable FP8 Muon optimizer, and MDP-based mid-training with context scaling across 16K, 64K, and 256K. On mathematical reasoning and agentic search suites it remains competitive with much larger frontier models such as Qwen3-235B-A22B and GLM-5.1. The recipe yields a ~4.2x improvement in 16K pre-training time-to-loss, and all weights, checkpoints, training code, data recipes, and W&B logs are open-sourced.

Hugging Face daily papers · 6d agoModel release

RMM Tools for MSPs: Features, Risks & How to Stay Secure

Threat actors continue abusing MSP remote monitoring and management tools to reach downstream customers, four years after the Kaseya supply chain attack.

Huntress examines how RMM platforms remain a favored gateway for attackers targeting managed service providers and their clients. A recent incident demonstrates that adversaries still successfully pivot from MSP RMM tooling into downstream customer environments. The piece also covers RMM features, associated risks, and hardening guidance for providers.

Huntress · 15d agoThreat actor in the wild

Top 10 Best Endpoint Privilege Management (EPM) Tools in 2026

A 2026 scorecard ranks ten endpoint privilege management tools, led by BeyondTrust, ThreatLocker and Delinea for elevation, coverage and policy depth.

The article ranks ten endpoint privilege management (EPM) tools using weighted criteria covering elevation workflow, platform coverage, policy depth, time-to-value and value. BeyondTrust scored highest overall (8.4) for cross-platform breadth, with ThreatLocker (8.2), Delinea (8.1) and Admin By Request (8.0) highlighted for allowlisting integration, cloud administration and deployment speed respectively. It also notes that Netwrix acquired CoSoSys in 2024, which affects bundling when shortlisting both EPM and device control.

Cyber Security News · 6d agoIndustry

LLMs and Contextual Integrity

Bruce Schneier highlights two papers: the CIMemories benchmark shows frontier LLMs leak memory attributes up to 69%, and an RL method reduces inappropriate disclosures.

Bruce Schneier discusses contextual integrity in LLMs, referencing the CIMemories benchmark, which uses synthetic profiles with 100+ attributes per user to test whether models with persistent memory disclose sensitive information appropriately. Evaluation showed frontier models exhibit up to 69% attribute-level violations, with GPT-5's violation rate rising from 0.1% to 9.6% across 40 tasks and reaching 25.1% with repeated prompting, showing unstable leakage behavior. A second paper introduces a reinforcement learning framework trained on a synthetic 700-example dataset that substantially reduces inappropriate disclosure while maintaining task performance, with improvements transferring to the human-annotated PrivacyLens benchmark.

Schneier on Security · 29d agoAI safety & security1

Evaluating Time-Series Foundation Models and Multimodal Dietary Context for CGM Forecasting

Study finds zero-shot time-series foundation models underperform on CGM forecasting; fine-tuned Chronos-Bolt cuts RMSE up to 18.4% and dietary context adds signal.

The paper evaluates time-series foundation models for continuous glucose monitoring forecasting across eight public datasets covering Type 1 diabetes, Type 2 diabetes, and non-diabetes populations. Under a unified protocol, zero-shot foundation models did not consistently outperform baselines like Elastic Net and PatchTST, but lightweight fine-tuning did, with fine-tuned Chronos-Bolt reducing RMSE by 6.5%-18.4% in the T1D cohort and 8.6%-18.2% in the non-diabetes/T2D cohort. A residual-based fusion framework adding dietary context from CGMacros reduced overall RMSE by about 3% and postprandial RMSE by about 15% versus CGM-only baselines.

arXiv cs.AI / cs.LG / cs.CL · 6d agoAI research

Top 10 Best Cloud Security Posture Management (CSPM) Tools in 2026

2026 CSPM comparison ranks Wiz atop cloud posture tools and recaps Google's pending roughly $32 billion acquisition of Wiz.

An editorial guide rates ten cloud security posture management (CSPM) tools, with Wiz ranked first for agentless visibility and attack-path context, Microsoft Defender for Cloud highlighted for Azure-centric economics, and Palo Alto Prisma Cloud noted for the broadest code-to-cloud module set. The article's biggest market note is Google's agreement to acquire Wiz for approximately $32 billion, described as the largest deal in security history, still progressing through regulatory review. It advises buyers to include roadmap-protection language in multi-year commitments and to press on multicloud neutrality post-close.

Cyber Security News · 5d agoIndustry

Molecular Déjà Vu: Digit-Level Retrieval of Published Values in Frontier Language Models

Audit of 22 frontier models finds widespread verbatim retrieval of published molecular property values, with higher reasoning increasing recall of memorized numbers.

An arXiv audit tests 22 frontier LLMs across 12 molecular regression benchmarks for verbatim retrieval of published values. More than 50% of the LLMs show verbatim retrieval on five datasets, and identical experiments are flagged 89% more often at a high reasoning level than at the lowest one. Suppressing retrieval moves model prediction errors closer together in relative terms, suggesting predictive capability is not determined solely by memorized values.

arXiv cs.AI / cs.LG / cs.CL · 12d agoAI research1

Ivanti Patches 10 EPMM, Neurons for ITSM and Sentry Flaws Enabling RCE and Admin Access

Ivanti patches 10 flaws in EPMM, Neurons for ITSM, and Sentry, including two 9.8-rated unauthenticated RCEs in ITSM.

Ivanti released fixes for 10 vulnerabilities across Endpoint Manager Mobile, Neurons for ITSM, and Sentry, and said it was not aware of active exploitation at disclosure. The most severe are CVE-2026-12744 and CVE-2026-12745, unauthenticated deserialization RCEs rated 9.8 in Neurons for ITSM, alongside authenticated deserialization and missing-authorization RCEs rated up to 9.9. CVE-2026-18851 is an 8.8-rated EPMM privilege escalation to administrator, and CVE-2026-83527 is an 8.1-rated unauthenticated authentication bypass in Sentry granting administrative access. Ivanti said the ITSM weaknesses were found using large language models; Cloud/SaaS fixes shipped August 9, 2026, and on-premises patches are available from September 2026.

IBM security advisory (AV26-922)

Canadian Cyber Centre relays IBM advisory for Langflow, MQ, and Sterling File Gateway flaws including MQ remote code execution (CVE-2026-13293).

Canadian Cyber Centre advisory AV26-922 relays IBM fixes for Langflow OSS (versions through 1.11.5 across release lines), IBM MQ (10.0.0.0 and 9.x LTS/CD through 9.4.5.1), and Sterling File Gateway (through 6.2.2.1). CVE-2026-13293 is a remote code execution flaw in IBM MQ Java messaging caused by an incomplete security scanner blocklist enabling network-based code execution. CVE-2026-19290 is an improper access control vulnerability in IBM Sterling File Gateway. Administrators are urged to review and apply the necessary updates.

llm 0.34

Version 0.34 of Simon Willison's llm CLI adds response-duration metrics to log output, plus bug fixes and faster log querying.

The open-source llm command-line tool for interacting with large language models released version 0.34. The headline change adds response duration in milliseconds and human-readable form to llm logs --usage Markdown output, plus a new duration_ms field in llm logs --short. The release includes several contributed bug fixes and a significant performance improvement to llm logs, alongside the related llm-openrouter 0.7.1 release.

Simon Willison · 14d agoAI tools & infra1

CVE-2026-52307: Stored XSS in 1CMS v5.6

CVE-2026-52307: authenticated stored XSS in 1CMS (ClassCMS) v5.6 Column Management lets attackers inject scripts via the title field.

ClassCMS 1CMS v5.6 contains an authenticated stored cross-site scripting vulnerability, CVE-2026-52307, in the Column Management component. Attackers can execute arbitrary web scripts or HTML by injecting a crafted payload into the title field. No CVSS score, patch information, or exploitation evidence was provided in the disclosure.

Full Disclosure · 8d agoVulnerabilityCVE-2026-52307

Hackers Use LLMs to Generate Exploit Scripts and Automate Post-Exploitation Across Latin America

Unit 42 says Latin American attackers used LLMs to automate post-exploitation in campaigns hitting Mexican government, water utilities, and Brazilian financial firms.

Unit 42 identified two campaigns in Latin America whose operators used commercial LLMs (Claude, GPT-4.1) behind a self-hosted NextChat interface to generate and debug post-exploitation scripts. Cluster CL-CRI-1131 compromised a transportation organization, Mexican federal ministries, and water utilities in Mexico and Ecuador, using native Windows tools and Volume Shadow Copies to dump the SAM registry hive and NTDS.dit. Cluster CL-CRI-1163 targeted Brazilian financial organizations with job-themed phishing, custom RATs, and a Go-based reverse SOCKS5 tunneling utility called SockTz, with nine versions deployed within roughly two hours. Trend Micro tracks related AI-augmented activity as SHADOW-AETHER-040 and SHADOW-AETHER-064.

GBHackersupdated · 6d agofirst · 6d agoThreat actor in the wild 2 sources1

Verifiable by Construction: Claim-Level Evaluation of Verbatim Citation in Clinical Question Answering

Evaluation of twelve LLMs on 222 clinical questions shows verbatim quotes rarely substantiate claims; claude-opus-5 fully substantiates only 37.1%.

The authors build a standardized harness over four clinical practice guidelines and evaluate twelve LLMs on 222 synthetic clinical questions, measuring citation attachment, verbatim quote production, and claim substantiation. Most models attach verbatim quotes to over 90% of claims from prompting alone, though lightweight models like claude-haiku-4.5 struggle. Quotes frequently fail to substantiate claims: claude-opus-5 quotes 98.0% of claims but fully substantiates only 37.1%, exposing a capability gap for verifiable clinical QA.

arXiv cs.AI / cs.LG / cs.CL · 2d agoAI research

US Defense Contractors Admit Their Rising CMMC Scores May Not Be Accurate

US defense contractors report rising self-assessed CMMC Phase I scores while doubting the accuracy of those self-evaluations.

Self-assessment scores under the US Department of Defense's Cybersecurity Maturity Model Certification (CMMC) Phase I have reportedly reached an all-time high. Contractors themselves doubt whether these self-reported scores accurately reflect their real security maturity. The uncertainty highlights concerns about the reliability of self-assessments in the program's initial phase.

Infosecurity Magazine · 27d agoPolicy & legal

Necessary or Sufficient? Evaluating LLM Explanations With Behavioural Evidence

New framework tests whether LLM-cited explanation factors are necessary or sufficient, finding weak correlation across Claude, GPT, and Gemini models.

An arXiv paper introduces black-box intervention tests measuring whether factors LLMs cite in their explanations are necessary or sufficient for their outputs in agent oversight workflows. Across eight models from the Claude, GPT, and Gemini families, Spearman correlations between cited rankings and measured influence ranged from 0.349-0.354 (advisor recommendation) to 0.431-0.580 (prompt monitoring). Uncited factors scored above the lowest cited factor in up to 57.6% of advisor responses, showing cited top-three factors do not reliably identify the most influential inputs.

Ivanti EPMM, Neurons and Sentry Vulnerabilities Enable Privilege Escalation and RCE Attacks

Ivanti patched ten CVEs across EPMM, Neurons for ITSM and Sentry, including critical unauthenticated deserialization RCE; no active exploitation reported.

On September 8, 2026, Ivanti disclosed advisories covering ten CVEs in Endpoint Manager Mobile (EPMM), Neurons for ITSM, and Sentry. The most severe are two unauthenticated deserialization RCE flaws in Neurons for ITSM, CVE-2026-12744 and CVE-2026-12745 (CVSS 9.8), plus three missing-authorization RCE bugs rated 9.9 and three authenticated deserialization RCE flaws. EPMM has CVE-2026-18851 (CVSS 8.8), an authenticated privilege escalation flaw, and Sentry has CVE-2026-83527 (CVSS 8.1), an authentication bypass. Ivanti reports no evidence of active exploitation; cloud/SaaS ITSM was patched on August 9, 2026, while on-premises 2025.2 through 2026.1 require September 2026 patches.

IBM security advisory (AV26-862)

Canada's Cyber Centre relayed an IBM advisory disclosing vulnerabilities across SPSS, MQ, Maximo, Instana, Concert and other IBM products.

The Canadian Centre for Cyber Security (AV26-862) published an IBM security advisory dated August 31, 2026, noting vulnerabilities affecting multiple IBM products as of August 28, 2026. Affected products include SPSS Collaboration and Deployment Services, IBM MQ Agent, Maximo Application Suite Monitor Component, Observability with Instana agent, Financial Transaction Manager for Red Hat OpenShift, Engineering Test Management, Control Center, Concert Software and Tivoli System Automation. No CVE identifiers, CVSS scores or exploitation details are provided in the relayed text.

Canadian Centre for Cyber Security · 16d agoAdvisory

Towards Scalable and Cost-Efficient Vulnerability Detection: A Study on Automatic Query Generation

A study finds LLM-synthesized CodeQL queries improve average F1-score by 82% over baseline queries, offering scalable vulnerability detection versus direct LLM scanning.

Researchers conducted an empirical study evaluating whether LLMs can synthesize executable CodeQL queries from National Vulnerability Database vulnerability data. LLM-generated queries significantly enhanced baseline CodeQL suites, yielding an 82% improvement in average F1-score across a diverse set of real-world vulnerabilities. A cost-benefit analysis shows direct LLM-based scanning of entire repositories is often computationally and financially prohibitive, while LLM query synthesis offers a scalable and cost-effective alternative for large-scale vulnerability detection.

arXiv cs.CR · 7d agoResearch1

Graph Machine: Towards Better Pretraining via Edges

Researchers propose Graph Machine, an O(n)-state sparse architecture that replaces 75% of Qwen3-0.6B dense layers with only slight loss change.

The paper introduces the Graph Machine (GM), an architecture that maintains an O(n)-sized state accessed through sparse, dynamic routing via pointer-like edges updated differentiably by a referral mechanism resembling pointer chasing. The authors replaced 75% of dense Transformer layers in Qwen3-0.6B with GM sparse layers and pretrained from scratch on 15.7B tokens. Retrieving 2 of 4,096 tokens per KV head in each sparse layer degrades loss only slightly, while retrieving 4 marginally improves loss over the dense baseline.

Hugging Face daily papers · 15d agoAI research

Towards a Deterministic Math Solver for Clinical Language Models

Paper shows handing arithmetic to a deterministic Python solver beats direct model calculation at 32B but not reliably at 7B on MedCalc-Bench.

Researchers test a Program-Solve interface where clinical LLMs write case-specific Python executed by a restricted local solver instead of doing arithmetic directly. On MedCalc-Bench Verified (1,100 cases, 55 calculators), Qwen2.5-32B-AWQ scored 90.53% with solver handoff versus 83.47% with direct arithmetic (+7.05 points), while Qwen2.5-7B gained an unreliable +3.29 points with a confidence interval spanning zero. The authors audited the benchmark against clinical guidelines and flagged 16 of 55 calculators for version, use, or coefficient concerns.

Hugging Face daily papers · 8d agoAI research