ZeroHour

Search: “GPT-4o-mini”

2 stories in the last 7d

PuzzleMask: Abusing Plain Prose as a Covert AI Attack Vector

Check Point details PuzzleMask, a plain-prose technique that bypasses LLM gatekeeper policy checks, letting hidden payloads reach target models unreviewed.

Check Point Research describes PuzzleMask, a prompt-crafting technique that hides policy-violating payloads inside plain-English prose wrappers, bypassing quick LLM-based policy checks without emojis, Base64, or invisible formatting. The researchers tested 23 automated prompts against gatekeepers including GPT-4o-mini, GPT-OSS-Safeguard 20b, Claude 3 Haiku, and Llama Guard 3, and all were classified as safe despite policies that flagged the plain versions. When submitted to GPT-5 in thinking-high mode with a Python interpreter, the target model extracted and acted on the payload in over 90% of trials. The technique is not itself a jailbreak but can carry a jailbreak prompt as payload; mitigations include input paraphrasing, hardened gatekeeper policies, and output monitoring.

Check Point Researchupdated · 6d agofirst · 6d agoAI safety & security 2 sources

Enhancing Accessibility of Medical Texts through Large Language Model-Driven Plain Language Adaptation

Study shows LLMs with Mixture-of-Agents and QLoRA finetuning effectively simplify medical texts into plain language while preserving content.

The paper evaluates Plain Language Adaptation (PLA) using GPT-4o-mini, Gemini-1.5-pro, and LLaMA in zero-shot and few-shot settings. It compares prompting strategies, QLoRA finetuning across models, and integrates Mixture-of-Agents (MoA) techniques for robustness. Results demonstrate LLM-driven PLA makes healthcare texts more comprehensible while retaining essential content.

arXiv cs.AI / cs.LG / cs.CL · 1d agoAI research