Bad Likert Judge: A Novel Multi-Turn Technique to Jailbreak LLMs by Misusing Their Evaluation Capability
Unit 42 details the Bad Likert Judge multi-turn jailbreak that abuses LLMs' evaluation capability, raising attack success rates over 60% across six frontier models.
Palo Alto Networks Unit 42 describes the Bad Likert Judge technique, a multi-turn jailbreak that asks a target LLM to act as a Likert-scale judge scoring the harmfulness of example responses. The highest-rated example in each scale can carry harmful content, bypassing the model's internal guardrails. Testing across six state-of-the-art text-generation LLMs showed an average attack success rate increase of more than 60% versus plain attack prompts, with tested models anonymized. The technique targets edge cases rather than typical use, and the article positions the work as guidance for defenders on potential jailbreak risks.
US Indicts 17 Iranians Over Years
US unsealed superseding indictment charging 17 Mabna Institute Iranians for IRGC-linked espionage stealing 31TB from universities, companies, and government agencies.
The Justice Department unsealed a superseding indictment charging 17 members of the Iran-based Mabna Institute, which conducted hacking campaigns since at least 2013 on behalf of the IRGC and other Iranian clients. The group compromised 144 US and 178 foreign universities, at least 42 US companies, and multiple government agencies, stealing over 31 terabytes of academic data and IP plus employee email inboxes. Hackers breached roughly 8,000 of 100,000 targeted professor accounts across 24 countries, selling stolen research through Megapaper.ir and Gigapaper.ir. Behzad Mesri, tied to the HBO breach and $6 million Bitcoin extortion, is among eight new defendants, and five defendants carry State Department Rewards for Justice bounties up to $10 million.