ClaimMirage: When Self-Claims in Domain Names Change LLM Threat Judgments
Self-claims in domain names can alter LLM threat judgments without explicit prompt-injection, reducing or increasing domain safety alerts.
A study on ClaimMirage shows that self-claims in domain names can significantly alter how LLMs judge domain safety and approval, without explicit prompt-injection, leading to increased or reduced alerts.
- Self-claims in domain names alter LLM threat judgments without explicit prompt-injection
- Self-claims can reduce or increase domain name alerts
- Risk-denial terms reduce alerts by 45.3 percentage points
- Endorsement terms increase alerts by 65.6 percentage points
- Findings motivate testing resistance to self-claims
Full article162 words · extracted from arxiv.org · click to collapse
Short claims such as not-phishing or official can change how a large language model (LLM) judges a domain name, without explicit prompt-injection commands. We study this manipulation as ClaimMirage: a name under inspection claims its own safety or approval. We analyze 622,080 judgments across 64 brands and five LLMs, comparing ten claims with length- and hyphen-matched controls in constructed brand-like names. Self-claims can substantially reduce or increase alerts, depending on the LLM and input setting. In one setting, risk-denial terms inside the registrable name reduce alerts by 45.3 percentage points even with a basic safeguard: the prompt supplies the potentially impersonated brand and its official domain for comparison. Without these references, endorsement terms at that position increase alerts by 65.6 points in the same LLM. References and component annotation remove some alert reductions but leave others or make them larger. These findings motivate testing resistance to self-claims and seeking independent evidence before treating a domain name under inspection as safe or authorized.
Text extracted automatically; images, tables and formatting may be missing. Original: https://arxiv.org/abs/2609.29130