JevOut: Natural Context Can Flip Decision Models
Natural context additions flip Jev and peer decision models onto high-confidence wrong choices.
JevOut studies decision models that map language to probabilities over finite choices used to route requests, select tools, and trigger actions. Short, natural-looking context additions, optimized while keeping the source, question, choices, and gold answer fixed, redirect Jev on 312 of 508 initially correct decisions (61.4%); in 229 cases the wrong option receives probability at least 0.7. Across seven datasets, three other decision systems show targeted flip rates of 64.9% to 73.2% on items they initially answer correctly. The authors argue this fragility makes raw probability outputs unreliable decision interfaces.
- Contexts redirect Jev on 312 of 508 initially correct decisions, or 61.4%.
- In 229 cases Jev assigns at least 0.7 probability to the fixed wrong option.
- Three other systems flip 64.9% to 73.2% of initially correct decisions.
- Authors warn these probability outputs are unreliable interfaces for routing and actions.
Full article202 words · extracted from arxiv.org · click to collapse
Dedicated decision models such as Jev map unstructured language to probability distributions over finite choices, allowing their outputs to directly route requests, select tools, and trigger actions. Yet real-world inputs rarely arrive in isolation: they come with background details and surrounding context. We find that short additions that fit naturally into this context can nevertheless redirect an otherwise correct decision, even when the correct answer remains unchanged. To study this behavior, we fix a wrong target option for each initially correct item and use the model's option probabilities to refine fluent context additions while preserving the source, question, choices, and gold answer. Within 64 accepted target evaluations, the optimizer identifies contexts that redirect Jev on 312 of 508 initially correct decisions (61.4%); in 229 cases, Jev assigns at least 0.7 probability to the fixed wrong option. Across seven datasets, three additional decision systems show targeted flip rates of 64.9%-73.2% on decisions they initially answer correctly. Taken together, these results expose a pronounced fragility in current decision models: short, ordinary-looking context can shift a correct choice to a high-confidence wrong one. Because these models turn language directly into downstream choices, this sensitivity raises concerns about treating their probability outputs as reliable decision interfaces.
Text extracted automatically; images, tables and formatting may be missing. Original: https://arxiv.org/abs/2609.30243