LLM-as-Jev: LLMs Are Already Jev-Style Decision Models -- When and How to Fine-Tune Them
LLM-as-Jev reads calibrated categorical decisions from token probabilities, matching specialized models on untuned Qwen3.5-4B.
LLM-as-Jev extracts categorical decisions from next-token probabilities over bracketed numeric option identifiers, without free-form text. It provides a training-free readout and a fine-tuning objective that uses a tree-factorized listwise loss plus KL divergence anchors to the base model. On Qwen3.5-4B and Qwen3-0.6B, the untuned 4B model matches community Jev-style models on the same backbone, beats letter-logit readouts, and handles variable option counts and image decisions. Fine-tuning mainly helps weaker models and many-option intent routing, while KL anchors and LoRA preserve conversational generation.
- Reads decisions from probabilities of bracketed numeric option tokens.
- Untuned Qwen3.5-4B matches community Jev models on the same backbone.
- Fine-tuning helps weaker models and many-option intent routing most.
- KL anchors and LoRA limit degradation of conversational generation.
Full article188 words · extracted from huggingface.co · click to collapse
Jev-style decision models return categorical probability distributions over predefined options without generating free-form text, enabling software systems to act on their outputs directly. In this work, we investigate the extent to which general-purpose LLMs already possess this capability out of the box, and when fine-tuning is actually necessary. We present LLM-as-Jev, an architecture-preserving framework that extracts calibrated decisions directly from next-token probabilities over bracketed numeric identifiers. LLM-as-Jev provides both a training-free inference recipe and a fine-tuning objective that optimizes candidate selection via a tree-factorized listwise loss while anchoring auxiliary predictions to the base model using KL divergence penalties. Evaluating on Qwen3.5-4B and Qwen3-0.6B, we find that modern LLMs are inherently effective decision models: without training, the 4B model matches community Jev-style models built on the same backbone, outperforms letter-logit readouts, supports arbitrary option counts, and natively handles multimodal decisions over images. Fine-tuning provides targeted rather than universal benefits -- substantially improving weaker models and specific tasks (such as many-option intent routing), but offering diminishing returns for strong backbones. Crucially, our KL anchors prevent behavioral degradation in conversational text generation, with LoRA delivering the strongest performance on capable models.
Text extracted automatically; images, tables and formatting may be missing. Original: https://huggingface.co/papers/2610.02076