Gacha Decoding: Eliciting Diverse Generations Through Instruction Following
Gacha Decoding uses instruction following and an external RNG to elicit diverse, high-quality language-model generations.
Researchers introduce Gacha Decoding, an inference-time method that treats diversity as an instruction-following problem rather than relying on token entropy. It pairs a language model's instruction following with randomness from an external RNG to realize distinct response modes. Across chat, creative writing, image-generation planning, and protein design, it reports up to 2.4x Vendi over the next-best prior approach and matches high-quality mode counts with 11.0x fewer samples. Diversity under the method improves as the underlying model becomes a better instruction follower.
- Treats generation diversity as an instruction-following problem.
- Reports up to 2.4x Vendi versus the next-best prior method.
- Reaches the same high-quality modes with 11.0x fewer samples.
- Diversity improves as instruction following improves, even as token entropy falls.
Full article178 words · extracted from arxiv.org · click to collapse
We introduce Gacha Decoding, an inference-time method for eliciting diverse language model generations that scales with model capability. Across open-ended domains (in-the-wild chat, creative writing, planning for image generation, and protein design), Gacha Decoding significantly outperforms existing generation diversity approaches at equal quality (up to 2.4x Vendi over the next-best prior approach), reaching the same number of high-quality modes with over an order of magnitude fewer samples (11.0x) and discovering novel modes that no other approach surfaces. Our key insight is to treat diversity as an instruction-following problem: rather than relying on the LM's token entropy, we combine its instruction-following capability with randomness from an external RNG tool to scalably identify and realize distinct modes of the response space. This approach of "planning with dice" enables Gacha to invert the long-observed tension between diversity and model capability. As the underlying LM becomes a better instruction follower, diversity under Gacha Decoding consistently improves--even as its token entropy and diversity under prior approaches decline. Together, our results highlight that instruction following, rather than token entropy alone, can drive generation diversity.
Text extracted automatically; images, tables and formatting may be missing. Original: https://arxiv.org/abs/2610.01382