ZeroHour
Story · 1 source · 1 articlefirst updated ()

Building Multilingual Bridges: Data Mixing as the Pillar of Generalization for In-Language Reasoning

infoAI researchimportance 45
What's new: This is the first merged summary for this story (no prior dashboard entry). New developments reported: announcement of Tiny Aya L2-Thinker's >93% in-language reasoning rate across 60 languages on 6 benchmarks, the finding that held-out-language generalization works without per-language reasoning supervision, and the public release of model weights and multilingual reasoning data.
Merged summary · glm-5.3-flash · rewritten as coverage arrives

Researchers train Tiny Aya L2-Thinker, a 3.35B model achieving over 93% in-language reasoning across 60 languages via optimized multilingual data mixing; model weights and reasoning data are released.

The paper addresses L2 reasoning, where models reason consistently in the language of the user's prompt rather than defaulting to English. Through data-centric SFT optimization of data composition and scheduling, the 3.35B Tiny Aya L2-Thinker reaches an in-language reasoning rate above 93% across 60 languages on six benchmarks spanning math, commonsense, instruction following, open-ended generation, and cultural reasoning. The authors find that generalization to held-out languages relies on broad language coverage, multilingual non-reasoning data, and a strong English reasoning backbone, without needing reasoning supervision in every target language, indicating that reasoning is a transferable, language-agnostic behavior. Model weights and multilingual reasoning data are publicly released. The two source reports agree on all headline figures (model size, 60-language coverage, >93% rate, 6 benchmarks) and on the release of weights and data; no disagreements were found.

  • Model: Tiny Aya L2-Thinker, a 3.35B-parameter model.
  • Performance: exceeds 93% in-language (L2) reasoning rate across 60 languages.
  • Evaluation: 6 benchmarks covering math, commonsense, instruction following, open-ended generation, and cultural reasoning.
  • Method: data-centric SFT optimization via multilingual data mixing (data composition and scheduling).
  • Held-out language generalization attributed to broad language coverage, multilingual non-reasoning data, and a strong English reasoning backbone.
  • Achieved without per-language reasoning supervision; reasoning shown to be transferable and language-agnostic.
  • Release: model weights and multilingual reasoning data made publicly available.
OrganizationsAya

Coverage timeline

  1. · 7d ago
    Hugging Face daily papers· 45
    Building Multilingual Bridges: Data Mixing as the Pillar of Generalization for In-Language Reasoning

    Researchers train Tiny Aya L2-Thinker, a 3.35B model achieving over 93% in-language reasoning across 60 languages via multilingual data mixing.