ZeroHour
arXiv cs.AI / cs.LG / cs.CLpublished ()ingested Yingzhi Wang1

Nuha-Speech: Building General-Purpose Arabic Speech-LLMs

infoAI researchimportance 30
AI summary · glm-5.3-flash

Nuha-Speech initiative builds general-purpose Arabic speech-LLMs using a 1.5M-sample speech QA corpus and fine-tuned Qwen-Omni variants.

The paper introduces Nuha-Speech, an initiative covering dataset construction, model training, and evaluation for Arabic speech large language models. The authors built an Arabic Speech Question-Answering corpus of over 1.5 million training samples and used it for supervised fine-tuning of Qwen-Omni model variants at multiple scales. A tailored evaluation framework with diverse tasks and metrics is designed to assess Arabic speech capabilities under limited resource constraints.

  • 1.5M-sample Arabic speech question-answering corpus for instruction tuning
  • Supervised fine-tuning of Qwen-Omni variants at different scales
  • Evaluation framework with diverse tasks and tailored metrics
  • Targets underrepresented Arabic language in multilingual speech-LLMs
VendorsAlibaba
ProductsNuha-Speech
OrganizationsarXiv
AI modelsQwen-Omni
Full article124 words · extracted from arxiv.org · click to collapse

As Speech Large Language Models (speech-LLMs) become increasingly multilingual, Arabic remains significantly underrepresented, highlighting the need for dedicated infrastructure to train and evaluate Arabic speech-LLMs. To address this gap, we introduce Nuha-Speech, a comprehensive initiative to develop general-purpose Arabic speech-LLMs spanning dataset construction, model training, and systematic evaluation. Specifically, we constructed a large-scale Arabic Speech Question-Answering (SQA) corpus comprising over 1.5 million training samples to allow instruction tuning over a broad range of core speech tasks. Then, the corpus was used for supervised fine-tuning based on Qwen-Omni model variants at different scales. Finally, we designed an evaluation framework featuring diverse tasks and tailored metrics. Through this work, we aim to establish foundational infrastructures for Arabic Speech-LLMs under constraints imposed by limited Arabic speech resources.

Text extracted automatically; images, tables and formatting may be missing. Original: https://arxiv.org/abs/2609.11892