HINTT Submission to the 2nd MLC-SLM Challenge: Comparing Cascaded and Unified Approaches to Diarization and ASR
HINTT's cascaded diarization and Qwen3-ASR pipeline beat a unified speech LLM on MLC-SLM Task 1.
HINTT submitted a multilingual speaker-attributed ASR system to the 2nd MLC-SLM challenge, which requires determining who spoke when and what was said. The submitted cascaded pipeline combines a fine-tuned DiariZen diarization model, a fine-tuned Qwen3-ASR model, and LLM-based generative error correction. A unified VibeVoice-ASR model was fine-tuned on the same official training data for comparison, without external data or pseudo-labels. Under Task 1 conditions, the cascaded system was more reliable, while unified speech LLMs remain a future direction.
- Cascaded system uses fine-tuned DiariZen, Qwen3-ASR, and LLM error correction.
- Unified comparison model is fine-tuned VibeVoice-ASR on the same official data.
- All fine-tuning used only official MLC-SLM data, with no external labels.
- Cascaded pipeline was more reliable under MLC-SLM Task 1.
Full article155 words · extracted from arxiv.org · click to collapse
This paper presents the HINTT system submitted to the 2nd Challenge and Workshop on Multilingual Conversational Speech Language Model (MLC-SLM). We address multilingual speaker-attributed ASR, where systems must determine who spoke when and what was spoken. We investigate two modeling strategies for this problem: a cascaded pipeline that combines speaker diarization with speech-LLM-based ASR, and a unified speech LLM that directly generates speaker labels, timestamps, and transcriptions. Our final submission is based on the cascaded pipeline, consisting of a fine-tuned DiariZen diarization model, a fine-tuned Qwen3-ASR model, and LLM-based generative error correction. For comparison, we also fine-tune VibeVoice-ASR as a unified model using the same official training data. All task-specific fine-tuning and model selection are performed using only the official MLC-SLM data, without external data or pseudo-labels. Experimental results demonstrate that the cascaded system remains more reliable under the MLC-SLM Task 1 conditions, while unified speech LLMs offer a promising direction for future speaker-attributed ASR.
Text extracted automatically; images, tables and formatting may be missing. Original: https://arxiv.org/abs/2610.08063