HINTT Submission to the 2nd MLC-SLM Challenge: Comparing Cascaded and Unified Approaches to Diarization and ASR
HINTT's cascaded diarization and Qwen3-ASR pipeline beat a unified speech LLM on MLC-SLM Task 1.
HINTT submitted a multilingual speaker-attributed ASR system to the 2nd MLC-SLM challenge, which requires determining who spoke when and what was said. The submitted cascaded pipeline combines a fine-tuned DiariZen diarization model, a fine-tuned Qwen3-ASR model, and LLM-based generative error correction. A unified VibeVoice-ASR model was fine-tuned on the same official training data for comparison, without external data or pseudo-labels. Under Task 1 conditions, the cascaded system was more reliable, while unified speech LLMs remain a future direction.