Align Then Reason: A Multimodal Lip-Sync Judge for Dubbing
Align Then Reason judges multilingual dubbing lip-sync by aligning lip frames to phonetics before LLM reasoning.
Dubbing quality control needs a reference-free judge of whether candidate text matches visible articulation in content and timing, using only silent video and text. Align Then Reason first builds a monotonic alignment between frame-level lip representations and phonetic units, then an LLM reasons over local evidence and a calibrated global score. On a seven-language benchmark, mean AUC improves 59.4%, 50.2%, and 50.8% over Qwen3.5 SFT baselines for 2B, 4B, and 9B reasoners. Gains also reach LLaMA-3.1-8B and Mistral-7B, transfer to unseen MuAViC languages, and improve dub-line reranking and script-to-clip assignment.
- ATR aligns lip frames to phonetic units before an LLM judges content and timing.
- Mean AUC rises about 50 to 59 percent over Qwen3.5 SFT baselines.
- Gains also appear for LLaMA-3.1-8B, Mistral-7B, and unseen MuAViC languages.
- ATR-9B beats lip-reading baselines by 52.0 percent on reranking and 17.7 percent on assignment.
Full article231 words · extracted from huggingface.co · click to collapse
Dubbing quality control requires a reference-free judge that can determine whether a candidate text line matches a speaker's visible articulation in both content and timing, using only silent video and text because dubbed audio may not yet exist. Existing visual speech recognizers and video-language models are poorly suited to this setting: even when fine-tuned to recover spoken content from lip motion, they remain largely insensitive to temporal errors. We introduce Align Then Reason (ATR), a multilingual lip-sync judge that first establishes a monotonic alignment between frame-level lip representations and the phonetic units of the candidate line, then reasons over this alignment to make the final judgment. An alignment scorer provides the LLM with both local evidence for each phonetic unit and a calibrated global alignment score, enabling it to reason jointly about content and timing. On a seven-language benchmark, our method improves mean AUC over the corresponding Qwen3.5 SFT baselines by 59.4%, 50.2%, and 50.8% with 2B, 4B, and 9B reasoners, respectively. The gains generalize across LLM families, reaching mean AUC improvements of 45.9% and 46.6% over the best baseline for LLaMA-3.1-8B and Mistral-7B, respectively. They also transfer across datasets to three unseen MuAViC languages. Furthermore, we evaluate on two downstream tasks built from real dubbing lines. On dub-line reranking, ATR-9B outperforms the best lip-reading baseline by 52.0%, while on script-to-clip assignment, ATR-9B improves over the best lip-reading baseline by 17.7%.
Text extracted automatically; images, tables and formatting may be missing. Original: https://huggingface.co/papers/2610.00825