MP-Bench: Evaluating Voice Agents as a Multiparty Conversation Participant
MP-Bench is the first benchmark for voice agents in multiparty conversations, finding real-time agents near chance on turn-taking.
MP-Bench is the first benchmark designed to objectively evaluate conversational speech systems as active participants in multi-party conversations. It assesses agents on turn-taking awareness and response appropriateness, with comprehension-based question-answering as a complementary evaluation. Benchmarking 12 voice agents shows real-time agents score at or below 22% on multiparty comprehension and remain near chance on multiparty turn-taking.