MUSE: Benchmarking Large Vision-Language Models on Multi-Modal Understanding in Situated Educationnew
Introduces MUSE, a twelve-task benchmark evaluating vision-language models on artistic image understanding in situated educational, Southeast Asian contexts.
MUSE is a benchmark assessing large vision-language models on artistic image understanding across twelve tasks spanning visual perception, semantic and affective interpretation, cultural understanding, and compositional reasoning. It decouples image annotation from question generation for controllable difficulty and curates images centering Singaporean and Southeast Asian multicultural contexts alongside Western art. Evaluations of open-source and proprietary models found substantial disparities, especially in affective interpretation and compositional reasoning.