arXiv cs.AI / cs.LG / cs.CL·2d agoThe Alignment Illusion in Multimodal Large Language Models#mllm#multimodal#representation-alignmentAI research
arXiv cs.CR·3d agoGUIAuditor: Enabling Post-hoc Child Safety Forensics via Action-Guided GUI Provenance on Mobile Devices#forensics#mobile#guiResearch
arXiv cs.AI / cs.LG / cs.CL·5d agoSLICEChat: Progressive In-Encoder Token Pruning for Whole-Slide Pathology Language Models#slicechat#pathology#token-pruningAI research1
Hugging Face daily papers·7d agoPackLab: A Comprehensive Framework for Developing, Training, and Evaluating MLLMs in Robotic Bin Packing#mllm#robotics#bin-packing
arXiv cs.CR·9d agoFingerprinting Multimodal Large Language Models#attention#distillation#fingerprintingAI safety & security
Hugging Face daily papers·10d agoRegion-Level Policy Optimization for Fine-grained MLLM Perception#fine-grained-understanding#mllm#reinforcement-learning
Hugging Face daily papers·10d agoUFO: Chain-of-Evaluation for Omni-Condition Alignment in Multi-Modal Image Generation#benchmark#evaluation-framework#image-generation2
Hugging Face daily papers·10d agoVABench: Measuring Embodied Spatial Intelligence through Visual Demonstrations, Active Perception, and Metric Control#benchmark#embodied-ai#evaluation
Hugging Face daily papers·13d agoHarnessVLN: Unifying Training-Free Embodied Navigation through an Agent Harness#agent-harness#embodied-ai#mllm
Hugging Face daily papers·14d agoOmniHarness: Harnessing Generalizable Visual Generation via Symbolic Policy Learning#agents#mllm#multimodal1
arXiv cs.AI / cs.LG / cs.CL·15d agoEmbodied-BenchForge: A Closed-Loop Agentic Workflow for Embodied Benchmark Construction#agents#automation#benchmarksAI research1
Hugging Face daily papers·20d agoReactVAU: A Slow-Fast Decoupled Framework for Streaming Video Anomaly Understanding#efficient-inference#mllm#real-time-inference
Hugging Face daily papers·22d agoVDiff-Bench: A Challenging Benchmark for Fine-Grained Image Difference Identification#benchmark#evaluation#image-difference
Hugging Face daily papers·22d agoMLLMs Hallucinate when Information Distribution Drifts in Synergy Heads#attention-heads#calibration#hallucination
Hugging Face daily papers·25d agoShallowStream: Index Shallow then Answer Deep for Streaming Video Understanding#inference-efficiency#kv-cache#mllm