Hugging Face daily papers·8d agoSpatial-Interactor: Learning Spatial Reasoning through Interaction with the Observable Physical World#spatial-reasoning#vision-language-models#curriculum-learning
arXiv cs.AI / cs.LG / cs.CL·8d agoDiaVLo: Diagnosing Behaviours of Vision-Language Models#vision-language-models#diagnostics#alignmentAI research
arXiv cs.AI / cs.LG / cs.CL·8d agoWhen Should a Failing Robot Ask? Initiating Corrective Human-Robot Dialogue from Audited Sensor Evidence#vision-language-models#human-robot#prompt-sensitivityAI research
arXiv cs.AI / cs.LG / cs.CL·9d agoCross-Modal Attention Acts as a Frequency Filter: Why Verbose Prompts Improve Robustness in Vision-Language Models#cross-modal-attention#image-corruption#llava-onevisionAI research1
arXiv cs.AI / cs.LG / cs.CL·10d agoPANORAMA: Panoptic Grounded Captioning via Mask Proposal Selection#benchmark#grounded-captioning#panocapsAI research
arXiv cs.AI / cs.LG / cs.CL·10d agoReporting Practice Matters: The Impact of Reference Choice on Chest X-ray Report Evaluation#benchmark#evaluation-metrics#mimic-cxrAI research
arXiv cs.AI / cs.LG / cs.CL·10d agoMUSE: Benchmarking Large Vision-Language Models on Multi-Modal Understanding in Situated Education#benchmark#cultural-understanding#educationAI research1
arXiv cs.AI / cs.LG / cs.CL·11d agoBridging the Confidence Gap: Temperature Scaling for Calibrating Test-Time Prompt Tuning#calibration#temperature-scaling#test-time-prompt-tuningAI research1
arXiv cs.CR·12d agoDon't Send What You Don't Need: Question-Guided Token Pruning as a Privacy Defense for Vision-Language Models#federated-learning#membership-inference#privacyResearch1
Hugging Face daily papers·14d agoE2A-Bench: Benchmarking Evidence-to-Action Reliability in Financial Chart Reasoning#benchmark#e2a-bench#financial-reasoning
arXiv cs.AI / cs.LG / cs.CL·16d agoCan Edge-Deployable Vision-Language Models Identify Species?#benchmark#bioclip#camera-trapsAI research1
Hugging Face daily papers·17d agoFeature Recovery for Object Understanding After Irreversible Fire Damage#benchmark#computer-vision#degradation
Hugging Face daily papers·18d agoShow-Harness: Just a VLM Agent Can Play Robots#agent-framework#embodied-agents#robot-control
Hugging Face daily papers·18d agoThink Before You Link: Rarity, Reasoning, and Retrieval in Multilingual Entity Linking#benchmark#entity-linking#knowledge-graph1
arXiv cs.AI / cs.LG / cs.CL·18d agoGoDeep: Annotation-Free Open-Vocabulary 3D Scene Understanding via Language-Space Lifting#3d-scene-understanding#clip#embeddingsAI research
Hugging Face daily papers·19d agoCoVeR: Coverage-Based Token Pruning for Multi-View 3D Reasoning in VLMs#3d-reasoning#efficiency#multimodal
Hugging Face daily papers·20d agoSpatialBlock: Enhancing Spatial Intelligence in LVLMs via Synthetic Block-Stacking Problem#3d-reasoning#dataset#lvlm
Hugging Face daily papers·21d agoReason Through the Latent! Making Latent Visual Reasoning Necessary#causal-analysis#latent-reasoning#multimodal
Hugging Face daily papers·23d agoKnowing What Not to Answer: Selective Non-Compliance in Vision-Language Models#benchmark#fine-tuning#k1
Hugging Face daily papers·Aug 24, 2026Real-World Knowledge-Guided Change Data Synthesis for Remote Sensing#change-detection#data-synthesis#pretrained-models