|国家预印本平台
首页|SURE-VQA: Systematic Understanding of Robustness Evaluation in Medical VQA Tasks

SURE-VQA: Systematic Understanding of Robustness Evaluation in Medical VQA Tasks

SURE-VQA: Systematic Understanding of Robustness Evaluation in Medical VQA Tasks

来源:Arxiv_logoArxiv
英文摘要

Vision-Language Models (VLMs) have great potential in medical tasks, like Visual Question Answering (VQA), where they could act as interactive assistants for both patients and clinicians. Yet their robustness to distribution shifts on unseen data remains a key concern for safe deployment. Evaluating such robustness requires a controlled experimental setup that allows for systematic insights into the model's behavior. However, we demonstrate that current setups fail to offer sufficiently thorough evaluations. To address this gap, we introduce a novel framework, called SURE-VQA, centered around three key requirements to overcome current pitfalls and systematically analyze VLM robustness: 1) Since robustness on synthetic shifts does not necessarily translate to real-world shifts, it should be measured on real-world shifts that are inherent to the VQA data; 2) Traditional token-matching metrics often fail to capture underlying semantics, necessitating the use of large language models (LLMs) for more accurate semantic evaluation; 3) Model performance often lacks interpretability due to missing sanity baselines, thus meaningful baselines should be reported that allow assessing the multimodal impact on the VLM. To demonstrate the relevance of this framework, we conduct a study on the robustness of various Fine-Tuning (FT) methods across three medical datasets with four types of distribution shifts. Our study highlights key insights into robustness: 1) No FT method consistently outperforms others in robustness, and 2) robustness trends are more stable across FT methods than across distribution shifts. Additionally, we find that simple sanity baselines that do not use the image data can perform surprisingly well and confirm LoRA as the best-performing FT method on in-distribution data. Code is provided at https://github.com/IML-DKFZ/sure-vqa.

Kim-Celine Kahl、Selen Erkan、Jeremias Traub、Carsten T. Lüth、Klaus Maier-Hein、Lena Maier-Hein、Paul F. Jaeger

医学研究方法

Kim-Celine Kahl,Selen Erkan,Jeremias Traub,Carsten T. Lüth,Klaus Maier-Hein,Lena Maier-Hein,Paul F. Jaeger.SURE-VQA: Systematic Understanding of Robustness Evaluation in Medical VQA Tasks[EB/OL].(2025-07-03)[2025-07-16].https://arxiv.org/abs/2411.19688.点此复制

评论