상세 보기
Concept-Grounded Detection of Vaccine Misinformation in Multimodal Content Using Interpretable Vision-Language Models
- Thapa, Laxmi;
- Jain, Aryaman;
- Koduru, Lakshmojee;
- Adhikari, Surabhi;
- Rashid, Junaid;
- ... Kim, Jungeun;
- 외 2명
SCOPUS
0초록
Vaccine misinformation poses a persistent public health challenge, particularly in visual formats such as memes and infographics that combine text, imagery, and rhetorical cues. While textual misinformation has been widely studied, image-based vaccine misinformation remains comparatively underexplored due to the difficulty of interpreting multimodal signals at scale. In this work, we evaluate how effectively multimodal Large Vision-Language Models (LVLMs) can (i) directly classify vaccination stance from images and (ii) extract interpretable concept-level representations that support more reliable and transparent prediction. Using the VaxMeme dataset of 10,244 annotated images, we compare direct zero-shot LVLM inference against a hybrid framework in which classical machine learning models are trained on LVLM-extracted binary concept features. Our results show that grounding stance prediction in structured concept representations consistently outperforms direct LVLM classification, yielding accuracy improvements of approximately 10 - 17% while enabling explicit inspection of the visual and rhetorical cues driving model decisions. These findings highlight the value of concept-grounded, neuro-symbolic approaches for interpretable multimodal misinformation detection. © 2026 Owner/Author.
키워드
- 제목
- Concept-Grounded Detection of Vaccine Misinformation in Multimodal Content Using Interpretable Vision-Language Models
- 저자
- Thapa, Laxmi; Jain, Aryaman; Koduru, Lakshmojee; Adhikari, Surabhi; Rashid, Junaid; Kim, Jungeun; Thapa, Surendrabikram; Naseem, Usman
- 발행일
- 2026-05
- 유형
- Conference paper
- 저널명
- WWW Companion 2026 - Companion Proceedings of the ACM Web Conference 2026
- 페이지
- 830 ~ 838