TY - GEN
T1 - RoI-MedCap
T2 - 2025 IEEE-EMBS International Conference on Biomedical and Health Informatics, BHI 2025
AU - Rubel, Al Shahriar
AU - Shih, Frank Y.
AU - Deek, Fadi P.
N1 - Publisher Copyright:
© 2025 IEEE.
PY - 2025
Y1 - 2025
N2 - Medical image captioning has gained significant attention due to the rapid advancements in Artificial Intelligence. However, existing research primarily focuses on global image captioning, lacking a mechanism for Region of Interest (RoI)-based captioning where users can specify an area and receive a caption centered on that specific region. In this paper, we propose a novel architecture with a vision encoder, a connector, and a Large Language Model (LLM) to generate captions for medical images with integrated RoI. We introduce a Multi-Stream Connector (MSC) to project visual features from a vision encoder to a representation that helps the LLM to generate captions centered on a specified region of an image indicated by a bounding box. We aim to generate captions with three aspects including the modality and structure, RoI analysis and lesion findings in RoI, and local-global relationship denoting impacts of findings in RoI to other regions. To achieve this goal, MSC incorporates three Cross Attentions focusing on three different aspects of generated captions. Our extensive experiments demonstrate that our method is well capable of generating captions highly aligned with human judgement, compared to existing related methods. The source code is available at https://github.com/alshahriarrubel/RoI-MedCap.
AB - Medical image captioning has gained significant attention due to the rapid advancements in Artificial Intelligence. However, existing research primarily focuses on global image captioning, lacking a mechanism for Region of Interest (RoI)-based captioning where users can specify an area and receive a caption centered on that specific region. In this paper, we propose a novel architecture with a vision encoder, a connector, and a Large Language Model (LLM) to generate captions for medical images with integrated RoI. We introduce a Multi-Stream Connector (MSC) to project visual features from a vision encoder to a representation that helps the LLM to generate captions centered on a specified region of an image indicated by a bounding box. We aim to generate captions with three aspects including the modality and structure, RoI analysis and lesion findings in RoI, and local-global relationship denoting impacts of findings in RoI to other regions. To achieve this goal, MSC incorporates three Cross Attentions focusing on three different aspects of generated captions. Our extensive experiments demonstrate that our method is well capable of generating captions highly aligned with human judgement, compared to existing related methods. The source code is available at https://github.com/alshahriarrubel/RoI-MedCap.
KW - Artificial Intelligence
KW - Cross Attention
KW - Large Language Model (LLM)
KW - Medical Image Captioning
KW - Multi-Stream Connector (MSC)
KW - Radiology Report Generation
KW - Region of Interest (RoI)
KW - Vision Language Model (VLM)
UR - https://www.scopus.com/pages/publications/105030450263
UR - https://www.scopus.com/pages/publications/105030450263#tab=citedBy
U2 - 10.1109/BHI67747.2025.11269516
DO - 10.1109/BHI67747.2025.11269516
M3 - Conference contribution
AN - SCOPUS:105030450263
T3 - BHI 2025 - IEEE-EMBS International Conference on Biomedical and Health Informatics, Conference Proceedings
BT - BHI 2025 - IEEE-EMBS International Conference on Biomedical and Health Informatics, Conference Proceedings
PB - Institute of Electrical and Electronics Engineers Inc.
Y2 - 26 October 2025 through 29 October 2025
ER -