Anatomy-Aware Vision-Language Learning for Medical Image Interpretation
ID:98 View Protection:ATTENDEE Updated Time:2026-07-22 16:09:56 Hits:25 Online

Start Time:2026-07-30 16:10(Asia/Kolkata)

Duration:15min

Session:S4 Computer Vision and Pattern Recognition » S4-2Computer Vision and Pattern Recognition

Video No Permission Presentation File

Tips: Only the registered participant can access the file. Please sign in first.

Abstract

Vision-language modeling has significantly advanced radiology by enabling models that jointly learn from medical images and radiology reports for tasks such as disease classification, report generation, and visual question answering. However, most existing approaches treat an entire medical image as a single entity during image-text alignment, overlooking the fine-grained anatomical reasoning process employed by radiologists. In clinical practice, radiologists systematically examine individual anatomical regions, associate findings with specific structures, and integrate these region-specific observations before arriving at a conclusion about the image as a whole. To address this limitation, we propose an anatomy-aware vision-language framework that learns anatomy-specific representations using dedicated anatomy tokens and anatomical segmentation masks. The framework further incorporates context-aware anatomical representations and jointly learns anatomical localization, aligns anatomical regions with their corresponding findings, and aligns global image representations with image-level disease categories within a unified vision-language framework. Extensive experiments on out-of-distribution datasets demonstrate the effectiveness of the proposed framework. The model achieves strong performance in zero-shot disease classification and anatomical segmentation, demonstrating robust generalization to unseen data and accurate localization of anatomical structures. Comprehensive ablation studies further validate the contribution of each component and the effectiveness of the proposed design choices.

Keywords
Vision-Language Models,Anatomy-Aware Learning,Medical Image Interpretation,Zero-Shot Classification,Medical Image Segmentation
Speaker
Shahab Ahmad Khan
Student University of Wisconsin–Madison

Submission Author
Shahab Ahmad Khan University of Wisconsin–Madison
Mohammad Afzal Aligarh Muslim University
Syed Mohammad Suhaib Aligarh Muslim University
Mohd Anas Aftab Aligarh Muslim University
Submit Comment
Verify Code Change Another
All Comments
Important Date
  • Conference Date

    Jul 30

    2026

    to

    Aug 01

    2026

  • Jul 28 2026

    Draft paper submission deadline

  • Jul 28 2026

    Registration deadline

Sponsored By
The United Societies of Science
Organized By
Kongunadu College of Engineering and Technology
Supported By
IEEE Section
IEEE Madras Section
Previous Conferences