0% Complete
صفحه اصلی
/
دومین همایش بین المللی هوش مصنوعی
Zero-Shot, Standard Fine-Tuning, and Curriculum Learning Approaches for VQA in GI Endoscopy
نویسندگان :
Mahdi Azmoodeh-Kalati
1
Mohammad Sadegh Maghareh
2
Reza Lashgari
3
1- دانشگاه شهید بهشتی
2- دانشگاه امیرکبیر
3- دانشگاه بهشتی
کلمات کلیدی :
Visual Question Answering،Gastrointestinal Endoscopy،Curriculum Learning،Fine-Tuning،Vision-Language Models
چکیده :
Visual Question Answering (VQA) on gastrointestinal (GI) endoscopy images is a challenging multimodal task due to complex visual content and specialized clinical language. This paper presents a comparative study of three inference strategies for GI endoscopy VQA using a large vision–language model: (1) zero-shot inference without domain training, (2) conventional fine-tuning on the target dataset, and (3) curriculum-guided fine-tuning with progressively harder training stages. Building on our prior work, we adapt a state-of-the-art 2-billion-parameter vision–language model to the Kvasir-VQA-x1 dataset (6,500 images, ~159k QA pairs) using Low-Rank Adaptation (LoRA). Experimental results on a held-out evaluation set demonstrate that domain-specific fine-tuning yields substantial performance gains over the zero-shot baseline (e.g., BLEU improves from 0.0700 to 0.2513). Moreover, curriculum learning further boosts accuracy across metrics (e.g., BLEU 0.2667), indicating better generalization to diverse question complexities. We describe the three-stage curriculum setup in detail, where each “easy” question is seen three times and “moderate” questions twice during training, potentially reinforcing core knowledge. The proposed curriculum-guided strategy achieves more balanced VQA performance across simple and complex questions than standard fine-tuning. We discuss the impact of repeated exposure to easier samples on learning and outline future steps to vary the curriculum schedule. Overall, our work highlights curriculum learning as an effective method to improve GI endoscopy VQA, outperforming both zero-shot and standard fine-tuning approaches in this domain.
لیست مقالات
لیست مقالات بایگانی شده
AI-CADx in Retinal Disease: Integrating Explainability, Privacy, and Scalability for Next-Generation Telemedicine
Mohammad Shojaeinia - Hamid Moghaddasi
Cross-domain prediction of ultimate strength in corroded reinforced concrete columns using domain adaptation techniques
Melina Sadeghi - Pooria Poorahad - Mahmoud R. Shiravand
Hybrid Deep Learning Models for Cardiovascular Disease Prediction: A Comprehensive Review of Convolution-Transformer Architectures
Ali Azimi Lamir - Masoud Bekravi - Babak Nouri Moghaddam
Creating a Foundation for Dynamic Difficulty Adjustment within PCG of games using Imitation Learning
Navid Siamakmanesh - Arian Ganji - Monireh Abdoos - Mojtaba Vahidi-Asl
Finite element model updating using computational intelligence - based methods: A Case study of the San Roque Valley bridge, California
Fatemeh Z Dokhanian - Negin G.sabagh - Mahmoud R Shiravand
Ensemble Machine Learning for Predicting Stroke Patient Survival
Seyedeh Maryam Mousavi - Samira Ahmadi - Solmaz Norouzi
A Master-Slave Approach for Simultaneously Controlling Two Drones when Carrying an Object
Seyyed Mohammad Ali Ardehali - Amin Faraji - Monireh Abdoos - Armin Salimi-Badr
Computational Complexity of Sentiment Analysis Algorithms in Natural Language Processing
Kiana Karimifard - Mohammad Ghasemzadeh
Robust Algorithmic Trading in Volatile Markets Using Pessimistic Transformer-Enhanced TD3 in Continuous Action Spaces
Amirhossein Ghozati - Armin Salimi-Badr
Potential of machine learning algorithms for predicting the properties of medium-density fiberboard (MDF): preliminary results
Rahim Mohebbi Gargari - Ali Shalbafan - Seyed Jalil Alavi - Maryam Amirmazlaghni - Seyed Hamzeh Sadatnejad - Heiko Thoemen
بیشتر
ثمین همایش، سامانه مدیریت کنفرانس ها و جشنواره ها - نگارش 44.5.0