Grounded Semantic Quality and Diversity Diagnostics for Image Caption Evaluation
ID:106
View Protection:ATTENDEE
Updated Time:2026-07-22 16:10:01 Hits:14
Online
Abstract
Image caption generation has advanced significantly with vision-language models which generate fluent, meaningful description for the Image. The conventional image captioning metrics like BLEU, METEOR, ROUGE and CIDEr mainly evaluate based on lexical or consensus-based similarity with the ground truth captions and provide limited explanation about the semantic grounding of the generated caption. To address this problem, Grounded Semantic Quality and Diversity diagnostic framework is proposed which not only evaluates the semantic grounding but also decompose them in categories to understand the reason of weak semantic grounding. To evaluate candidate set caption, Grounded Semantic Diversity metric is introduced. In addition to evaluation, the experiments to use the metrics for reranking and optimization are demonstrated to show its effectiveness for the purpose. Experimental results show that when GSQ is incorporated as a reward component in self-critical sequence training, the fine-tuned model improves CIDEr by 19.3% and object-F1 by 10.6%, while reducing the unsupported semantic rate by 19.6% compared with the BLIP baseline.
Keywords
Image captioning Metrics,Knowledge Graph,Reranking,,Entity awareness,Hallucination mitigation,Context awareness,Diversity in image captioning
Submission Author
Sharmila Kharat
MIT Academy of Engineering Alandi Pune
Sunita Barve
MIT Academy of Engineering Alandi Pune
Submit Comment