[2602.04587] VILLAIN at AVerImaTeC: Verifying Image-Text Claims via Multi-Agent Collaboration
Summary
The paper presents VILLAIN, a multimodal fact-checking system that verifies image-text claims through collaborative agents, achieving top performance in the AVerImaTeC task.
Why It Matters
With the rise of misinformation, tools like VILLAIN are crucial for enhancing the reliability of information by automating the verification of image-text claims. This research contributes to the fields of AI and fact-checking, offering a scalable solution that can be applied in various contexts, such as news media and social platforms.
Key Takeaways
- VILLAIN employs multi-agent collaboration for fact-checking.
- The system ranked first in the AVerImaTeC evaluation metrics.
- It retrieves and analyzes both textual and visual evidence.
- Modality-specific agents generate reports to identify inconsistencies.
- The source code is publicly available for further research.
Computer Science > Computation and Language arXiv:2602.04587 (cs) [Submitted on 4 Feb 2026 (v1), last revised 20 Feb 2026 (this version, v2)] Title:VILLAIN at AVerImaTeC: Verifying Image-Text Claims via Multi-Agent Collaboration Authors:Jaeyoon Jung, Yejun Yoon, Kunwoo Park View a PDF of the paper titled VILLAIN at AVerImaTeC: Verifying Image-Text Claims via Multi-Agent Collaboration, by Jaeyoon Jung and 2 other authors View PDF HTML (experimental) Abstract:This paper describes VILLAIN, a multimodal fact-checking system that verifies image-text claims through prompt-based multi-agent collaboration. For the AVerImaTeC shared task, VILLAIN employs vision-language model agents across multiple stages of fact-checking. Textual and visual evidence is retrieved from the knowledge store enriched through additional web collection. To identify key information and address inconsistencies among evidence items, modality-specific and cross-modal agents generate analysis reports. In the subsequent stage, question-answer pairs are produced based on these reports. Finally, the Verdict Prediction agent produces the verification outcome based on the image-text claim and the generated question-answer pairs. Our system ranked first on the leaderboard across all evaluation metrics. The source code is publicly available at this https URL. Comments: Subjects: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Computers and Society (cs.CY) Cite as: arXiv:2602.04587 [cs.CL] (or ar...