[2512.03794] AdaptVision: Efficient Vision-Language Models via

[2512.03794] AdaptVision: Efficient Vision-Language Models via Adaptive Visual Acquisition

arXiv - Machine Learning March 03, 2026 4 min read

About this article

Abstract page for arXiv paper 2512.03794: AdaptVision: Efficient Vision-Language Models via Adaptive Visual Acquisition

Computer Science > Computer Vision and Pattern Recognition arXiv:2512.03794 (cs) [Submitted on 3 Dec 2025 (v1), last revised 28 Feb 2026 (this version, v2)] Title:AdaptVision: Efficient Vision-Language Models via Adaptive Visual Acquisition Authors:Zichuan Lin, Yicheng Liu, Yang Yang, Lvfang Tao, Deheng Ye View a PDF of the paper titled AdaptVision: Efficient Vision-Language Models via Adaptive Visual Acquisition, by Zichuan Lin and 4 other authors View PDF HTML (experimental) Abstract:Vision-Language Models (VLMs) have achieved remarkable success in visual question answering tasks, but their reliance on large numbers of visual tokens introduces significant computational overhead. While existing efficient VLM approaches reduce visual tokens through fixed-ratio compression, they operate passively and lack the ability to adapt to varying task requirements. This motivates a fundamental question: Can VLMs autonomously determine the minimum number of visual tokens required for each sample? Inspired by human active vision mechanisms, we introduce AdaptVision, an efficient VLM paradigm that enables adaptive visual token acquisition through a coarse-to-fine approach. Our model initially processes compressed visual tokens from low-resolution images and selectively acquires additional visual information by invoking a bounding box tool to crop key regions when necessary. We train AdaptVision using a reinforcement learning framework that carefully balances accuracy and efficiency. Cen...

Originally published on March 03, 2026. Curated by AI News.

Llms

Apple to open Siri to rival AI services beyond ChatGPT

Apple plans to open its Siri voice assistant to rival artificial intelligence (AI) services, moving beyond its partnership with OpenAI, a...

AI Tools & Products · 4 min · 8 minutes ago

Llms

Claude's scheduled tasks finally fixed what ChatGPT, Gemini, and every other AI tool got wrong

The boring stuff finally does itself.

AI Tools & Products · 9 min · 8 minutes ago

Llms

ChatGPT Just Got 33% More Accurate (The AI News You Missed)

ChatGPT has improved its accuracy by 33%, marking a notable enhancement for users of the AI platform.

AI Tools & Products · 1 min · 8 minutes ago

Llms

Exclusive | The Sudden Fall of OpenAI’s Most Hyped Product Since ChatGPT

The content discusses the sudden decline of OpenAI's most anticipated product since ChatGPT.

AI Tools & Products · 1 min · 9 minutes ago

[2512.03794] AdaptVision: Efficient Vision-Language Models via Adaptive Visual Acquisition

About this article

Related Articles

Apple to open Siri to rival AI services beyond ChatGPT

Claude's scheduled tasks finally fixed what ChatGPT, Gemini, and every other AI tool got wrong

ChatGPT Just Got 33% More Accurate (The AI News You Missed)

Exclusive | The Sudden Fall of OpenAI’s Most Hyped Product Since ChatGPT

No comments

Stay updated with AI News