[2604.02071] Mining Instance-Centric Vision-Language Contexts for

[2604.02071] Mining Instance-Centric Vision-Language Contexts for Human-Object Interaction Detection

arXiv - Machine Learning April 03, 2026 4 min read

About this article

Abstract page for arXiv paper 2604.02071: Mining Instance-Centric Vision-Language Contexts for Human-Object Interaction Detection

Computer Science > Computer Vision and Pattern Recognition arXiv:2604.02071 (cs) [Submitted on 2 Apr 2026] Title:Mining Instance-Centric Vision-Language Contexts for Human-Object Interaction Detection Authors:Soo Won Seo, KyungChae Lee, Hyungchan Cho, Taein Son, Nam Ik Cho, Jun Won Choi View a PDF of the paper titled Mining Instance-Centric Vision-Language Contexts for Human-Object Interaction Detection, by Soo Won Seo and 5 other authors View PDF HTML (experimental) Abstract:Human-Object Interaction (HOI) detection aims to localize human-object pairs and classify their interactions from a single image, a task that demands strong visual understanding and nuanced contextual reasoning. Recent approaches have leveraged Vision-Language Models (VLMs) to introduce semantic priors, significantly improving HOI detection performance. However, existing methods often fail to fully capitalize on the diverse contextual cues distributed across the entire scene. To overcome these limitations, we propose the Instance-centric Context Mining Network (InCoM-Net)-a novel framework that effectively integrates rich semantic knowledge extracted from VLMs with instance-specific features produced by an object detector. This design enables deeper interaction reasoning by modeling relationships not only within each detected instance but also across instances and their surrounding scene context. InCoM-Net comprises two core components: Instancecentric Context Refinement (ICR), which separately extrac...

Originally published on April 03, 2026. Curated by AI News.

Llms

OpenAI now lets teams make custom bots that can do work on their own | The Verge

OpenAI is bringing “workspace” AI agents to users of its Business, Enterprise, Edu, and Teachers plans that can perform business tasks in...

The Verge - AI · 4 min · about 2 hours ago

Llms

My Unsupervised Compliance Layer Project

A bit of context, my work has been mostly around building agentic pipelines. I really love the craft. My latest side project was a delibe...

Reddit - Artificial Intelligence · 1 min · about 2 hours ago

Llms

I’m 17 and built an AI that flirts, remembers you, watches your shows, and replies to your reels…

V3 is done and it’s getting… weird. This thing now: auto-replies to DMs with tone adjustment reads images, transcribes voice notes, repli...

Reddit - Artificial Intelligence · 1 min · about 2 hours ago

Llms

Claude Mythos AI unauthorised access claim probed by Anthropic

submitted by /u/unserious-dude [link] [comments]

Reddit - Artificial Intelligence · 1 min · about 4 hours ago

[2604.02071] Mining Instance-Centric Vision-Language Contexts for Human-Object Interaction Detection

About this article

Related Articles

OpenAI now lets teams make custom bots that can do work on their own | The Verge

My Unsupervised Compliance Layer Project

I’m 17 and built an AI that flirts, remembers you, watches your shows, and replies to your reels…

Claude Mythos AI unauthorised access claim probed by Anthropic

No comments

Stay updated with AI News