[2603.27593] STRIDE: When to Speak Meets Sequence Denoising for

[2603.27593] STRIDE: When to Speak Meets Sequence Denoising for Streaming Video Understanding

arXiv - AI March 31, 2026 3 min read

About this article

Abstract page for arXiv paper 2603.27593: STRIDE: When to Speak Meets Sequence Denoising for Streaming Video Understanding

Computer Science > Computer Vision and Pattern Recognition arXiv:2603.27593 (cs) [Submitted on 29 Mar 2026] Title:STRIDE: When to Speak Meets Sequence Denoising for Streaming Video Understanding Authors:Junho Kim, Hosu Lee, James M. Rehg, Minsu Kim, Yong Man Ro View a PDF of the paper titled STRIDE: When to Speak Meets Sequence Denoising for Streaming Video Understanding, by Junho Kim and 4 other authors View PDF HTML (experimental) Abstract:Recent progress in video large language models (Video-LLMs) has enabled strong offline reasoning over long and complex videos. However, real-world deployments increasingly require streaming perception and proactive interaction, where video frames arrive online and the system must decide not only what to respond, but also when to respond. In this work, we revisit proactive activation in streaming video as a structured sequence modeling problem, motivated by the observation that temporal transitions in streaming video naturally form span-structured activation patterns. To capture this span-level structure, we model activation signals jointly over a sliding temporal window and update them iteratively as new frames arrive. We propose STRIDE (Structured Temporal Refinement with Iterative DEnoising), which employs a lightweight masked diffusion module at the activation interface to jointly predict and progressively refine activation signals across the window. Extensive experiments on diverse streaming benchmarks and downstream models demonst...

Originally published on March 31, 2026. Curated by AI News.

Llms

What if Claude purposefully made its own code leakable so that it would get leaked

What if Claude leaked itself by socially and architecturally engineering itself to be leaked by a dumb human submitted by /u/smurfcsgoawp...

Reddit - Artificial Intelligence · 1 min · about 2 hours ago

Llms

Observer-Embedded Reality

Observer-Embedded Reality Consciousness, Complexity, Meaning, and the Limits of Human Knowledge A Conceptual Philosophy-of-Science Paper ...

Reddit - Artificial Intelligence · 1 min · about 2 hours ago

Llms

I think we’re about to have a new kind of “SEO”… and nobody is talking about it.

More people are asking ChatGPT things like: “what’s the best CRM?” “is this tool worth it?” “alternatives to X” And they just… trust the ...

Reddit - Artificial Intelligence · 1 min · about 6 hours ago

Llms

Why would Claude give me the same response over and over and give others different replies?

I asked Claude to "generate me a random word" so I could do some word play. Then I asked it again in a new prompt window on desktop after...

Reddit - Artificial Intelligence · 1 min · about 6 hours ago

[2603.27593] STRIDE: When to Speak Meets Sequence Denoising for Streaming Video Understanding

About this article

Related Articles

What if Claude purposefully made its own code leakable so that it would get leaked

Observer-Embedded Reality

I think we’re about to have a new kind of “SEO”… and nobody is talking about it.

Why would Claude give me the same response over and over and give others different replies?

No comments

Stay updated with AI News