[2603.27942] JaWildText: A Benchmark for Vision-Language Models on

[2603.27942] JaWildText: A Benchmark for Vision-Language Models on Japanese Scene Text Understanding

arXiv - AI March 31, 2026 4 min read

About this article

Abstract page for arXiv paper 2603.27942: JaWildText: A Benchmark for Vision-Language Models on Japanese Scene Text Understanding

Computer Science > Computer Vision and Pattern Recognition arXiv:2603.27942 (cs) [Submitted on 30 Mar 2026] Title:JaWildText: A Benchmark for Vision-Language Models on Japanese Scene Text Understanding Authors:Koki Maeda (1 and 2), Naoaki Okazaki (1 and 2) ((1) Institute of Science Tokyo, Tokyo, Japan, (2) Research and Development Center for Large Language Models, National Institute of Informatics, Tokyo, Japan) View a PDF of the paper titled JaWildText: A Benchmark for Vision-Language Models on Japanese Scene Text Understanding, by Koki Maeda (1 and 2) and 7 other authors View PDF HTML (experimental) Abstract:Japanese scene text poses challenges that multilingual benchmarks often fail to capture, including mixed scripts, frequent vertical writing, and a character inventory far larger than the Latin alphabet. Although Japanese is included in several multilingual benchmarks, these resources do not adequately capture the language-specific complexities. Meanwhile, existing Japanese visual text datasets have primarily focused on scanned documents, leaving in-the-wild scene text underexplored. To fill this gap, we introduce JaWildText, a diagnostic benchmark for evaluating vision-language models (VLMs) on Japanese scene text understanding. JaWildText contains 3,241 instances from 2,961 images newly captured in Japan, with 1.12 million annotated characters spanning 3,643 unique character types. It comprises three complementary tasks that vary in visual organization, output forma...

Originally published on March 31, 2026. Curated by AI News.

Llms

What if Claude purposefully made its own code leakable so that it would get leaked

What if Claude leaked itself by socially and architecturally engineering itself to be leaked by a dumb human submitted by /u/smurfcsgoawp...

Reddit - Artificial Intelligence · 1 min · 18 minutes ago

Llms

Observer-Embedded Reality

Observer-Embedded Reality Consciousness, Complexity, Meaning, and the Limits of Human Knowledge A Conceptual Philosophy-of-Science Paper ...

Reddit - Artificial Intelligence · 1 min · 18 minutes ago

Llms

I think we’re about to have a new kind of “SEO”… and nobody is talking about it.

More people are asking ChatGPT things like: “what’s the best CRM?” “is this tool worth it?” “alternatives to X” And they just… trust the ...

Reddit - Artificial Intelligence · 1 min · about 5 hours ago

Llms

Why would Claude give me the same response over and over and give others different replies?

I asked Claude to "generate me a random word" so I could do some word play. Then I asked it again in a new prompt window on desktop after...

Reddit - Artificial Intelligence · 1 min · about 5 hours ago

[2603.27942] JaWildText: A Benchmark for Vision-Language Models on Japanese Scene Text Understanding

About this article

Related Articles

What if Claude purposefully made its own code leakable so that it would get leaked

Observer-Embedded Reality

I think we’re about to have a new kind of “SEO”… and nobody is talking about it.

Why would Claude give me the same response over and over and give others different replies?

No comments

Stay updated with AI News