[2504.04372] Assessing the Impact of Code Changes on the Fault

[2504.04372] Assessing the Impact of Code Changes on the Fault Localizability of Large Language Models

arXiv - Machine Learning March 06, 2026 4 min read

About this article

Abstract page for arXiv paper 2504.04372: Assessing the Impact of Code Changes on the Fault Localizability of Large Language Models

Computer Science > Software Engineering arXiv:2504.04372 (cs) [Submitted on 6 Apr 2025 (v1), last revised 5 Mar 2026 (this version, v4)] Title:Assessing the Impact of Code Changes on the Fault Localizability of Large Language Models Authors:Sabaat Haroon, Ahmad Faraz Khan, Ahmad Humayun, Waris Gill, Abdul Haddi Amjad, Ali R. Butt, Mohammad Taha Khan, Muhammad Ali Gulzar View a PDF of the paper titled Assessing the Impact of Code Changes on the Fault Localizability of Large Language Models, by Sabaat Haroon and 7 other authors View PDF HTML (experimental) Abstract:Generative Large Language Models (LLMs) are increasingly used in non-generative software maintenance tasks, such as fault localization (FL). Success in FL depends on a models ability to reason about program semantics beyond surface-level syntactic and lexical features. However, widely used LLM benchmarks primarily evaluate code generation, which differs fundamentally from semantic program reasoning. Meanwhile, traditional FL benchmarks such as Defect4J and BugsInPy are either not scalable or obsolete, as their datasets have become part of LLM training data, leading to biased results. This paper presents the first large-scale empirical investigation into the robustness of LLMs fault localizability. Inspired by mutation testing, we develop an end-to-end evaluation framework that addresses key limitations in existing LLM evaluation, including data contamination, scalability, automation, and extensibility. Using real-...

Originally published on March 06, 2026. Curated by AI News.

Llms

Claude Max 20x usage hit 40% by Monday noon — how does Codex CLI compare?

I'm on Claude Max (the $100/mo plan) and noticed something that surprised me. By Monday noon I had already used 40% of the 20x monthly li...

Reddit - Artificial Intelligence · 1 min · about 2 hours ago

Llms

How to use the new ChatGPT app integrations, including DoorDash, Spotify, Uber, and others | TechCrunch

Learn how to use Spotify, Canva, Figma, Expedia, and other apps directly in ChatGPT.

TechCrunch - AI · 10 min · about 5 hours ago

Llms

Anthropic Restricts Claude Agent Access Amid AI Automation Boom in Crypto

AI Tools & Products · 7 min · about 11 hours ago

Llms

Is cutting ‘please’ when talking to ChatGPT better for the planet? An expert explains

AI Tools & Products · 5 min · about 11 hours ago

[2504.04372] Assessing the Impact of Code Changes on the Fault Localizability of Large Language Models

About this article

Related Articles

Claude Max 20x usage hit 40% by Monday noon — how does Codex CLI compare?

How to use the new ChatGPT app integrations, including DoorDash, Spotify, Uber, and others | TechCrunch

Anthropic Restricts Claude Agent Access Amid AI Automation Boom in Crypto

Is cutting ‘please’ when talking to ChatGPT better for the planet? An expert explains

No comments

Stay updated with AI News