[2508.18473] Principled Detection of Hallucinations in Large Language

[2508.18473] Principled Detection of Hallucinations in Large Language Models via Multiple Testing

arXiv - Machine Learning April 29, 2026 3 min read

About this article

Abstract page for arXiv paper 2508.18473: Principled Detection of Hallucinations in Large Language Models via Multiple Testing

Computer Science > Computation and Language arXiv:2508.18473 (cs) [Submitted on 25 Aug 2025 (v1), last revised 28 Apr 2026 (this version, v3)] Title:Principled Detection of Hallucinations in Large Language Models via Multiple Testing Authors:Jiawei Li, Akshayaa Magesh, Venugopal V. Veeravalli View a PDF of the paper titled Principled Detection of Hallucinations in Large Language Models via Multiple Testing, by Jiawei Li and 2 other authors View PDF HTML (experimental) Abstract:While Large Language Models (LLMs) have emerged as powerful foundational models to solve a variety of tasks, they have also been shown to be prone to hallucinations, i.e., generating responses that sound confident but are actually incorrect or even nonsensical. Existing hallucination detectors propose a wide range of empirical scoring rules, but their performance varies across models and datasets, and it is hard to determine which ones to rely on in practice or to treat as a reliable detector. In this work, we formulate the problem of detecting hallucinations as a hypothesis testing problem and draw parallels with the problem of out-of-distribution detection in machine learning models. We then propose a multiple-testing-inspired method that systematically aggregates multiple evaluation scores via conformal p-values, enabling calibrated detection with controlled false alarm rate. Extensive experiments across diverse models and datasets validate the robustness of our approach against state-of-the-art m...

Originally published on April 29, 2026. Curated by AI News.

Llms

[2604.16909] PRISM: Probing Reasoning, Instruction, and Source Memory in LLM Hallucinations

Abstract page for arXiv paper 2604.16909: PRISM: Probing Reasoning, Instruction, and Source Memory in LLM Hallucinations

arXiv - AI · 4 min · about 1 hour ago

Llms

[2604.07802] Latent Anomaly Knowledge Excavation: Unveiling Sparse Sensitive Neurons in Vision-Language Models

Abstract page for arXiv paper 2604.07802: Latent Anomaly Knowledge Excavation: Unveiling Sparse Sensitive Neurons in Vision-Language Models

arXiv - AI · 4 min · about 1 hour ago

Llms

[2602.07605] Fine-R1: Make Multi-modal LLMs Excel in Fine-Grained Visual Recognition by Chain-of-Thought Reasoning

Abstract page for arXiv paper 2602.07605: Fine-R1: Make Multi-modal LLMs Excel in Fine-Grained Visual Recognition by Chain-of-Thought Rea...

arXiv - AI · 4 min · about 1 hour ago

Llms

[2602.07096] RealFin: How Well Do LLMs Reason About Finance When Users Leave Things Unsaid?

Abstract page for arXiv paper 2602.07096: RealFin: How Well Do LLMs Reason About Finance When Users Leave Things Unsaid?

arXiv - AI · 3 min · about 1 hour ago

[2508.18473] Principled Detection of Hallucinations in Large Language Models via Multiple Testing

About this article

Related Articles

[2604.16909] PRISM: Probing Reasoning, Instruction, and Source Memory in LLM Hallucinations

[2604.07802] Latent Anomaly Knowledge Excavation: Unveiling Sparse Sensitive Neurons in Vision-Language Models

[2602.07605] Fine-R1: Make Multi-modal LLMs Excel in Fine-Grained Visual Recognition by Chain-of-Thought Reasoning

[2602.07096] RealFin: How Well Do LLMs Reason About Finance When Users Leave Things Unsaid?

No comments

Stay updated with AI News