[2603.24999] Efficient Detection of Bad Benchmark Items with Novel

[2603.24999] Efficient Detection of Bad Benchmark Items with Novel Scalability Coefficients

arXiv - AI March 27, 2026 4 min read

About this article

Abstract page for arXiv paper 2603.24999: Efficient Detection of Bad Benchmark Items with Novel Scalability Coefficients

Statistics > Applications arXiv:2603.24999 (stat) [Submitted on 26 Mar 2026] Title:Efficient Detection of Bad Benchmark Items with Novel Scalability Coefficients Authors:Michael Hardy, Joshua Gilbert, Benjamin Domingue View a PDF of the paper titled Efficient Detection of Bad Benchmark Items with Novel Scalability Coefficients, by Michael Hardy and 2 other authors View PDF Abstract:The validity of assessments, from large-scale AI benchmarks to human classrooms, depends on the quality of individual items, yet modern evaluation instruments often contain thousands of items with minimal psychometric vetting. We introduce a new family of nonparametric scalability coefficients based on interitem isotonic regression for efficiently detecting globally bad items (e.g., miskeyed, ambiguously worded, or construct-misaligned). The central contribution is the signed isotonic $R^2$, which measures the maximal proportion of variance in one item explainable by a monotone function of another while preserving the direction of association via Kendall's $\tau$. Aggregating these pairwise coefficients yields item-level scores that sharply separate problematic items from acceptable ones without assuming linearity or committing to a parametric item response model. We show that the signed isotonic $R^2$ is extremal among monotone predictors (it extracts the strongest possible monotone signal between any two items) and show that this optimality property translates directly into practical screening...

Originally published on March 27, 2026. Curated by AI News.

Llms

Google Launches Gemini Import Tools to Poach Users From Rival AI Apps

Anyone looking to switch their AI assistant will find it surprisingly easy, as it only takes a few steps to move from A to B. This is not...

AI Tools & Products · 4 min · about 3 hours ago

Ai Startups

Could factories run faster and greener? How AI 'digital twins' reshape production

Researchers at Örebro University have developed a new production system that uses artificial intelligence (AI) to improve efficiency and ...

Reddit - Artificial Intelligence · 1 min · about 4 hours ago

Llms

[2603.11687] SemBench: A Universal Semantic Framework for LLM Evaluation

Abstract page for arXiv paper 2603.11687: SemBench: A Universal Semantic Framework for LLM Evaluation

arXiv - AI · 4 min · about 8 hours ago

Llms

[2603.11413] Evaluation format, not model capability, drives triage failure in the assessment of consumer health AI

Abstract page for arXiv paper 2603.11413: Evaluation format, not model capability, drives triage failure in the assessment of consumer he...

arXiv - AI · 4 min · about 8 hours ago

[2603.24999] Efficient Detection of Bad Benchmark Items with Novel Scalability Coefficients

About this article

Related Articles

Google Launches Gemini Import Tools to Poach Users From Rival AI Apps

Could factories run faster and greener? How AI 'digital twins' reshape production

[2603.11687] SemBench: A Universal Semantic Framework for LLM Evaluation

[2603.11413] Evaluation format, not model capability, drives triage failure in the assessment of consumer health AI

No comments

Stay updated with AI News