BenchMIRT: What are LLM benchmarks actually measuring?
Allen AI releases BenchMIRT, a method for auditing LLM benchmarks using multidimensional Item Response Theory. It analyzes individual prompts to separate underlying capabilities driving scores. The project includes a technical report, public dataset, and code repository. This is a research tool for evaluating benchmark validity rather than a new model or runtime.
