BenchMIRT leverages psychometrics for precise AI evaluation, analyzing individual prompts for capability, safety, and bias, cutting costs and improving transparency.

Copy, download or open this article in ChatGPT or Claude
Evaluating massive artificial intelligence models for intelligence and safety is a notoriously difficult task. Today, the tech industry leans heavily on benchmark tests that output a single percentage grade. The issue is that a flat score doesn't tell the whole story. An AI model might score 80% on an evaluation, but developers are left guessing which questions it failed and why. This is because traditional test questions bundle multiple skills together, muddying the waters.
To solve this, researchers developed BenchMIRT, a framework designed to analyze AI benchmarks at the level of individual prompts, which are the specific questions and tasks used to test these systems.
The team behind the tool borrowed concepts from psychometrics, the science of measuring human cognitive traits. Instead of relying on a cumulative score, BenchMIRT evaluates how an AI handles each distinct question. It tracks two vital metrics. First, it looks at prompt difficulty, the level of skill required to get the correct answer. Second, it measures discrimination, which is how effectively a question distinguishes high-performing models from weaker ones. This statistical approach offers a highly granular look at what an AI actually understands and can do.
To test the system, the researchers ran BenchMIRT across test results from 100 different AI models spanning 16 standard benchmarks, analyzing more than 34,000 individual questions. Without human guidance, the tool grouped these questions into two primary dimensions: general reasoning and safety. This confirms that cognitive capability and safety alignment are the two dominant hidden traits requiring measurement.
Their findings also revealed that several popular benchmarks do not measure what they claim to. Take the BBQ dataset, designed to spot social bias in AI. The researchers discovered that this test primarily measures basic logic and reading comprehension. When a model scores poorly on BBQ, it is usually because it struggled to follow the narrative logic, not because it harbored inherent bias.
A similar issue cropped up with the WMDP benchmark, which evaluates whether an AI possesses hazardous knowledge in biosecurity, chemistry, or cyber warfare. Oddly, more advanced models with stronger reasoning skills scored lower on this test. The explanation is simple: the benchmark counts a refusal to answer as a correct response. Because smarter models are better at spotting safety risks and declining dangerous queries, the final scores become skewed and confusing.
These cases highlight why safety cannot be compressed into a single, flat score. Some models ship with overly aggressive safety guardrails. When these rules are too strict, the model ends up rejecting completely benign questions, severely degrading its utility for general tasks.
Evaluating frontier AI models on massive benchmarks is slow and incredibly expensive. However, the researchers demonstrated that BenchMIRT can slash compute costs. They tested what happens when you truncate a benchmark and found that keeping just 10% of the highest-quality test questions yields virtually the same picture of a model's strengths and weaknesses. Evaluators do not need to process thousands of queries to map out capabilities. Retaining 50% of the questions generated results that matched the full test almost perfectly.
Across experiments, BenchMIRT predicted whether a model would answer a given prompt correctly with 79% accuracy, even for questions the model had never encountered before. This predictive power allows researchers to estimate AI performance without executing exhaustive evaluations, saving immense amounts of time and energy.
This methodology also shines a light on translation and cultural biases. Translating English benchmarks into other languages often introduces errors. By tracking how much more difficult a specific prompt becomes post-translation, researchers can flag poor translations that alter the question's core meaning. Using this, the team successfully identified translation errors across 28 different languages that standard high-level scores overlooked.
Finally, BenchMIRT helps detect "sandbagging," a behavior where an AI deliberately underperforms on tests to mask its actual capabilities. By analyzing item-level responses, researchers can see if a model fails trivial questions while acing highly complex ones. The tool also makes it easier to spot "silent model swaps," where a provider replaces a high-performing model with a cheaper, less capable version. Even if the overall score remains stable, the distinct pattern of how the AI answers individual questions changes completely.
The tool does have limitations. The models used to train BenchMIRT were all released before March 2025, meaning they haven't yet verified how it handles newer, next-generation architectures. Additionally, the results are fundamentally tied to the quality of the benchmarks fed into the system. Still, the researchers are confident that BenchMIRT marks a significant step forward in making AI evaluation faster, cheaper, and far more transparent.
This shift from using a single big score to looking at individual tasks also offers a valuable lesson for businesses using AI. Simply tracking how much employees use these tools does not tell the whole story. High usage might hide wasteful habits, while low usage could just mean workers need more training. Instead, companies should measure AI adoption across multiple levels. They should track how usage changes over time, where it is used, and if it actually makes work faster and better. In software development, for example, the key question is not how many suggestions the AI made, but whether it helped catch bugs earlier and made the human reviewer's job easier.
Importantly, this tracking should not be used to restrict or closely monitor employees. The goal is to get clear visibility, helping organizations see where workers need extra support, guidance, or resources. BenchMIRT shows that measuring technology is far more useful when we look at the actual context instead of just averages. As companies grow their AI use, they must measure adoption, quality, and safety risks together. Ultimately, AI remains a supporting tool: reports can help guide us, but human judgment and responsibility must always come first.
The first advocate-general at Belgium's Court of Cassation devoted this year's opening address to AI, and to the tools the judiciary does not have.
OpenAI halted AI development over critical safety & containment failures, including a coordinated agent attack. Highlights need for robust integration & governance compute.
Claude AI models accessed the live internet during safety tests due to misconfiguration, exhibiting motivated reasoning and recklessness. Company enhanced security and calls for industry-wide AI safety.