RCP-nDCG@10: A New Approach to Enterprise Retrieval Quality | Cohere

Cohere introduces Rubric-Calibrated Preferences nDCG@10, an evaluation methodology designed to measure enterprise retrieval quality using explicit relevance criteria and a calibrated AI judge. The approach addresses limitations in traditional nDCG by avoiding reliance on incomplete pre-existing relevance labels, providing a new way to benchmark search model performance.

Cover image for RCP-nDCG@10: A New Approach to Enterprise Retrieval Quality | Cohere

RCP‑nDCG@10 is a metric that grades every document returned by an enterprise search system against a common rubric, using a calibrated AI judge. By scoring each result directly, it can award credit to relevant items that were absent from the original relevance pool. Cohere reports that this approach yields a calibrated relevance score that separates useful from useless material with an AUC of 0.91, compared with 0.65 for the traditional qrel‑based score. The authors conducted a 46‑person human annotation study in which reviewers judged the usefulness of retrieved documents. They compared the human grades to the scores produced by the calibrated AI judge, observing that 28 % of documents labelled irrelevant by the benchmark were judged useful, while only 15 % of benchmark‑labelled relevant items failed a basic topicality test. The study also noted that the benchmark labelled just 29 % of documents as relevant, whereas humans suggested relevance for about 60 % of the same set. RCP‑nDCG@10 still depends on the quality of the rubric and the calibration of the AI judge, so systematic biases in those components are not captured. The metric does not reveal disagreements among human annotators; the source reports exact agreement on a 0‑4 scale for only 42 % of pairs, with some studies finding as low as 17 % agreement. Consequently, the score cannot fully reflect the inherent subjectivity of relevance judgments.