• Decrease Text SizeIncrease Text Size

Benchmark Contamination

Public benchmarks circulate widely and end up in training corpora, after which a high score may reflect memorization rather than capability. This matters for procurement, because a model selected on published benchmarks may perform very differently on an organization's actual work. The defence is evaluating on private material — a set of questions drawn from the organization's own domain with answers its experts approved — which cannot have leaked into any training set.

Because Centralpoint retains the assembly of each execution, an organization's own evaluation set can be run across candidate models with full context held constant, so what is compared is the model rather than the scaffolding. Runtime selection makes that comparison practical: switching models for an evaluation run requires no re-indexing.


{0}