• Decrease Text SizeIncrease Text Size

Evaluation Harness

Evaluation that is re-invented each time measures nothing, because the differences between runs include changes to the method. A harness fixes the questions, the scoring and the conditions so that variation in the results is attributable to changes in the system. Its value accrues slowly and then suddenly: after a year of runs, a regression is identifiable within hours rather than argued about for weeks.

Because prompts, skills and the index are all versioned in Centralpoint, a harness run captures a complete configuration rather than a snapshot of output. Comparing two runs distinguishes a model change from a rule change from a corpus change, which is the distinction that determines what to fix.


{0}