Golden Set

Without a fixed reference, quality assessment is anecdote. A golden set gives a stable measurement: the same questions, re-run after any change to a prompt, a skill, a model or the corpus, compared against answers a human approved. Its value depends on curation — the questions must cover the cases that matter, including the ones where the correct behaviour is refusal or escalation, not only the ones the system handles well. A golden set composed of easy questions measures nothing useful.

Because prompts and skills are versioned records and the Interaction Log retains what each execution assembled, a golden-set run produces a comparison across full context rather than across output text alone — which skills loaded, what was retrieved, what it cost. That distinguishes a change in the model from a change in the governance around it.


{0}