• Decrease Text SizeIncrease Text Size

Confidence Calibration

Language models express uncertainty poorly. Fluency is constant whether the answer is well supported or invented, which means the reader's most natural signal of reliability carries no information. Calibration measures the gap: across a set of answers, how often statements delivered with apparent certainty are correct. Poorly calibrated systems are dangerous in proportion to how articulate they are, because confident prose invites acceptance. The remedies are structural rather than linguistic — grounding answers in retrieved sources so certainty can be checked, and declining where support is absent rather than producing a plausible completion.

Governance-tier skills in Centralpoint can require that answers stay within retrieved material and escalate rather than extrapolate, so an unsupported question routes to a person instead of producing a confident guess. Because the Interaction Log retains what was retrieved for each execution, a confident answer with thin support is identifiable after the fact rather than indistinguishable from a well-grounded one.


{0}