What is DPO?

DPO simplifies preference training by optimizing directly against pairs of preferred and rejected responses, avoiding the reward model and reinforcement loop that earlier alignment approaches required. It is cheaper, more stable, and has become common for tuning models toward particular response styles. The governance question it raises is worth stating plainly: preference alignment encodes judgements about what a good answer looks like, and those judgements are baked into weights that cannot be inspected, versioned or audited afterwards. Centralpoint's position is the opposite arrangement — organizational judgement lives in skills and prompts held as records in the organization's own SQL environment, version-controlled, audience-scoped and owned by named people. A rule can be read, disputed, amended and rolled back. An alignment baked into weights can only be replaced by retraining.