Automated Classification at Scale
Manual classification works for tens of thousands of records and fails at millions, which is the scale at which conversational and historical content actually arrives. The alternatives are to classify nothing, to classify only new material, or to automate. Automation raises the question of what is doing the classifying: a model-based approach introduces probabilistic judgement and inference cost into a control that ought to be deterministic, while a rule-based approach requires the organization to articulate its vocabulary explicitly.
Centralpoint uses the rule-based path. Data Cleaner applies the organization's own dictionary — its terms, statutes and categories — during ingestion, so classification is deterministic and repeatable, no model is paid to judge sensitivity, and no risk is inherited from a model judging it wrongly. What was applied during any past period is establishable, because the dictionary carries version history.