• Decrease Text SizeIncrease Text Size

What is data mining in the Centralpoint context?

The term is used here in its original sense rather than the statistical one: examining accumulated content to discover what is actually there. Most organizations cannot state what their repositories hold. Content has built up across decades, classification is inconsistent where it exists, duplication is extensive, and the people who understood particular collections have moved on. Indexing that estate without examining it first produces a retrieval surface whose contents are unknown — which is the opposite of governance regardless of what controls sit above it. Centralpoint's ingestion characterizes as it goes: Data Transfer reads from the source systems, Data Cleaner evaluates content against the organization's own dictionary, and taxonomy assignment places records in a structure the business recognizes. Duplication, orphaned material, unclassified sensitive content and stale records surface at that point, when they are cheap to address rather than after they have been embedded.