Sensitive Data Discovery for AI
This has moved from an emerging practice to a mainstream control at the upper end of the market, which is a useful signal for organizations still treating it as optional. The reasoning is straightforward: general-purpose models cannot be relied upon to recognize what a particular organization considers sensitive, so the recognition has to happen before content reaches them. Discovery answers what exists, classification answers what category it belongs to, and the pair determine what may enter a retrieval surface at all.
In S&P Global's Voice of the Enterprise: Data & Analytics, Data Governance & Privacy 2026 survey, 37.2% of organizations reported using discovery and classification of sensitive data to safeguard it from generative AI and large language models, rising above 47% among businesses with more than $1 billion in revenue. Centralpoint performs both during ingestion, using the organization's own dictionary rather than a generic pattern set.