What is AWQ?

AWQ reduces the precision of a model's weights while identifying and protecting the small proportion that matter most to output quality, which allows substantial compression without the degradation cruder quantization produces. The result is a model that runs on less hardware and responds faster, which is what makes local inference practical on infrastructure an ordinary organization already owns. That is where the governance relevance lies. Quantization is what turns on-premises inference from a theoretical option into a deployable one, and on-premises inference is the only arrangement in which prompts containing regulated content never leave the environment. Centralpoint supports embedded local models alongside cloud providers, with selection made per request — so an organization can route sensitive categories to local inference and everything else to whichever provider is most economical, without maintaining two separate deployments.