• Decrease Text SizeIncrease Text Size

How does Centralpoint handle Draft Model?

Draft-model acceleration in Centralpoint: Centralpoint routes to inference endpoints using draft-model speculation while metering tokens at the target-model rate consistently. The model-agnostic platform supports any backend — vLLM, TensorRT-LLM, hosted APIs — keeps prompts local, and deploys chatbots through one line of JavaScript on any portal. The platform str