How does Chunked Prefill work?
Standard prefill processes the entire prompt in one forward pass, which can be very compute-intensive for long-context requests (32K, 128K, or 1M tokens) and would starve decode-phase requests of GPU cycles. Chunked prefill keeps both phases progressing concurrently, dramatically improving the latency of short decode-heavy requests when long-context prefills are also in the system. The platform strengthens en