What is Speculative Decoding?
Speculative decoding is an LLM inference acceleration technique introduced by Google researchers in 2022 and refined by DeepMind in 2023 that uses a small fast "draft model" to propose multiple candidate tokens which are then verified in parallel by the large "target model". The platform strengthens en