Skip to content

Tutorials

Step-by-step tutorials to guide you through complete workflows, from data preparation to serving trained models in production.

Train a Speculator

The main end-to-end walkthrough: prepare data, generate hidden states, train, and serve. Covers Eagle-3, P-EAGLE, DFlash, DSpark, and MTP, in online, offline, or hybrid mode -- pick your algorithm and mode at the top of the page.

Response Regeneration

Regenerate dataset responses using your target model for improved drafter alignment. Recommended before training.

Evaluating Model Performance

Benchmark and evaluate your trained speculator models.

Serve in vLLM

Deploy your trained speculator models in vLLM for production inference.