Deployment Overview#
The vLLM Production Stack provides three primary deployment options to suit different use cases and infrastructure requirements. Each option offers unique capabilities and advantages:
Deployment Options#
- Helm Chart Deployment
The standard deployment method using Helm charts for Kubernetes. This provides a streamlined way to deploy vLLM with configurable parameters for models, resources, and routing logic.
- Custom Resource Definitions (CRD)
Deploy using Kubernetes CRDs for more advanced configurations and operator-based management. This option provides greater flexibility and integration with Kubernetes-native workflows.
- Gateway Inference Extension
Advanced deployment option that uses agentgateway, the Gateway API Inference Extension, and the llm-d Router to route requests across pools of vLLM model servers.
Choose the deployment option that best fits your infrastructure requirements and use case complexity.