Skip to content
vLLM
attn_res
Initializing search
GitHub
Home
User Guide
Developer Guide
Benchmarking
API Reference
CLI Reference
Community
vLLM
GitHub
Home
User Guide
User Guide
Getting Started
Getting Started
Quickstart
Installation
Installation
GPU
CPU
TPU
Examples
Examples
Applications
Applications
API Server
Chatbot
Rag
Basic
Basic
Offline Inference
Online Serving
Deployment
Deployment
Async LLM Streaming
Helm Charts
LLM Engine Example
Sagemaker-Entrypoint
Disaggregated
Disaggregated
Disaggregated Encoder
Disaggregated Serving
Ec Both Encoder
Disaggregated Prefill V1
Flexkv Connector
KV Load Failure Recovery Test
LMCache Examples
Mooncake Connector
Features
Features
Automatic Prefix Caching
Batch Invariance
Context Extension
Data Parallel
Kv Events
Logging Configuration
Custom Logits Processors
LoRA
Offline Inference with the OpenAI Batch file format
Pause Resume
Profiling
Prompt Embed
Reset Kv
Sharded State
Speculative Decoding
Structured reads on DiffusionGemma
Structured Outputs
Tensorize vLLM Model
Torchrun
Generate
Generate
Batched Chat Completions Online
Multimodal
Qwen 1M Offline
Trace Replay Offline
Observability
Observability
Monitoring Dashboards
Metrics
Setup OpenTelemetry POC
Prometheus and Grafana
Pooling
Pooling
Classify
Embed
Plugin
Reward
Score
Token Classify
Token Embed
Ray Serving
Ray Serving
Batch LLM Inference
Elastic Ep
Multi-Node-Serving
Ray Serve Deepseek
Run Cluster
Reasoning
Reasoning
OpenAI Chat Completion Tool Calls With Reasoning
OpenAI Chat Completion With Reasoning
OpenAI Chat Completion With Reasoning Streaming
OpenAI Responses Client
RL
RL
Rdt vLLM Serve
Rdt Weight Source
RLHF Async New APIs
RLHF Http IPC
RLHF Http NCCL
RLHF IPC Fsdp Ep
RLHF M2N
RLHF NCCL Fsdp Ep
RLHF Sharded Rdt Small Ep
RLHF Sparse NCCL
Routed Experts E2E
Skip Loading Weights In Engine Init
Scale Out
Scale Out
Init
Example Mm Serve
Token Generation Client
Speech To Text
Speech To Text
OpenAI
Realtime
Tool Calling
Tool Calling
Chat With Tools Offline
OpenAI Chat Completion Client With Tools
OpenAI Chat Completion Client With Tools Required
OpenAI Chat Completion Client With Tools Xlam
OpenAI Chat Completion Client With Tools Xlam Streaming
OpenAI Responses Client With Mcp Tools
OpenAI Responses Client With Tools
General
General
vLLM V1
Frequently Asked Questions
Production Metrics
Reproducibility
Security
Troubleshooting
Usage Stats Collection
Inference and Serving
Inference and Serving
Offline Inference
Online Serving
Online Serving
Derenderer APIs
Generative Scoring
OpenAI-Compatible Server
Renderer APIs
Speech to Text APIs
Trace Replay
Context Parallel Deployment
Data Parallel Deployment
Troubleshooting distributed deployments
Expert Parallel Deployment
Parallelism and Scaling
Integrations
Integrations
Claude Code
Codex
LangChain
LlamaIndex
Deployment
Deployment
Using Docker
Using Kubernetes
Using Nginx
Frameworks
Frameworks
Anyscale
AnythingLLM
AutoGen
BentoML
Cerebrium
Chatbox
Crusoe
Dify
dstack
Haystack
Helm
Hugging Face Inference Endpoints
LiteLLM
Lobe Chat
LWS
Modal
Nebius Serverless AI
Open WebUI
Retrieval-Augmented Generation
RunPod
SkyPilot
Streamlit
NVIDIA Triton
Integrations
Integrations
AIBrix
NVIDIA Dynamo
KAITO
KServe
Kthena
KubeAI
KubeRay
Llama Stack
llm-d
llmaz
Production stack
Training
Training
Async Reinforcement Learning
What is Layerwise (Re)loading?
Prompt Token ID Logprobs (Teacher Scoring)
Reinforcement Learning from Human Feedback
Sampling Mask (Distribution Replay)
Transformers Reinforcement Learning
Weight Checker
Weight Transfer
Weight Transfer
Base Classes and Custom Engines
IPC Engine
NCCL M2N Engine
NCCL Engine
Sharded RDT Engine
Configuration
Configuration
Conserving Memory
Engine Arguments
Environment Variables
Model Resolution
Optimization and Tuning
Server Arguments
TPU
Models
Models
Supported Models
Generative Models
Pooling Models
Pooling Models
Classification Usages
Embedding Usages
Reward Usages
Scoring Usages
Specific Model Examples
Token Classification Usages
Token Embedding Usages
Extensions
Extensions
Loading model weights with fastsafetensors
Loading Model Weights with InstantTensor
Loading models with Run:ai Model Streamer
Loading models with CoreWeave's Tensorizer
Hardware Supported Models
Hardware Supported Models
CPU - Intel® Xeon®
XPU - Intel® GPUs
TPU
Features
Features
Automatic Prefix Caching
Batch Invariance
Context Extension
Cross-Encoder Output Reuse
Custom Arguments
Custom Logits Processors
Disaggregated Encoder
Disaggregated Prefilling (experimental)
CPU EC Connector Usage Guide
Engram: conditional memory via n-gram lookups
IndexCache
Initialized engine snapshots
Interleaved Thinking
KV Offloading Usage Guide
LoRA Adapters
MooncakeConnector Usage Guide
MooncakeStoreConnector Usage Guide
MoRIIOConnector Usage Guide
Multimodal Inputs
NixlConnector Compatibility Matrix
NixlConnector Usage Guide
Per-Request Metrics
Preload
Prompt Embedding Inputs
Reasoning Outputs
Sleep Mode
Structured Outputs
Tool Calling
Text watermarking
Quantization
Quantization
AutoAWQ
b12x Linear and MoE Backends
BitsAndBytes
FP8 ViT Encoder Attention
GGUF
GPTQModel
Intel Quantization Support
NVIDIA Model Optimizer
Online Quantization
Quantized KV Cache
AMD Quark
TorchAO
LLM Compressor
LLM Compressor
FP8 W8A8
INT4 W4A16
INT8 W4A8
INT8 W8A8
Speculative Decoding
Speculative Decoding
Per-Request Acceptance Metrics
Adaptive Verification
Draft Models
Dynamic Speculative Decoding
EAGLE Draft Models
Hidden State Extraction
LiLiCorr
MLP Draft Models
MTP (Multi-Token Prediction)
N-Gram Speculation
Parallel Draft Models
vLLM-Project/Speculators
Suffix Decoding
Developer Guide
Developer Guide
General
General
Deprecation Policy
Dockerfile
Editing Agent Instructions
Incremental Compilation Workflow
JIT Kernel Warmup
Labels
Profiling vLLM
Releasing vLLM
Vulnerability Management
Model Implementation
Model Implementation
Basic Model
Registering a Model
Unit Testing
Multi-Modal Support
Speech-to-Text (Transcription/Translation) Support
CI
CI
CI Failures
Nightly Builds of vLLM Wheels
Update PyTorch version on vLLM OSS CI/CD
Design Documents
Design Documents
Plugins
Plugins
Endpoint Plugins
IO Processor Plugins
LoRA Resolver Plugins
Plugin System
Architecture Overview
Attention Backend Feature Support
CUDA Graphs
Vision Encoder (ViT) CUDA Graphs
CustomOp
Dual Batch Overlap
How to debug the vLLM-torch.compile integration
Fused MoE Modular Kernel
Fusion torch.compile passes
HiSparse local KV offload architecture
Integration with Hugging Face
Hybrid KV Cache Manager
Logits Processors
Metrics
Multi-Modal Data Processing
Model Runner V2 Design Document
Fused MoE Kernel Features
Python Multiprocessing
NIXL KV Cache Lease Renewal
NIXL push-mode KV transfer
Optimization Levels
Paged Attention
Automatic Prefix Caching
torch.compile integration
torch.compile with Multimodal Encoders
vLLM IR: Functional Intermediate Representation
Benchmarking
Benchmarking
Benchmark CLI
Parameter Sweeps
Performance Dashboard
API Reference
API Reference
vllm
vllm
collect_env
connections
env_override
envs
exceptions
forward_context
logger
logits_process
logprobs
model_inspection
outputs
pooling_params
sampling_params
scalar_type
scripts
sequence
tasks
version
assets
benchmarks
compilation
config
cute_utils
device_allocator
distributed
engine
entrypoints
inputs
ir
kernels
logging_utils
lora
model_executor
models
models
common
deepseek_v4
deepseek_v32
deepseek_v41
dots3_note
glm5next
hy_v4
inkling
kimi_k3
kimi_k3
amd
amd
kda
kda_metadata
latent_moe_runner
linear
mla
model
mtp
ops
ops
attn_res
kda_chunk
kda_decode
kda_prefill
third_party
common
nvidia
minimax_m3
qwen4_exp
multimodal
parser
platforms
plugins
profiler
ray
reasoning
renderers
snapshot
tilelang_utils
tokenizers
tool_parsers
tracing
transformers_utils
triton_utils
usage
utils
v1
CLI Reference
CLI Reference
vllm
vllm
chat
complete
preload
run-batch
serve
bench
bench
latency
mm-processor
serve
startup
throughput
sweep
sweep
plot
plot_pareto
serve
serve_workload
startup
launch
launch
render
Community
Community
Contact Us
Meetups
Sponsors
Governance
Governance
Collaboration Policy
Committers
Governance Process
Blog
Forum
Slack
Home
API Reference
vllm
models
kimi_k3
amd
ops
vllm.models.kimi_k3.amd.ops.attn_res
¶
Back to top