vllm.v1.metrics.cache_hit_source
¶
Classes:
-
CacheHitSource–Where cached prompt tokens' KV came from.
CacheHitSource
¶
Where cached prompt tokens' KV came from.
Members are listed in the order blocks are offloaded: accelerator, then
host memory, then secondary tiers. Per-source token counts are plain
dict[CacheHitSource, int] mappings holding only non-zero entries.
Methods:
-
outermost–The tier in
sourcesfarthest from the accelerator.
Source code in vllm/v1/metrics/cache_hit_source.py
outermost(sources)
classmethod
¶
The tier in sources farthest from the accelerator.