Tokens In <> Tokens Out API¶
/inference/v1/generate takes prompt token IDs and returns generated token IDs. It is the generation step of the render → generate → derender pipeline used in disaggregated serving (see the Renderer APIs and Derenderer APIs).
The endpoint is registered on vllm serve with --enable-scale-out, and always with --tokens-only. --tokens-only also skips loading the tokenizer, so the server can't turn token IDs back into text.
Output modes¶
The request field output_mode selects how much of the postprocessing the server does for the caller.
output_mode |
Response adds | Server needs |
|---|---|---|
tokens (default) |
nothing, token IDs only | nothing |
text |
text on every choice and decoded tokens with bytes in logprobs |
a tokenizer |
text returns text and token IDs in a single call, without a separate derender hop per chunk. This suits latency sensitive streaming, where the derender hop lands on every output token.
Every response and every stream chunk echoes output_mode. Servers that predate the field ignore it and return token IDs only with a 200, so check that the output_mode in the response matches the one you sent. Older servers don't return it.
import httpx
response = httpx.post(
"http://localhost:8000/inference/v1/generate",
json={
"token_ids": [151644, 872, 198],
"sampling_params": {"max_tokens": 32},
"output_mode": "text",
},
).json()
assert response["output_mode"] == "text"
print(response["choices"][0]["text"])
print(response["choices"][0]["token_ids"])
Text¶
textcomes from the request's own detokenizer, soskip_special_tokens,spaces_between_special_tokens,stopandinclude_stop_str_in_outputbehave as they do on/v1/completions.- A matched stop string is cut from
textunlessinclude_stop_str_in_outputis set.token_idskeeps every generated token, so decodingtoken_idsyourself brings the stop string back. /derenderonly acceptsoutput_mode: "tokens"responses and returns a 400 for text responses which are already detokenized.- With
stream: true, each choice'stextis the delta since the previous chunk. A chunk is sent whenever the engine output carries new text or afinish_reason, even with no new token IDs. That covers text held back for stop string matching and the final output after an abort.
Logprobs¶
With output_mode: "tokens", logprob entries carry token_id:N placeholders. With output_mode: "text", they carry the decoded token strings and their UTF-8 bytes for the sampled token and every entry in top_logprobs. A server started with --return-tokens-as-token-ids always returns the placeholders, as /v1/completions does.
Errors¶
The server returns a 400 for:
output_mode: "text"on a server without a tokenizer (--tokens-onlyor--skip-tokenizer-init).output_mode: "text"withsampling_params.detokenize: false.- An unsupported
output_modevalue.
For using output_mode with separate prefill and decode pools, see Disaggregated Prefilling.
Aborting requests¶
POST /inference/v1/abort_requests aborts in-flight requests. It is registered wherever /inference/v1/generate is and requires the API key when --api-key is set, like /inference/v1/generate.
curl -X POST http://localhost:8000/inference/v1/abort_requests \
-H "Authorization: Bearer $VLLM_API_KEY" \
-H "Content-Type: application/json" \
-d '{"request_ids": ["generate-tokens-42"]}'
The request ID is the request_id the server returned: the X-Request-Id header or body request_id you sent, prefixed with generate-tokens-. A missing request_ids returns a 400. The response is empty and the abort finishes in the background.
With --tokens-only, the same handler is also served at POST /abort_requests for existing deployments. That path doesn't require the API key, even when --api-key is set. See Security.