Running vLLM on TPU#
This page shows how to run vLLM inference on a Cloud TPU with Kinetic. You add one dependency file, forward one environment variable, and run the example script. Read this page if you want to generate text with a large language model, such as Llama 3.1, on a TPU slice.
Before you start#
You need these things:
A Kinetic cluster and an active profile.
kinetic initcreates both. See Getting Started.A node pool that matches the accelerator of the example. The example uses
accelerator="tpu-v5litepod", a 4-chip TPU v5e slice on one host. Runkinetic pool list. If the list has no matching pool, add one:kinetic pool add --accelerator tpu-v5litepod-4
A Hugging Face token if the model is gated. The example uses
meta-llama/Llama-3.1-8B, which is a gated model. Request access on the model page. Then create a token in your Hugging Face account settings.
Step 1: Add vllm-tpu to a dependency file#
Save a file with the name requirements.txt in the same directory as
your script:
vllm-tpu
When you call the decorated function, Kinetic looks for a dependency
file. Kinetic starts the search in the directory of the script and
continues in the parent directories. A requirements.txt next to the
script is therefore the first file that Kinetic finds. See
Dependencies for the search rules.
Kinetic then builds a container image that contains vllm-tpu and
caches the image. The first run waits for that build. Later runs with
an unchanged file reuse the image.
Note
Do not add jax, jaxlib, or libtpu to the file. The image that
Kinetic builds already contains JAX with the TPU runtime. If the file
has such a line, Kinetic removes the line and logs a warning. See
JAX and accelerator runtimes.
Step 2: Review the environment variables#
The example uses two kinds of environment variable. The vLLM variables are constants, and the script sets them inside the function. The Hugging Face token is a secret, and the script forwards it from your shell. You do not set anything in this step. Step 4 sets the token.
The vLLM variables#
The function sets three variables with os.environ before the
from vllm import LLM, SamplingParams line. The values therefore exist
in the pod process when vLLM loads.
Variable |
Value in the example |
Purpose |
|---|---|---|
|
|
Selects the TPU backend of vLLM. |
|
|
Lists the JAX backends. Kinetic sets |
|
|
Selects the vLLM engine version. |
You do not set these variables in your shell, and you do not forward
them with capture_env_vars. The code carries them.
The Hugging Face token#
One value comes from your shell: HF_TOKEN. The decorator lists the
name in capture_env_vars=["HF_TOKEN"]. When you call the function,
Kinetic reads HF_TOKEN from your environment and copies the value into
the job payload. The pod applies the value before the pod calls your
function, so vLLM can download a gated model. If your shell has no
HF_TOKEN, Kinetic captures nothing and logs no error. The pod then
has no token, and the download of a gated model fails. See
Forward Environment Variables.
Warning
The name HF_TOKEN contains TOKEN, so Kinetic logs a warning when it
captures the value. Kinetic stores the value in plaintext inside the job
payload in the jobs bucket, and every job pod in the cluster can read
that bucket. Kinetic deletes the payload when a blocking call or a
result() call collects a successful result with the default cleanup.
Otherwise the payload stays until the 30-day lifecycle rule of the
bucket deletes it. If you intend to forward the token, the warning needs
no action. See Secrets and
Security.
Step 3: Read the example#
Save this script as vllm_demo.py, next to the requirements.txt from
step 1:
# Installation
# before you begin please install VLLM in your remote environment
# pip install vllm-tpu
import os
import kinetic
# We use a TPU accelerator. Adjust as needed (e.g., 'tpu-v5e-1', 'tpu-v5litepod-4')
@kinetic.run(
accelerator="tpu-v5litepod",
# Capture Hugging Face token if using a gated model like Gemma
capture_env_vars=["HF_TOKEN"],
)
def run_vllm_inference():
# Imports must happen inside the decorated function because they need to run
# in the remote container where vllm is installed.
os.environ["VLLM_TARGET_DEVICE"] = "tpu"
os.environ["JAX_PLATFORMS"] = "tpu,cpu"
os.environ["VLLM_USE_V1"] = "0"
from vllm import LLM, SamplingParams
model_id = "meta-llama/Llama-3.1-8B"
print(f"Initializing vLLM with model: {model_id}")
# We use arguments matching the quickstart
llm = LLM(model=model_id, tensor_parallel_size=4, max_model_len=2048)
sampling_params = SamplingParams(temperature=0.8, top_p=0.95)
prompts = [
"Hello, my name is",
"The capital of France is",
"The president of the United States is",
]
print("Generating completions...")
outputs = llm.generate(prompts, sampling_params)
# Print the results
for output in outputs:
prompt = output.prompt
generated_text = output.outputs[0].text
print(f"\nPrompt: {prompt!r}")
print(f"Generated text: {generated_text!r}")
if __name__ == "__main__":
run_vllm_inference()
Four points in the script matter:
The accelerator.
accelerator="tpu-v5litepod"resolves to the default v5e slice: 4 chips on one host, topology2x2. The stringtpu-v5litepod-4names the same slice. See Accelerators.The parallelism.
tensor_parallel_size=4matches the 4 chips of the slice.The import. The
from vllm import ...line is inside the function. The pod runs that line, and the image on the pod contains vLLM. Your machine does not need vLLM.The result. The function prints the completions in the pod, and Kinetic streams the pod log to your terminal. If you want the completions in your local process, return them from the function.
Step 4: Run the script#
Set HF_TOKEN for the command and run the script:
HF_TOKEN=your-hf-token python vllm_demo.py
The active profile supplies the project, the zone, and the cluster. You do not pass them. Kinetic then does these things:
Package. Kinetic captures
HF_TOKEN, serializes the function, and archives the directory of the script.Build. Kinetic builds an image with
vllm-tpuon the first run, or reuses the cached image.Schedule. The cluster autoscaler starts a TPU v5e node in the matching node pool if no free node exists.
Run. The pod downloads the model weights from Hugging Face, loads the model on the TPU, and prints the completions. Kinetic streams the log lines to your terminal.
Collect. The call returns when the function ends. Kinetic deletes the job resources.
Note
The first run waits for the image build. Every run waits for a node
start if the node pool has scaled to zero. Every run also downloads the
model weights again, because the pod filesystem does not persist between
jobs. If you expect a long run, call run_vllm_inference.run_async()
instead of the blocking call, and collect the result later. See
Detached Jobs.
Change the model or the slice#
The model. Change
model_idto another Hugging Face model. If the model is gated, your token must have access to the model.The slice. Change
acceleratorto another single-host slice, for exampletpu-v5litepod-8(8 chips on one host). Settensor_parallel_sizeto the number of chips in the slice. Add a node pool that matches the new accelerator withkinetic pool add.Larger slices. A v5e slice with 16 or more chips spans more than one host. Kinetic runs such a job on the Pathways backend, one pod per host. See Distributed Training.