Skip to content

Kubernetes

The chart in deploy/helm/toolrank runs toolrank serve with its embedding backbone. It is not published to a chart repository yet: install it from a checkout.

helm install toolrank deploy/helm/toolrank -n toolrank --create-namespace \
  --set auth.apiKey="$TOOLRANK_API_KEY"
kubectl -n toolrank port-forward svc/toolrank 8765:8765

The backbone

embedding.mode picks where the backbone runs:

Mode What runs Needs
vllm (default) a vLLM pod next to toolrank, weights on their own volume a GPU node (nvidia.com/gpu)
external nothing: embedding.url and embedding.model name your endpoint an OpenAI-compatible /v1/embeddings
bundled one pod with the toolrank-vllm image a GPU node, and the image built from deploy/docker/Dockerfile.vllm and pushed to your registry (embedding.bundled.image)

embedding.backbone names the weights: toolrank's LoRA-trained Qwen3-Embedding-8B by default (yasinyaman/toolrank-emb-8b at v0.2, served as toolrank-emb-v0.2), or {repo: Qwen/Qwen3-Embedding-8B, revision: "", name: qwen3-emb} for the base model with the packaged heads. embedding.profile sets the precision: fp8 (the default, quantized at load, about 8 GB, served as <name>-fp8) or bf16 (about 16 GB, <name>). In our runs FP8 stays within a point of bf16. The embedding cache is keyed by the served name, so switching profiles re-embeds the catalogue once. The first start of a vLLM pod downloads 16 GB; the probes allow 30 minutes for it. On k3s, and on other clusters where the NVIDIA runtime is not the default, set embedding.runtimeClassName=nvidia (and install the NVIDIA device plugin, which advertises nvidia.com/gpu).

Your tools

config is the toolrank.json of Serve them to agents: MCP servers and OpenAPI base URLs and headers. An init container runs toolrank ingest mcp on every start, so helm upgrade with a new server rolls the pod and the server appears in the catalogue. Helm merges your values into the chart's, so drop the example server explicitly (mcpServers: {time: null}). Credentials go in a Secret listed under envFrom and are referred to as ${VAR} in the config. OpenAPI specs are ingested with ingest.args, or with kubectl exec into the pod: the server re-reads the catalogue without a restart.

Keys and tenants

At least one key is required, since the server listens on the pod network: auth.apiKey (or auth.apiKeySecret) for one token, auth.apiKeys (or auth.apiKeysSecret, a Secret holding the file) for tenants, each with its own sources and credentials. The Service's names (toolrank, toolrank.<namespace>.svc, ...) are allowed as Host; add an ingress host to serve.allowedHosts. Terminate TLS at the ingress.

What it creates

One toolrank Deployment (one replica, Recreate: the index snapshot, the usage log and the learned heads live on one ReadWriteOnce volume), its Service, a ConfigMap, a Secret when the keys are given inline, the data volume (and a weights volume for vllm and bundled). The volumes are kept on helm uninstall: they hold the embedding cache and the usage log. serve.args passes further flags (--server-weight 0.2, --co-use 2, --allow-write). Prometheus scrapes /v1/metrics with a key that reaches every source (Metrics).

Trying it

scripts/helm_smoke.py installs the chart on any cluster without a GPU: a fake embedding server in the cluster, external mode, then a search, a call to a stdio MCP server inside the pod, a refused token, the metrics, a key limited to one source, an upgrade that adds a server, and an uninstall that keeps the data volume. On a laptop, kind gives a cluster in Docker:

kind create cluster --name toolrank
uv run python scripts/helm_smoke.py                    # the published image
docker build -f deploy/docker/Dockerfile -t toolrank:dev . && kind load docker-image toolrank:dev --name toolrank
uv run python scripts/helm_smoke.py --image toolrank:dev --tenants
kind delete cluster --name toolrank

The GPU modes have been checked by rendering and by the API server's validation, not yet on a GPU cluster.