Skip to content

Quick start

toolrank needs an embedding model to rank tools: its backbone, Qwen3-Embedding-8B trained further on tool retrieval, served by vLLM on a GPU with at least 16 GB of memory (or any OpenAI-compatible /v1/embeddings endpoint serving it). The Docker path starts both; the pip path uses an endpoint you already run.

With Docker (one GPU host)

git clone https://github.com/yasinyaman/toolrank && cd toolrank/deploy/docker
cp .env.example .env              # set TOOLRANK_API_KEY

List the MCP servers you want to put behind toolrank in toolrank.json (the same format as an MCP client's config file). It starts with one:

{
  "mcpServers": {
    "time": {"command": "uvx", "args": ["mcp-server-time"]}
  }
}

Index them, then start the stack:

docker compose run --rm toolrank ingest mcp --config /config/toolrank.json --out /data
docker compose up -d

The first start downloads the model (16 GB). toolrank then listens on 127.0.0.1:8765: MCP at /mcp, REST at /v1, both behind the key from .env.

curl -s -H "Authorization: Bearer $TOOLRANK_API_KEY" \
  -d '{"query": "what time is it in Tokyo?"}' http://127.0.0.1:8765/v1/search

Docker covers both images, OpenAPI specs, GPU memory and running behind a proxy.

With pip

pip install "toolrank[mcp]"           # add ,openapi for YAML specs; ,stem for a stemmed BM25 fallback

Serve the backbone on a GPU host, as toolrank-emb-v0.2 on port 8091 (the default toolrank uses; --quantization fp8 halves its memory, served then as toolrank-emb-v0.2-fp8):

vllm serve yasinyaman/toolrank-emb-8b --revision v0.2 --served-model-name toolrank-emb-v0.2 \
  --runner pooling --max-model-len 8192 --port 8091

With the base Qwen/Qwen3-Embedding-8B served as qwen3-emb instead, toolrank heads pull fetches the heads that go with it (60 MB) and --emb-model qwen3-emb selects it.

On another machine, point toolrank at it with --emb-url http://HOST:8091/v1 or TOOLRANK_EMB_URL; an endpoint that wants a key gets TOOLRANK_EMB_API_KEY (toolrank sends OPENAI_API_KEY only to api.openai.com). Then index your tools and search them:

toolrank ingest mcp --server time="uvx mcp-server-time" --out tools/
toolrank ingest openapi https://petstore3.swagger.io/api/v3/openapi.json --name petstore --out tools/
toolrank search --data tools/ "list the pets tagged dog"

And serve them to agents:

toolrank serve --data tools/ --config toolrank.json     # MCP at http://127.0.0.1:8765/mcp

--config tells serve how to reach the MCP servers when an agent calls one of their tools; the catalogue itself comes from ingest. See Serve them to agents.

Connect an agent

Any MCP client can use toolrank as one server:

claude mcp add --transport http toolrank http://127.0.0.1:8765/mcp \
  --header "Authorization: Bearer $TOOLRANK_API_KEY"

Desktop clients start servers themselves, so serve runs with --stdio; every path must be absolute (Claude Desktop starts servers in /):

{
  "mcpServers": {
    "toolrank": {
      "command": "/path/to/venv/bin/toolrank",
      "args": ["serve", "--stdio", "--data", "/path/to/tools",
               "--config", "/path/to/toolrank.json", "--emb-url", "http://127.0.0.1:8091/v1"]
    }
  }
}

Platforms that run tools themselves call POST /v1/search and then POST /v1/call (or run the tool on their side); see the REST API.

The agent sees two tools, search_tools and call_tool, whatever the size of the catalogue.