Concepts¶
The problem¶
An agent with hundreds of tools cannot see all of them: their definitions crowd out the task, cost tokens on every turn, and make the model pick worse. Tool retrieval puts a search in front: the agent describes what it wants to do, and only the tools that fit are loaded.
Backbone and heads¶
toolrank ranks with an embedding model, the backbone, served by vLLM (or any OpenAI-compatible
embeddings endpoint). A request and a tool's text become vectors, and the tools whose vectors are
closest to the request's come first. The default backbone is Qwen3-Embedding-8B with a LoRA trained
on ToolRet's training pairs (yasinyaman/toolrank-emb-8b, served as toolrank-emb-v0.2;
model card): the best single stage we measured. Its gain is in requests
that need several tools (on generated two- and three-tool tasks over an API catalogue of our own,
NDCG@10 71 → 83 and 55 → 75); a request for one tool it finds about as well as the base model.
The base Qwen3-Embedding-8B can be served instead (as qwen3-emb). On it sit two small heads,
one for requests and one for tools (29.9M parameters together). Each is x + MLP(x): it starts as
the identity, so it can only move the backbone's vectors where training showed it helps. The
packaged heads (v0.1) were trained on ToolRet's training pairs; they run in numpy, without torch,
and ship as a 60 MB .npz (model card). toolrank applies them only on the
base model they were trained on: on the LoRA backbone they cost 1–2 points, and heads trained on
top of it learn nothing. toolrank finetune and toolrank learn train your own.
The backbone reads an instruction in front of each request
(Instruct: Given an agent's request for a tool, retrieve the MCP tool that fulfills it), and each
tool as JSON with its server first ({"server", "name", "description", "inputSchema"}): the server
name alone is worth several points on MCP benchmarks.
Adaptive K¶
A fixed top-k either misses tools that a multi-step request needs or wastes context on tools it
does not. By default toolrank returns the tools within 0.2 cosine of the best one, at most 10:
a precise request gets one or two tools, a broad one more. --k N asks for a fixed number instead.
The ingest directory¶
toolrank ingest writes one directory, which search and serve read:
| Path | What it holds |
|---|---|
tools.jsonl |
the catalogue: one tool per line, id <source>/<name> |
sources.json |
each source (MCP server or OpenAPI spec) and when it was listed; never secrets |
cache/ |
embeddings, keyed by endpoint, model, truncation and text |
index/ |
the tool vectors after the heads, rebuilt only for new or changed tools |
usage/ |
the usage log (serve only) |
Re-running an ingest replaces exactly the sources it lists, keeps a source whose listing failed,
and embeds only new or changed tools. A running server notices the new tools.jsonl and swaps its
index without a restart.
Because the embedding cache is keyed by the endpoint's URL as written, ingest --emb-url must use
the same URL string as search and serve for them to find the vectors (TOOLRANK_EMB_URL sets it
for all three).
Search, rank, call¶
- Search: a request against the whole catalogue (
toolrank search, MCPsearch_tools,POST /v1/search). Each search gets asearch_id. - Rank: tools the caller brings, scored for a request, not the catalogue (
POST /v1/rank, which the LiteLLM filter uses). - Call: runs a catalogue tool (MCP
call_tool,POST /v1/call) on its MCP server or as an HTTP request to its OpenAPI operation. OpenAPI calls are GET and HEAD only unlessserve --allow-write.
On the agent APIs a tool goes by its api name: github/issues/create becomes
github__issues__create, and ids that don't fit the APIs' character rules get a short hash.
While the index builds¶
At the first start, tools missing from the cache are embedded in the background (1,862 tools took 154 s with a remote backbone). Meanwhile searches get BM25 keyword matches, marked as such, and calls work; the server reports healthy once the semantic index is up.
The usage log¶
serve appends every search and call to usage/usage-YYYY-MM-DD.jsonl: what was retrieved, what
the agent then called and how it went, with each call tied to the search that found its tool.
Requests and arguments are stored as keyed digests, not text, unless --log-text. The log is what
the next heads can learn from: tools retrieved but not called, and tools called successfully.