Serve them to agents¶
toolrank serve --data tools/ --config toolrank.json # http://127.0.0.1:8765/mcp and /v1
toolrank serve --data tools/ --config toolrank.json --stdio # for desktop clients
Needs the [mcp] extra. The agent sees two tools instead of hundreds:
search_tools(query, k?)returns the matching tools with their input schemas (complete for the first three, shortened for the rest) and asearch_id;call_tool(name, arguments, search_id?)forwards the call to the tool's MCP server or OpenAPI operation. A call that fails validation returns the tool's full schema, so the agent can retry.
The config file¶
--config tells serve how to reach the sources when a tool is called: an MCP client file
(mcpServers) plus an optional openapi section with each API's base URL and headers. ${VAR} is
read from the environment.
{
"mcpServers": {"time": {"command": "uvx", "args": ["mcp-server-time"]}},
"openapi": {"stripe": {"base_url": "https://api.stripe.com",
"headers": {"Authorization": "Bearer ${STRIPE_KEY}"}}}
}
An API's headers go only to its configured base_url, and redirects are not followed. OpenAPI
calls are GET and HEAD only unless you pass --allow-write.
Keys and hosts¶
On 127.0.0.1 (the default) no key is needed. Any other address needs one:
--api-keyorTOOLRANK_API_KEY: one bearer token for everyone;--api-keys keys.json: one named token per client or team ({"ci-agent": "${CI_AGENT_KEY}"}); the name goes into the usage log as the tenant.
Tenants¶
A named key can also be limited to some sources and carry its own credentials:
{
"ops": "${OPS_KEY}",
"team-a": {
"key": "${TEAM_A_KEY}",
"sources": ["github", "time"],
"headers": {"github": {"Authorization": "Bearer ${TEAM_A_GITHUB_TOKEN}"}},
"env": {"time": {"TZ": "Europe/Istanbul"}}
}
}
sources: the key sees and calls only these sources' tools. Searches,/v1/tools,/v1/rankby id and calls leave the others out, and a tool outside them is answered like a tool that does not exist. Withoutsourcesthe key reaches everything.headers: sent with this key's calls to that source on top of the config's: an OpenAPI source (to its configuredbase_urlonly) or a streamable HTTP MCP server.env: added to a stdio MCP server's environment.
A source for which a key has headers or env gets a connection (for a stdio server, a process)
of that key's own, so one team's token never carries another team's call. The server refuses to
start when a key has credentials for a source it cannot send them to. Each key also has its own
co-use table and its own heads (DATA/heads/tenants/<name>/); the catalogue, the index and the
embedding cache are shared. /v1/metrics is server-wide, so a key limited by sources cannot read
it.
The server checks the Host header, so a web page cannot reach it through a rebound domain. It
answers to its bind address and, when bound to 0.0.0.0, to localhost. Add the names it is
reached by, such as a compose service or a proxy's host, with --allowed-host NAME or
TOOLRANK_ALLOWED_HOSTS=a,b. A name without a port also matches requests with no port, as a
reverse proxy on 80 or 443 sends them.
The server speaks plain HTTP. To reach it from another machine, put a reverse proxy that terminates TLS in front of it: without one the bearer token, the requests and the tools' results cross the network in the clear.
REST¶
The same port serves REST for platforms that search and call tools themselves:
| Endpoint | What it does |
|---|---|
POST /v1/search |
a search over the catalogue |
POST /v1/rank |
scores up to 200 tools you pass in, or catalogue ids |
POST /v1/call |
runs a catalogue tool |
GET /v1/tools, GET /v1/tools/{id} |
the catalogue (?full=true with schemas) |
GET /v1/metrics |
Prometheus metrics |
GET /openapi.json, GET /healthz |
the API description; readiness |
See the REST reference.
Many servers, several tools, nothing that fits¶
Four options for catalogues where plain ranking leaves something on the table. All are off by default; the numbers are from the benchmarks (Benchmarks explains the sets).
--server-weight 0.2: a vote for the right server. Each server is embedded once as a summary
(its name and tool names), and a tool's score becomes its own cosine plus 0.2 of the request's
cosine with its server. On catalogues of many servers this lifts the first hit: MCP-Zero (293
servers) top-1 79.9 → 81.0, LiveMCPBench NDCG@10 54.0 → 55.1, with fewer tools returned at a higher
recall. Choosing servers first and searching only those loses points everywhere (the right server
ranks first only 70–85% of the time), so the vote is soft. On a catalogue whose "servers" are a few
huge groups it does nothing useful (ToolRet's three categories: −0.1 NDCG@10, −0.8 averaged by
category), which is why it is not the default.
--co-use 2: tools that are called together. The usage log knows which tools agents called
after the same request. With --co-use N, a result gains up to N tools that were called along with
one of its tools in at least two requests and at least half of that tool's requests; they come
last, marked used_with. On a simulated log this changed one list in twenty and raised the share
of requests that got every tool they needed by 0.6 points for 0.05 more tools per list; making
the ranked list longer buys a seventh of that per tool. The table is rebuilt from the last 30 daily
log files every five minutes, counts all API keys together, and is not applied to a request that
names its own k.
--rerank cross or --rerank jev: a second stage. The first stage scores the request and each
tool apart; a second stage reads the request together with each of the top 20 tools' full
documentation (cut to 3,000 characters) and reorders them. How many tools a search returns still
comes from the first stage's cosines (adaptive K); the order comes from the second stage. Measured
over the released heads, it is the largest single gain on catalogues the models never saw:
LiveMCPBench NDCG@10 54.0 → 62.7 with the local reranker, 64.0 with Jev; MCP-Zero top-1 79.9 →
91.3 / 92.3; ToolRet 54.0 → 58.1 / 57.7 (the comparison).
--rerank cross --rerank-emb-url http://HOST:PORT/v1: Qwen3-Reranker-8B behind vLLM's score API (about 16 GB more GPU memory, 0.3–0.6 s more per search).deploy/spark/compose.yaml'srerankprofile shows the vLLM flags.--rerank jev: TypeSafe AI's hosted Jev, withTYPESAFE_API_KEYset (about 0.3 s per search). The request text and the top tools' text are sent to TypeSafe; the server says so when it starts. Answers are cached inDATA/cache.
--rerank-depth, --rerank-tool-format and --rerank-max-chars change the setting; the defaults are
the one that measured best. Requests that name their own k are reranked too.
--cut-threshold T --cut-min 0: say so when nothing fits. By default a search returns at least
one tool. With a threshold and a minimum of zero, a request whose best score is below T gets an
empty list and a note telling the agent to answer without a tool or rephrase (add --cut-margin 0.2
to keep the usual cut above the threshold). Choose T on your own traffic: the scores of requests
that have a tool and of those that do not overlap, and their scale moves with the catalogue and the
way requests are written. At a T that turns away 1% of answerable requests, MCP-Zero catches a
quarter of the unanswerable ones; on LiveMCPBench no threshold is that cheap. The log keeps what a
turned-away search would have shown (results, with shown: 0), which is what to calibrate on.
Metrics¶
GET /v1/metrics answers in Prometheus' text format, behind the same bearer token as the rest of
/v1:
scrape_configs:
- job_name: toolrank
metrics_path: /v1/metrics
authorization: {credentials: "<TOOLRANK_API_KEY>"}
static_configs: [{targets: ["127.0.0.1:8765"]}]
| Metric | What it tells you |
|---|---|
toolrank_searches_total{via,mode,arm}, toolrank_search_duration_seconds |
traffic and ranking latency (the request's embedding included) |
toolrank_search_tools_returned, toolrank_search_empty_total, toolrank_search_co_use_added_total |
how many tools a search hands over |
toolrank_search_returned_tokens_total, toolrank_search_saved_tokens_total, toolrank_catalog_tokens |
the token estimate (below) |
toolrank_calls_total{kind,outcome,via}, toolrank_call_duration_seconds |
calls forwarded and how they ended |
toolrank_calls_linked_total{link}, toolrank_called_tool_rank |
whether calls can be tied to a search, and where the called tool stood in it |
toolrank_embedding_texts_total{kind,source}, toolrank_embedding_tokens_total |
embedding-cache hits (source="cache") against texts sent to the endpoint |
toolrank_catalog_tools, toolrank_catalog_sources, toolrank_index_ready, toolrank_heads{arm}, toolrank_build_info |
what is being served |
The token estimate answers "what did searching save over loading every tool?". A tool counts as
its name, description and input schema in JSON at four characters a token; a search returned the
tokens of the tools it handed over and saved the catalogue's total minus that. So
saved / (saved + returned) is the reduction: on a catalogue of 1,862 tools (910k tokens) a search
returns about 5k, a reduction above 99%. It is an estimate and a floor: schemas are counted in full
although later hits are shortened, and the first searches after a catalogue change claim nothing
while the new catalogue is being sized. The share of calls in the top five of their search is
toolrank_called_tool_rank_bucket{le="5"} / toolrank_called_tool_rank_count.
Labels never carry request text, tool arguments or API key names (arm says tenant for any
key's own heads), and the counters are server-wide: any valid key can read them. They count what
the usage log sees and keep counting under --no-usage-log; they start at zero with the process.
The usage log¶
Every search and call goes to DATA/usage/ (or --usage-log DIR), one JSON line per event in a
file per day, each call tied to the search that found its tool. It is what
toolrank learn trains the heads on.
The server also follows DATA/heads while it runs: current.npz there replaces the served heads,
candidate.npz takes --candidate-share of the requests, and tenants/<name>/ holds the same for one
API key. See Learn from the usage log.
What it holds, and what it does not:
- Requests and call arguments are keyed digests (HMAC-SHA256 under
DATA/usage/.key, a random 32-byte key made on first use, mode 0600): repeats are recognisable, guesses are not, and the log alone reconstructs nothing. The request's embedding-cache key is a digest too (emb_hmac), which is howlearnfinds its vector without its text. - Tool names, scores, outcomes, latencies and the server's own instruction are text; so are the
tools a search added by co-use (
added). An instruction sent with a request is a digest. The client (API key name, client app, remote address) is a digest; the session id and the tenant (the API key's name) are text. --log-textadds the request text and the error text of failed calls;--mask-piithen replaces e-mail addresses, phone, card and IBAN numbers in them with tags (a pattern, not an understanding).--no-usage-logturns the log off.
For data-protection purposes (GDPR, KVKK): without --log-text the log holds no personal data of
the requests themselves, only digests under a key you hold; the .key file and the embedding cache
(DATA/cache, which holds the requests' vectors) are the parts to protect and to delete with the
log. To forget a period, delete its daily files; to forget one API key's traffic, filter its
tenant lines out. An embedding vector can be matched to text only by embedding a guess and
comparing, which needs the same model and the cache.