Changelog¶
Notable changes, newest first. The format follows Keep a Changelog, and versions follow semantic versioning; before 1.0, a minor version may change behaviour.
[Unreleased]¶
[0.2.0] - 2026-10-02¶
Learning from the usage log, a second stage, tenants, metrics, a Helm chart, and a LoRA-trained default backbone.
Added¶
toolrank eval --rerank jev: TypeSafe AI's Jev reorders the top--rerank-depthtools of any scorer with one Choice question per query;--scorer jevranks a corpus with Jev alone (chunked Choice questions, the chunk winners re-ranked once). The key is read fromTYPESAFE_API_KEY; answers are cached injev.sqlitenext to the embedding cache;scripts/jev_compare.shruns the comparison rows.toolrank eval --rerank dense|clm|cross: a second scorer over the first one's shortlist, with its own--rerank-*endpoint, formats, heads and text cut (--rerank-max-chars);--scorer crossand--rerank crossrun a cross-encoder behind vLLM's score API (Qwen3-Reranker and bge-reranker-v2-gemma prompt formats, scores cached inscores.sqlite); compose profilererankserves both on the GB10;scripts/clm_rerank.shandscripts/cross_rerank.shrun the rows.scripts/lora_train.py([lora]extra): LoRA fine-tuning of Qwen3-Embedding-8B on the fine-tuning data path, a parity check against the served vectors, the best adapter picked on a dev set and merged into weights that compose profileloraserves asqwen3-emb-lora.toolrank learn: heads trained from whattoolrank servelogged (calls that endedokas positives,tool_erroras weak positives, tools shown but not called as hard negatives), with the requests' vectors found through the log's key and never their text; the newest requests are the dev set, and the heads are written only when they beat the served ones there.- Heads that change while serving:
toolrank servefollowsDATA/heads(current.npzreplaces the served heads,candidate.npzanswers a sticky--candidate-shareof the requests,tenants/<name>/the same per API key), and the usage log records which arm answered.toolrank learnwrites its result as the candidate and takes--replay pairs.jsonlagainst forgetting;toolrank abcompares the arms on the log and promotes the candidate, sets it aside, or waits. toolrank serve --mask-pii: with--log-text, e-mail addresses, phone, card and IBAN numbers are replaced by tags before the request or error text is written.--server-weight W(eval,search,serve): each server is embedded as a summary of its tools and a tool's score gains W times the request's cosine with its server. Off by default; 0.2 lifts the first hit on catalogues of many servers (MCP-Zero top-1 79.9 → 81.0) and does nothing for a catalogue of a few huge groups.toolrank search/serve --co-use N: a result gains up to N tools that the usage log shows were called together with one of its tools; hits carryused_with, the log's search eventsadded.search_toolssays so when a threshold (--cut-threshold T --cut-min 0) turned every tool away, instead of returning an empty list without a word.scripts/routing_sweep.pyandscripts/couse_sweep.py: the measurements behind the two options and the no-tool gate.toolrank search/serve --rerank cross | jev: a second stage reorders the top 20 tools with the request (a local cross-encoder behind vLLM's score API, or TypeSafe AI's hosted Jev); the number of tools returned still comes from the first stage's cosines.toolrank data gen-queries: a selection set of your own. A chat model writes one request per sampled tool of any catalogue (four styles; tools sampled evenly over the sources), giving a benchmark-format directory foreval,finetune --devandlearn --devthat shares no query with a benchmark.scripts/lora_train.py --keep-allkeeps every evaluated adapter, so another dev set can pick again.- Tenants: an
--api-keysentry can be{key, sources, headers, env}.sourceslimits the key to those sources' tools (search, catalogue and calls);headersandenvare its own credentials for a source, sent only with its calls over a connection of its own. Co-use tables are per key. - A Helm chart (
deploy/helm/toolrank): toolrank with its backbone as a vLLM pod, an external endpoint or the bundled image, FP8 or bf16 profiles, MCP sources ingested by an init container, keys and tenants from values or Secrets;scripts/helm_smoke.pyinstalls it on a cluster without a GPU. GET /v1/metrics: Prometheus metrics of a running server (searches and calls with latency histograms, where called tools stood in their search, embedding-cache hits, and an estimate of the tokens searching saved over loading the whole catalogue).scripts/learn_sim.pyandscripts/learn_sim.sh: the learning loop measured on simulated traffic. A benchmark is served as a catalogue, a share of its queries is logged by an agent that calls the gold tools it is shown,toolrank learntrains on that log and the queries never served are the test; the learn guide's "What to expect" carries the numbers.
Changed¶
- The default backbone is Qwen3-Embedding-8B with a LoRA trained on ToolRet's training pairs
(
yasinyaman/toolrank-emb-8batv0.2, served astoolrank-emb-v0.2,-fp8in FP8): ToolRet NDCG@10 58.90 against 54.03 for the base model with heads, in one stage and without heads. On an API catalogue of our own it gains on tasks that need several tools and ties with the base model on requests for one tool (see its model card). The packaged heads are applied only on the base model they were trained on (qwen3-emb,qwen3-emb-fp8;TOOLRANK_HEADSstill forces them);toolrank learnstarts from no heads on the new backbone. Docker, compose and the Helm chart serve it by default;TOOLRANK_BACKBONE/embedding.backboneselect the base model again. - The documentation lives at https://yaman.dev/toolrank/ (the old address redirects there).
Fixed¶
- Heads given vectors of another width (an embedding endpoint serving another model) now say so, naming both widths, instead of failing inside the index build with a bare shape error.
- OpenAPI calls no longer keep cookies: one API response's
Set-Cookiewas sent with every later call to that host, whoever made it.
[0.1.0] - 2026-09-30¶
The first public release.
Added¶
toolrank ingest mcp | openapi | drop: MCP servers (stdio and streamable HTTP) and OpenAPI 3.x specs into an ingest directory; a re-run syncs only what changed and embeds only new or changed tools.toolrank search: Qwen3-Embedding-8B with the packaged heads, adaptive K, and a persistent vector index (numpy, FAISS HNSW or pgvector).toolrank serve: an MCP proxy with two tools,search_toolsandcall_tool, in front of every ingested tool; a REST API (/v1/search,/v1/rank,/v1/call,/v1/tools); API keys; a usage log that ties each call to the search that found the tool.- Tool search for Claude's Messages API (
tool_reference) and OpenAI's Responses API (client-sidetool_search). - Framework adapters: LangGraph (langgraph-bigtool retrieval), LlamaIndex (a tool retriever) and the LiteLLM proxy (a tool filter).
toolrank finetune: heads trained on your own request-to-tool pairs, the epoch picked on a dev set.toolrank evalandtoolrank compare: ToolRet, LiveMCPBench and MCP-Zero under ToolRet's protocol.toolrank heads pull; thetoolrankDocker image (amd64, arm64), a Dockerfile that bundles vLLM, and compose examples for both.