toolrank heads v0.1 for Qwen3-Embedding-8B¶
Two small learned heads that sit on top of Qwen3-Embedding-8B
sentence embeddings for tool retrieval. The state head reads the request, the action head reads a
tool; each computes x + MLP(x) (a skip head: 4096 → 1536 → 1536 → 4096, GELU, LayerNorm), then
L2-normalises, and a tool's score is the cosine of the two. The heads start as the identity, so
training only adds a correction on top of the base model.
Code, documentation and benchmarks: github.com/yasinyaman/toolrank,
yaman.dev/toolrank; the package:
pip install toolrank.
| File | toolrank-heads-qwen3-emb-8b-v0.1.npz, float16, 59.8 MB |
| sha256 | f3c101251b9c23925e2925bc02c4492e6c3dea79bcfa1b56715affb20f9f72f0 |
| Parameters | 29.9M (both heads) |
| Format | numpy .npz, loaded with allow_pickle=False; runs without torch (toolrank.adapters.heads_np.NumpyHeads) |
| Source checkpoint | qwen_full_skip_neg0_e5.pt (torch), exported with toolrank heads export --dtype float16 |
| License | Apache-2.0, like toolrank; see Training data for the data it was trained on |
Serving (stored in the checkpoint's cfg, used as defaults by toolrank)¶
- Backbone:
Qwen/Qwen3-Embedding-8B, last-token pooling, 4096-d, rows L2-normalised,truncate_prompt_tokens8192. bf16 or FP8: with the backbone in FP8 (vLLM's--quantization fp8) every benchmark stays within a query or two of bf16 (ToolRet 53.94 / 47.27, LiveMCPBench 53.48, MCP-Zero 79.51). - Tools: toolrank's
documentationtext. For ingested MCP and OpenAPI tools that is{"server", "name", "description", "inputSchema"}as JSON. - Requests:
Instruct: {instruction}\nQuery: {request}(instruct_query). The default instruction, chosen among three on LiveMCPBench and MCP-Zero, isGiven an agent's request for a tool, retrieve the MCP tool that fulfills it.
Training¶
- Data: 206K request–tool pairs from ToolRet-Training-20w
(
mangopy/ToolRet-Training-20w). Pairs whose request equals a ToolRet benchmark request were dropped. - Objective: InfoNCE with in-batch negatives only. The dataset's mined negatives cost up to 10 points and were not used.
- Optimisation: lr 1e-5, batch 512, 5 epochs; epoch 4 was picked on held-out training pairs.
- The backbone stays frozen. Its vectors come from a cache, so training takes minutes.
Training data license¶
ToolRet-Training-20w states no license: its dataset card has no license field and no license
text. The ToolRet code repository (mangopy/tool-retrieval-benchmark) is Apache-2.0, but that does
not cover the data, and the data itself is drawn from earlier benchmarks with their own terms. The
toolrank maintainers publish these heads accepting that risk.
If you need clean data provenance, use Qwen3-Embedding-8B without heads (Apache-2.0; 1–3 points lower) or train your own heads on your data.
Results¶
All numbers are w/ inst, and each set is scored under its own protocol (benchmarks):
| ToolRet NDCG@10 (micro / cat-macro) | LiveMCPBench Recall@5 | MCP-Zero top-1 | |
|---|---|---|---|
| Qwen3-Embedding-8B | 51.11 / 46.54 | 50.82 | 78.19 |
+ these heads (torch .pt) |
54.03 / 47.14 | 53.03 | 79.87 |
+ these heads (this .npz, numpy) |
54.03 / 47.13 | 53.03 | 79.87 |
- ToolRet has 44,453 tools and 7,961 queries.
- LiveMCPBench has 525 tools and 94 tasks; the server name is in the tool text.
- MCP-Zero has 2,792 tools and one LLM-written request per tool; the server name is in the tool text.
- float16 against the torch checkpoint: projection cosine ≥ 0.999999, top-10 overlap ≥ 99.89% on all three sets.
Use¶
toolrank search --data data/mytools "create an invoice for this customer" --clm-ckpt default
toolrank eval --data data/toolret --scorer clm --clm-ckpt path/to/toolrank-heads-qwen3-emb-8b-v0.1.npz \
--emb-url http://127.0.0.1:8091/v1 --emb-model qwen3-emb --with-inst
--clm-ckpt default reads TOOLRANK_HEADS, else ~/.cache/toolrank/heads/, else downloads the file
and checks its sha256. The file is downloaded from https://huggingface.co/yasinyaman/toolrank-heads-qwen3-emb-8b/resolve/v0.1/toolrank-heads-qwen3-emb-8b-v0.1.npz.
TOOLRANK_HEADS_URL points the download at a mirror; the sha256 is checked all the same.
Limitations¶
- English requests and tool texts only.
- Trained on ToolRet's task mix. The gain is mostly on large ToolRet tasks and small elsewhere (+2 to +3 points out of domain).
- The ToolRet gate threshold (50 cat-macro) is not reached.
- The heads only fit Qwen3-Embedding-8B vectors. Another backbone needs its own heads.