Teaching local models to call tools they were not trained for.
A service agent in our runtime failed its very first tool call on every run. Then it spent 1,937 output tokens reasoning about which tools it thought it did not have, and invented two tools that exist nowhere in the codebase.
That was one live run, one local model, one prompt. Here is what the run produced:
tool not found: document [1,937 tokens of the model reasoning about which tools it must be missing, concluding wrongly that Document is not in this environment, and inventing generate_tool_suggestion and search to work around it]
The agent's def was fine. It granted tools: [Document]. The prompt told the model to call document, lowercase. Tool dispatch is an exact-match map lookup. The call never reached the tool and never could. The agent looked correctly configured in every listing, and the only symptom was a service agent producing confident nonsense.
That was an internal bug: a bundle prompt written wrong. We caught it and added a test that scans every bundled agent's prompt for the "call this tool" pattern and asserts the name is one the agent is actually granted. Done. Not the story.
The story is that even after we fixed the lowercase-name bug, small local models kept doing versions of the same thing. They would fail a tool call, then spend hundreds of tokens reasoning about what to do next, then either invent a tool that does not exist, or format a call in a shape the runtime does not accept, or give up and answer without the tool. This post is about what we did about it.
Why this matters. We run local models actively in the runtime. Not as an experiment. As the default for several service agents, including the memory consolidator's memory/ontologist (which decides where facts go), memory/extractor (which turns transcripts into facts), the document manager (which structures corpora on write), and the low-tier chat path (chat/local). Local models save cost, run offline, and stop a self-hoster from being tied to a cloud vendor. But every service-agent path in loomcycle needs to reliably call tools, and small local models are exactly the class that gets tool calling wrong.
Attempt 1: inject the tool-usage guide into the prompt
Our first read was: the model does not know how to call these tools, so let us tell it. Every tool ships a JSON schema, but small local models under-attend to the schema array on every call and end up guessing which operation to use and which fields are required.
So we built a compact digest. Context op=guide returns, for every tool granted to the run, a block like this:
{
"name": "Memory",
"side_effect_class": "mutating",
"ops": ["set", "get", "recall", "search", ...],
"required": {"set": ["scope", "key", "value"], ...},
"hint": "Use scope: user unless the fact is shared across users."
}
The digest is parsed from the tool's own input schema, so it cannot drift from what the model is actually being sent. We added an injectable prompt reference, {{tool:Context.guide}}, that expands the digest into the system prompt at request time. And we added a per-agent flag, inject_tool_guide, that appends the reference automatically so an operator does not have to hand-place it. The bundled LLM agents opted in.
Then we watched a live run.
A local model on Ollama wanted to search the web. It emitted its call as prose. The prose looked like this:
<tool_result> {"type":"function","function":{"name":"WebSearch","arguments":{"query":"..."}}} </tool_result>
Ollama's tool-call extractor expects <tool_call>. It saw <tool_result> and recovered nothing. The run ended with the un-executed block sitting in the answer field, and the operator wondering why no search ran.
The framing was ours. Our injected reference block used <tool_result> tags to demonstrate what a tool result looks like coming back from the runtime. The small local model looked at the reference block, decided that the loomcycle way to call a tool must be to emit blocks in that shape, and copied it. The injected help was teaching the model to format its calls the wrong way.
We removed inject_tool_guide from the bundled agents. The mechanism stayed (an operator on a strong-model agent can still opt in) but it is no longer a default. That was the first thing we learned: help injected into the prompt of a small local model gets read as an example to imitate, not as a spec to follow.
Attempt 2: proper help articles, and a pointer on every failed call
If in-prompt help does more harm than good on small models, help has to arrive somewhere else. Two changes went in.
First, we wrote proper per-operation help articles. One Markdown file per tool operation, under tools/<Tool>/<op>.md. Each article covers the arguments and their defaults, the actual result keys the runtime returns, the errors a model is likely to hit and what to do about each, and one to three call examples. There are 159 of these now, covering Memory (25), Document (47), History (11), Channel (8), Context (14), Skill (2), plus Recall and Agent. A lint asserts every call example validates against the tool's live input schema, so an article cannot teach a shape that would refuse.
Second, we changed what happens when a tool call fails. A failed call now ends with the exact Context op=help call for the operation that failed:
Path.mv: source "/entities/person/alice" does not exist
How to call it: call Context with {"op":"help","topic":"Path/mv"}
The idea is that a small model does not need to remember the help API. When it fails, the error message hands it the exact next call it should make. Learn the shape, then retry.
This narrowed the problem. Models with tool-call training that had merely forgotten the right op signature would follow the pointer, read the help, and retry with the right shape. But a lot of small models still did not. They would see the pointer, ignore it, and hallucinate a similar-looking call. Or worse, they would call the help topic, get a good response, and then still submit the wrong call shape on the next turn because the help had rolled out of the effective context.
Which brought us to the third attempt.
Attempt 3: put the help in the error, and keep the pair together in context
Two changes together, both required.
The first change is that a failed call now carries its own help. The pointer is still there for models that want to fetch a broader article, but the immediate response also carries the specific operation's schema shape and one worked example verbatim, inlined into the error itself. A model that failed Path.mv with an unknown source gets, in the same turn:
Path.mv: source "/entities/person/alice" does not exist
Schema for Path.mv:
op: required, must equal "mv"
path: required, absolute path to the resource to move
to: required, absolute destination path
scope: optional, one of agent | user | tenant
Example:
{"op":"mv","path":"/documents/draft","to":"/documents/final","scope":"user"}
For the whole article: call Context with {"op":"help","topic":"Path/mv"}
This is not just "structured errors" (a category, a retryable flag, a retry-after). Those we already had: every tool result carries errorCategory (one of transient | validation | business | permission), isRetryable, and retryAfterSeconds. The new part is that the schema and the example are IN the error body. A small local model that would not have followed a pointer to the help article now has the help sitting directly in the last thing it read.
The second change is stateful context management that keeps tool calls and tool results together across turns.
Loomcycle has a stateful context mode where each turn's input is a compact structured state rather than the full transcript (see the compact-context post). Before this arc, the stateful memo dropped tool call arguments and tool call results early to save space. That was the right call for prose-shaped tasks. It was the wrong call for a tool-calling loop, because when the memo dropped the failed call and the help response, the model on turn N+1 had no memory of what had gone wrong on turn N. It would fail the same way, get the same help, and repeat.
The fix keeps the tool call arguments and the last help response in the stateful memo across at least the next turn, so a correction is visible when the model needs it. The memo still evicts older tool calls to keep the state bounded, but the recent pair (the failing call and the help that would fix it) stays.
flowchart TB
subgraph BEFORE["Before"]
T1["Turn N: model calls Tool"] --> E1["Error: bad shape"]
E1 --> H1["Help pointer
(and article on request)"]
H1 --> M1["Stateful memo:
drops the call
drops the help"]
M1 --> T2["Turn N+1: model retries
with no memory of the correction"]
T2 --> E2["Same error"]
end
subgraph AFTER["After"]
T3["Turn N: model calls Tool"] --> E3["Error: bad shape
+ schema inlined
+ example inlined"]
E3 --> M2["Stateful memo:
keeps the call args
keeps the help response"]
M2 --> T4["Turn N+1: model retries
with the correction visible"]
T4 --> S1["Call succeeds"]
end
style BEFORE fill:#fff5f5
style AFTER fill:#f0fff4
style E2 fill:#ffe8e8,stroke:#c86a5c
style S1 fill:#e8ffef,stroke:#4a9e60
Along the way: what does not work
Three shapes of intervention we tried and either reverted or discarded. Naming them because they were reasonable, and someone will try them.
Injecting the tool guide as a bundle default. Reverted, as described above. The reference-block format gets copied as tool-call format on small local models. The mechanism stays for opt-in on strong-model agents.
Making tool usage an imperative protocol in the extractor prompt. We tried it once for a different feature (making the model always attach source snippets to a recall). Result: accuracy dropped from a live 0.396 to 0.0034 at 1.0 abstention, because the mandatory protocol consumed all of the model's attention and it stopped answering the question. A mandatory protocol competes with the task for attention on a small model. See the retrieval-quality post for the full story.
Trusting the model to elect the corrective call itself. On the same test set, three models given the same tool, same prompt, and same store diverged wildly in how often they used a helpful optional tool: cloud deepseek issued the call on 214 out of 150 questions (it re-fetched on many); qwen3.6 local issued 7 calls; the agentic-tuned ornith-1.5:35b local issued 17 for a statistically-insignificant +2.4 percentage points. Optional tools on small models are effectively unused unless the runtime insists.
The final benchmark
Four models on the same tool-calling bench, run against loomcycle v1.100.0 on 2026-09-29. Each model ran ten tasks (five tasks, two runs each). The tasks exercise the tool-call paths a real service agent hits: create a Document, upsert facts, search across scopes, walk a Path tree, find and summarise a previous chat. Cloud deepseek-v4-flash in stateful mode is the baseline ceiling; the other three are local models running in the service configurations we ship (medium, small, coder).
The "previous arm" column is arm K, the same benchmark run against loomcycle v1.99.0 before the stateful-memo change landed. It shows what changed between the two runs.
| Model (config) | YES | PARTIAL | NO | YES in K | Input tokens (K → L) | Tool calls (K → L) | Help calls (K → L) | Failed calls (K → L) |
|---|---|---|---|---|---|---|---|---|
deepseek-v4-flash (cloud, stateful) | 10 | 0 | 0 | 10 | 544K → 339K | 86 → 39 | 28 → 13 | 4 → 0 |
qwen3.6 (medium, append) | 9 | 1 | 0 | 10 | 662K → 588K | 21 → 19 | 1 → 0 | 3 → 3 |
gpt-oss (small, append) | 8 | 0 | 2 | 7 | 485K → 633K | 19 → 25 | 0 → 2 | 5 → 8 |
ornith-1.5:35b (coder, stateful) | 10 | 0 | 0 | 8 | 301K → 307K | 71 → 36 | 20 → 12 | 8 → 0 |
Two things stand out.
First, both stateful models hit 10 out of 10. Cloud deepseek-v4-flash and local ornith-1.5:35b, on the same tasks. That is what closes the arc: on the tasks this benchmark tests, a local model in the coder configuration running in stateful mode matched the cloud ceiling.
Second, the stateful-memo change is what carried both stateful models over the line. ornith went from 8 in arm K to 10 in arm L; the two arm K runs that had failed had both timed out at 600 seconds after re-reading the same help article multiple times. On arm L, no help topic was read twice in any stateful run; in arm K the same article had been re-read up to four times. deepseek T3a alone dropped from 35 iterations and 193K input tokens in arm K to 5 calls and 42K input tokens in arm L. Across the whole deepseek configuration, input fell 38 percent and tool calls fell 55 percent.
The cost of keeping Context results in the stateful memo shows up on every step. A deepseek run's last call now carries about 6.9K input tokens at the median, up from about 5.2K in arm K. That is the price of not repeating help fetches. On this benchmark, it more than paid for itself.
Two append-mode notes worth naming. qwen3.6 dropped one task from 10 to 9 (a "find and summarise the previous chat" task where the model named the right chat but gave no topic; the chat had no summary and the model did not recap it). gpt-oss scored 8 of 10, and its failures are the kind a small local model produces when it under-attends to the schema: 2 Ollama parse errors, 1 call to a tool named ?…?, ids with junk characters, and a document that ended up with only half the intended sub-chunks because the create for the other half failed on an id string c?? too short to count as a truncated id.
Two follow-on fixes landed after arm L, and would show their effect in a re-run on a later release. The first covers ids ending in U+FFFD (the Unicode replacement character): the cut-id guard fired correctly once and the model retried with the full id, but a second case slipped through because the id ended in U+FFFD, which the guard did not recognise as truncation. The second covers a duplicate-document create: deepseek on one run created its document twice because the create_document result had been evicted from the stateful memo when the kept help article pushed it out; the fix keeps the last real action's result alongside the kept help. Neither is a benchmark result. They are what a real benchmark on a new release turns up: the seams that the previous release did not stress hard enough. Any future arm will find its own.
What this does not fix
Three limits worth naming.
Models with no tool-call training at all. A base model that has never seen a tool-call token in training will not learn from an inlined schema. This arc is about small local models that have SOME tool-call training but not for loomcycle's specific tool inventory. If your model has zero tool-call training, the closest workable path is a heavily structured output format the model can produce as text, then parse server-side.
Non-stateful mode. The help-in-error change works everywhere. The keep-the-pair-together change only works in stateful mode. In the default append mode, the failing call and the help are already visible in the running transcript, which is what makes append expensive. In compaction mode, the pair may get summarised before the model retries. If you are running a small local model on a long compaction-mode chat, prefer stateful.
Operators who write bundle prompts wrong. The lowercase-document bug that started this post was not a model bug. It was a bundle-prompt bug. A prompt that names a tool the agent does not have, or that names it with the wrong case against an exact-match dispatch, will fail no matter how much help the runtime injects into the error afterwards. The bundle test that catches this pattern runs in CI now, but any custom agent def with the same mistake will hit the same wall.
What comes next
The tool-help arc opens a three-post series on the September retrieval-quality work. Two follow-on posts (documents indexing, then memory retrieval) cover the surfaces this arc runs underneath. Everything three posts describe together is what a self-hoster on a local model can now run reliably.
The next arc is agent teams: state machines for multi-agent workflows, review gates for team outputs, and a visual canvas for constructing and debugging them. A different problem shape.
Companion reading: what structured memory costs, and what closes the gap; when one embedding per chunk stops working; compact context, still recall; the memory architecture drawn.