The model only knows a tool exists if it is declared in its catalogue, with a name, a description and parameters. A tool is an action the agent can ask the app to perform (find a file, look at the calendar, write some text); if it isn't declared, the feature can be flawless in the app and never get used. If it is, the catalogue travels with every turn, next to what you say and what the screen shows (the three inputs arrive together), and it is paid for on every sentence.
01A list that lived outside git
The list of what Saelyx can do wasn't in the repository. It lived in the agent's configuration, in an outside provider's console, with no tests, no review and no diff. The app's parity check covered dispatch (what it runs when the model asks for something) and couldn't cover what wasn't in git.
On 25 July 2026 I saw the gap in full: the app could run 52 tools and the agent declared 24. The 28 missing ones included things as ordinary as finding a file or checking the calendar. The model didn't know they existed, so it never called them.
The first thing I did was dump the live configuration into the repository, with one rule: every change in the console gets committed in the same motion. In September the repository became the source, with one file per tool and one line per property, so that trimming a paragraph doesn't look like twenty lines of noise.
02The bill: 145 ms and 6,860 tokens
I measured what declaring costs. On 25 July 2026, against the model that was then in production, I timed how long it takes to start answering with three different catalogues. From 24 to 52 tools, the first token slipped 145 ms per turn and the input grew by 6,860 tokens. With a conversation budget of 300 to 800 ms, 145 doesn't disappear into the noise.
The 28 were worth the price, but it is paid every time, even on a bare “yes”: it carries the same catalogue as the summary of a meeting.
| Change to the catalogue | First token | Input tokens | How I know |
|---|---|---|---|
| +28 tools declared (24 to 52) | +145 ms | +6,860 | Measured |
| Dropping the phone-only tools on the computer | −23 ms | −1,109 | Measured |
| Descriptions under 300 characters | ≈ −25 ms | ≈ −1,200 | Estimated; not done |
03Dropping the phone tools from the computer: 23 ms
The obvious idea: the computer doesn't need the tools that only make sense on the phone. I tried taking them out and won back 23 ms. Almost nothing.
The two measurements, the 145 ms rise and the 23 ms drop, fall on the same slope: about 21 ms per thousand tokens of catalogue. Two points can't prove more than that, but they suggest weight is what matters, not the count. Cutting every description under 300 characters would be worth, at that slope, about 25 ms (an estimate). I didn't: those descriptions tell the model when to use each tool and, in several, to treat what it reads as data and never as orders. You don't touch what holds the behaviour up for 25 ms.
On the phone the arithmetic looked better: it carried, on every turn, plenty of tools that only the computer can run, and July's estimate was to win back nearly the whole surcharge. But that needed a different list per device, and the app couldn't send its own when it opened the conversation. I had written down that what the app does choose is which agent to connect to.
That was false, and of everything here it's the one that stings most: the agent travels inside the conversation's credential, which the server generates. On 8 September I split the tools between two agents assuming each device would go to its own; both went to the same one and the Mac lost part of its abilities, the whole terminal included. I fixed it on the 9th: the server now splits the agents by kind of device. And before splitting anything now, I measure who each agent serves: an agent with zero conversations serves nobody, whatever it is called “phone”.
04A default value that left the tools mute, twice
When I declared the 28 I didn't say whether each tool expects a response, and the API set “no” on its own. Fourteen return a piece of data, and they came out mute: the model called them, the app did the work and the data never came back, so the agent kept talking without mentioning it. I spotted it by cross-checking the configuration dump against the tools that did work: the 24 good ones were exactly the 24 with that field set to “yes”.
I wrote the rule and thought that was enough. On 14 September it came back: a copy of the catalogue for the Mac agent had been born with the same value, on nearly half of its tools. I asked for the summary of the last meeting and the agent said there were no meetings. The app had answered correctly and its log said so; what failed was what the model received, a generic success phrase instead of the result, and that only showed in the conversation transcript, on the other side.
This time I closed it with a test: since 21 September that field has been part of each tool's versioned definition, and the test fails if a trim touches it.
05The September diet, and the target I missed
With the list in git I could read it whole. More than half of the tool text that can reach the model was descriptions. I trimmed three repetitions: an inventory of operations that appeared twice (in the description and in the parameter that selects them), a voice-confirmation protocol copied word for word into more than ten tools, and descriptions that repeated the agent's prompt.
Result, on the Mac agent: −23 % in what can reach the model, −16 % in what the API returns and −43 % in descriptions. It is the same cut with two denominators: the API adds to each property bookkeeping fields that don't reach the model and can't be removed, about half the characters of the schema. I trust the first figure.
Before touching anything I saved what each tool declared, and a test compares it with today's catalogue: no operation, parameter or closed list lost. The target, on the other hand, I didn't hit: I got a little over halfway. The schema also grew slightly, absorbing text from the descriptions, and I don't count that as a saving. The budget check fails at a cap set just above what's there now, not at the target, because a cap that fails on the day it ships gets switched off; the target is printed, with the distance still to go, every time I measure. The cuts have been in production since 21 September.
06What I don't know
- The measurements are from one day in July and one model, the one of the time; with today's model the slope could be different.
- I haven't measured on a real phone whether dropping the surplus wins back what I estimated, nor how far the first token fell with the diet: from the size of the block to latency there is a calculation, not a measurement.
07Frequently asked questions
What is a tool in a voice agent?
An action the agent can ask the app to perform, like finding a file or looking at the calendar. To the model it only exists if it is in its catalogue.
Why does every declared tool add latency?
Because the catalogue travels with every turn and the model reads it before answering: in my July measurements, about 21 ms per thousand tokens.
What happens if the app can do something the agent hasn't declared?
It will never be done, because the model doesn't know it exists. It happened to me with 28 tools until 25 July 2026.
Have you measured how much the September diet sped things up?
No. I measured the size of the block (23 % less on the Mac agent, in what can reach the model), not its effect on the first token in a real session.
Does the catalogue contain my data?
No. It describes what the agent can do, not what it sees of you. What travels and what is kept: what leaves your computer and where your data lives.
— Adianny
Senior DevOps engineer. I've spent about four years working with machine learning and MLOps, now as a tech lead; Saelyx is my side project.