Skip to content
How it works

The answer that took 102 seconds

A question typed in the middle of a meeting took 102 seconds to start being answered; the model took between 1.7 and 2.3 s. Here is where the rest went and what I changed.

On 5 October, in a real meeting, I typed a question into Saelyx's floating panel. It was the kind that asks for a fact and nothing else. The answer started to arrive 102 seconds later. Sixteen seconds after that the session ended, with the answer half-delivered.

01The model was fine

When something is slow, you look at the model. I looked there first: between 1.7 and 2.3 seconds to the first fragment of the answer, on every turn of that meeting. The voice engine itself reports it, turn by turn, and not one of them falls outside that range. A model that starts in about two seconds doesn't account for the other hundred.

02Three clocks

Between the keystroke and the answer there are three clocks, each in a different place and read with its own instrument. The timeline of that question:

t+0 s    I type the question
t+7 s    it reaches the voice engine
t+102 s  the answer starts
t+118 s  the session ends, mid-answer
  • The model: 1.7 to 2.3 s to the first fragment, as in the section above.
  • My queue: 7 s. It holds back whatever gets typed while the agent is busy, because a new message makes the engine throw away the answer it was halfway through. There's a cap on it; here 7 seconds went. I can see that in my own log.
  • The voice turn: 95 s. From the moment the question reaches the engine until it starts answering, according to the engine's own transcript. It is nearly all of it.

In a normal conversation the engine knows whose turn it is: you talk, you stop, it answers. In a meeting the microphone picks up the whole room, and I also pass the agent what the others are saying so it has context. To the engine, all of that is one user who never stops talking, in blocks of several thousand characters. In a meeting someone is almost always talking; that's what they're for.

While that turn stays open, the agent doesn't answer. And every new line that reaches it as a turn throws away whatever answer it was preparing. The question never found a gap in 95 seconds. It found one when the meeting ended.

03Two paths, and who decides

I left the voice turn as it was. What I changed is the route: whatever doesn't need the voice no longer goes through it.

Whatever doesn't ask to act (a question, a summary, some ideas) now goes to the model in writing, with the last lines of the meeting as context, and the answer is shown in the panel. It isn't also sent by voice, so it never gets answered twice. The three questions I typed that same afternoon, in another meeting, were answered, in full, in 2.4, 1.3 and 2.5 seconds.

Whatever asks for an action stays on the voice path, which is the one with the tools: closing the meeting, noting something down, searching the web. A "yes" or a "no" to something the agent said, redoing the last line ("make it shorter") and anything that needs tools only that path has go there too. If the written path fails (for example, no connection or no session), the same question goes out by voice, once. In the worst case, it behaves as it did before.

Where the 102 seconds went, and how each question is now routed Top, a bar drawn to scale: 7 seconds in my queue and 95 in the voice turn add up to 102 seconds from keystroke to answer. Below it, on the same scale, the model (1.7 to 2.3 seconds to the first fragment) and the written path (1.3 to 2.5 seconds) are barely visible. Bottom, a diagram: what gets typed goes through a decision. If it doesn't ask to act it takes the written path to the floating panel. If it asks for an action, is a yes or a no, asks to redo the last line or needs tools, or there is doubt, it takes the voice path. If the written path fails, the question goes out once, by voice. From keystroke to answer 102 s in total, all on the same scale t+0 t+102 s voice turn · 95 s 7 s · my queue For comparison, on the same scale: The model · 1.7 to 2.3 sto the first fragment Written path · 1.3 to 2.5 sfull answer 0 50 s 100 s Who decides the route What gets typed Asks to act? no yes, or in doubt Written path → floating panel 1.3 to 2.5 s Voice path commands · yes or no redo · tools if it fails: once, by voice
Top, to scale: the morning's question, with 7 s in my queue and 95 s of voice turn, and below it the model and the written path (three questions from that afternoon). Bottom, who decides the route of whatever gets typed.

When in doubt, voice wins, because the two mistakes don't cost the same. A question that lands on the voice path by mistake is slow, as it used to be. A command that lands on the written path by mistake doesn't get executed, because that path has no tools. That's why I made the list of what counts as a command generous on purpose.

04What I discarded, and why

The simplest thing was to send everything in writing, and it answered in two seconds. But that leaves commands without tools, and I'd rather have a slow answer than a command that doesn't happen.

The first version, the one from that morning, sent in writing only what looked like a question: a question mark or a question word. "Summarize the last bit" or "give me three ideas" didn't look like questions and sat waiting for a gap on the voice path. In the afternoon meeting, with no gap, the voice took 26 seconds or didn't answer, and over about thirteen minutes six typed messages got no answer. Whoever typed them saw nothing, not even that they were waiting. That same afternoon I flipped the rule: everything that isn't a command, a yes or a no, or a redo of the last line goes in writing. And what still goes by voice no longer stays silent: if no signal from the agent arrives within a few seconds, the panel says the voice is waiting for a gap.

I had already tried a more impatient engine: it split a sentence into three requests and I took it off on 25 September. In a meeting, with the room talking non-stop, the likely result is that it would answer the room. That last part is reasoning: I haven't measured it.

05It wasn't the first time

On 18 September, in Meeting Mode, a simple question took between 30 and 45 seconds. The model answered in 2.5 to 3.4 s per call, and the engine waited 19.7 and 26.6 s before taking the floor in the two waits I measured. That time the cause I found was mine: an "I'm still here" signal, sent every few seconds so the call wouldn't expire, told the engine the user was still active. I made it stay quiet while a reply is pending, the same day.

Both times I started the same way: separating the clocks before touching the model. It's my rule now: if a meeting latency can't be split into model, queue and turn, I'm not measuring anything yet.

06What is still not fixed

First, the voice: with no gap it took 26 seconds or didn't answer, and I haven't fixed that. I took one class of questions off that path, and commands still go through it.

It is one case of 102 seconds, measured to the second with two sources (my log and the engine's transcript), so the split between 7 and 95 could be off by one. For the written path there are three measured answers, and for the voice path with no gap, six typed messages; one meeting in the morning and another in the afternoon. I have no percentiles. And the two timings don't come from the same instrument: the 1.7 to 2.3 s are to the first fragment of the voice model, the 1.3 to 2.5 s are the full answer on the written path, which is a different path.

The written path answers in one or two sentences, with the last lines of the meeting and no tools, which is what I ask of it. And all of this is what happens on the Mac: I haven't measured it on the iPhone.

07Frequently asked questions

Why not make the voice engine take the floor sooner?

I already tried a more impatient setting and it split a sentence into three requests; I took it off on 25 September. In a meeting, with the room talking, the likely result is that it would answer the room, though I haven't measured that. I'd rather take whatever doesn't need that path off it.

What still goes through the voice path?

Commands (closing the meeting, noting something down, searching the web), a "yes" or a "no" to something the agent said, redoing the last line, and anything that needs tools only that path has. When in doubt, it goes by voice.

What happens if the written path fails?

The same question goes out by voice, once. If the voice doesn't find a gap within a few seconds either, the panel says so instead of staying silent.

Does it read my screen every time I ask?

No. Only if the question names it ("what does the screen say", "read this"). Then your own Mac reads the text of the window in front of you or, if the one in front is Saelyx itself, the text of the screen your cursor is on, leaving out Saelyx's own windows. That text travels with the question, trimmed, with any keys and passwords it recognizes masked and marked as data rather than instructions. If you are sharing a single window, or haven't granted the screen recording permission, nothing is read. The full detail is in what leaves your computer.

Does it work the same on the iPhone?

What I describe is what the Mac does today, which is where I measured it. I haven't measured it on the iPhone and I don't take it for granted.


— Adianny
Senior DevOps engineer. I've spent about four years working with machine learning and MLOps, now as a tech lead; Saelyx is my side project.

Invite-only access

Ask for access for your next meeting.

Saelyx sees your screen, hears you and acts with you at once. I read every request by hand and I write to you as soon as there is room for you.