Skip to content
How it works

Exactly ten minutes: two failures in disguise

A voice session that cut out at exactly 600 seconds, and a fallback that dropped turns to the small model without telling anyone. In both cases the mechanism built to recover from the failure is what hid the cause.

On 15 September, in the middle of a meeting, a voice session cut out and came back on its own three seconds later. It came back with a blank conversation. It had been open for 600 seconds, not one more, and carried 58 messages.

A week later, on the 22nd, I found a second failure that had made no noise at all. In about fifteen hours of log, the fallback had dropped turns to the small model 96 times when the large one was supposed to answer. I saw it while reviewing spend, not because anything warned me.

Two September failures in two stripsTop, the duration cap: a voice session that closes at 10 minutes with the reason “maximum duration”, another opened 4 minutes 31 seconds later whose close time is calculated at 14 and a half minutes, and today's cap of 2 hours. Bottom, the fallback: a task turn that the large model rejects, the fallback drops to the small model 96 times in about 15 hours, and the log keeps a reason that does not match the model that answered. After the fix, models that reject sampling no longer receive it; the fallback alarm is still pending.1 · The capmin051015One session600 sclose reason: maximum durationThe next one4 min 31 s later · close time calculatedAfter2 hthe most the engine let me set2 · The fallbackVoice turntask in progressLarge modelrejects samplingFallback96 times in about 15 hSmall modelanswersTurn logmodel: small≠reason: task in progressthe event is logged; no alarm watches itsampling no longer sent to models that reject itfallback alarm: pending
Top: the cap, one session that closed at its limit and the next one, whose closing time could be calculated. Bottom: the fallback and what the log recorded. Each figure comes from a single case.

01“It died and came back”

A session that cuts out and returns by itself, with a blank conversation, looks like a network drop. That's how I read it at first: the network, the computer, or the voice engine.

What didn't fit was the arithmetic. A network that fails doesn't usually do it at exactly ten minutes. And the automatic reconnect, which took three seconds, left everything as new except what had been said.

02A round number is a cap

Before touching anything I asked the voice engine, the external provider I describe in what leaves your computer, how that conversation had ended. It keeps a closing reason for each one, and this one said it had exceeded its maximum duration. The two voice configurations I had in production, the Mac's and the one for iPhone, iPad and watch, both carried a cap of 600 seconds.

That was enough. Another session, opened four and a half minutes later, had its closing time calculable in advance: its opening time plus 600 seconds. A failure you can predict with an addition is a setting.

What made it look random was the automatic reconnect. It did its job well, bringing the voice back in three seconds, and in doing so wiped the trail: all that was left to see was “it dropped and came back”.

That same day I raised it to two hours on both configurations, the most the engine would let me set. On 25 September I read them back from the service and they were still at two hours.

03The fallback that made no noise

On 18 September I moved to a new version of the large model. That day a set of meeting minutes came out with a visible error: the new model rejected a sampling parameter (one of those that govern how much risk it takes when choosing each word) that the minutes route was sending. I fixed that route that day.

The voice route builds its own request and had the same defect, but that one made no noise. With a task in progress and no extended reasoning, it sent the same parameter, the new model rejected it, and the fallback did what it had been told to do, drop that turn to the small model and carry on. No visible error. In the log, between 21 and 22 September, there are 96 of those drops in about fifteen hours.

The log of every turn has one field with the model that answered and another with the reason it was chosen. The fallback changed the first and left the second alone: the reason kept saying “task in progress”, which is the reason for using the large model. The fallback's own event was written down, but no metric or alarm watched it.

The fix was a list of the models that reject those parameters, which no longer get them. First I measured it with a real one-token call to the service: the large model rejected two of the parameters and the small one accepted both. The test that protects it fails against the old code, which is the only thing that gives it value: a test that passes first time and has never been seen to fail proves nothing.

It wasn't enough. In a little over an hour, another 18 turns fell to the small model because of a third parameter: the request carries three of these, spread over two different fields, and my first fix had covered two. Second fix, the same day.

Two rules came out of this. A model change gets tested with a real call at every place that builds the request, with its real parameters: fixing only the route that had failed in plain sight was the mistake, because it left the voice route untouched. And a fallback that changes what you had decided has to write that down where someone will read it, and have an alarm.

I didn't remove the fallback to make errors show: I'd rather have a voice that answers with another model than one that goes quiet. What's missing is a way to see when it acts.

04What is still open

As of 6 October 2026, I still have three things open.

  • The ceiling exists: it sits at two hours, and a longer meeting would reach it just as the September one reached ten minutes. What would fix it is renewing the session before it expires, carrying over what was said. It has been on my list since 15 September and is still open.
  • The fallback still has no alarm. The event is in the log and only shows up if someone looks for it; I got to it through spend. I've had the rule written down since 22 September; the alarm, not yet.
  • Two cases are not statistics: for the 96 answers I know which model answered, not whether it answered worse.

05Frequently asked questions

Why did it cut out at exactly ten minutes?

Because both voice configurations carried a cap of 600 seconds per session. It was a value set in the voice engine, not a network failure.

Is it still happening?

The ten-minute cap, no. Since 15 September both configurations are set to two hours, and on 25 September I read them back from the service and they were still that. The two-hour ceiling does exist.

What happens if a meeting runs past two hours?

It reaches the two-hour ceiling, the most the voice engine would let me set. Renewing the session before it expires, carrying over what was said, would fix it and is still open.

How did you find out the fallback was dropping to another model?

By reviewing spend, not through an alarm: task-in-progress turns showed up answered by the small model.

Is there an alarm on that fallback yet?

Not yet. The event is in the log, but no metric or alarm watches it. The rule is written down; the alarm is pending.


— Adianny
Senior DevOps engineer. I've spent about four years working with machine learning and MLOps, now as a tech lead; Saelyx is my side project.

Invite-only access

Ask for access before your next meeting.

Saelyx sees your screen, hears you and acts with you at once. I read every request by hand and I write to you as soon as there is room for you.