How it works

Three inputs at once, one single line of reasoning

Screen, camera and microphone open at the same moment and arriving at the same place. It is the hardest part of the product to explain, because from outside it looks like a list of features and inside it is the opposite.

When someone reads that Saelyx "sees your screen, sees you through the camera and hears you", the natural reading is three features living in the same program. Like saying a Swiss army knife has scissors, a corkscrew and a file.

It isn't that. And the difference is the whole product.

01What "at once" actually means

It means the three inputs arrive at a single line of reasoning, at the same moment, and what comes out accounts for all three. There is no turn for the screen, another for the camera and another for what you said.

The proof is boring to describe and obvious to use: you say "this one here", pointing with your voice while looking at a spreadsheet, and hold a piece of paper up to the camera in the same sentence. If each input had its turn, that sentence could not be answered. You would have to say it three times, in three steps, re-explaining the previous one each time.

02Why it isn't three features side by side

Doing it in turns is easier and in a demo you cannot tell. In a demo everything is slow, the person showing it knows what they are about to say, and nobody interrupts.

You can tell in real work, and it shows up as tiredness, not as an error. It is the toll of putting the machine back in context every time: cropping the screen, describing what you are looking at, repeating the figure you just read. Each round costs little and there are fifty a day.

When asking costs a ritual, you don't ask: you guess.

03The hard part isn't capturing. It's the rhythm

Each of the three inputs arrives at a different speed and means a different thing:

  • Voice comes in turns, and you have to know when you've finished —and when you're just thinking, which is not the same.
  • The screen changes on its own, without warning, and almost always without mattering. A blinking cursor is not information.
  • The camera depends on what you do with your hands, and it usually brings the most valuable thing exactly when least expected.

The hard part is deciding which of the three matters this second. A system that attends to everything equally drowns; one that attends to a single thing makes you repeat yourself. And it has to be settled while the conversation carries on, not afterwards.

04And then there's knowing when to shut up

With three inputs open, the temptation to talk is constant: something is always happening. An agent that comments on every screen change is unbearable within four minutes.

So the other half of the work isn't perceiving, it is deciding when to say nothing. And being able to cut it off mid-sentence, which is the only thing that stops it being frightening to let it start: you talk over it and it stops dead, without finishing the thought out of politeness.

05What leaves your computer with all three open

This goes here and not hidden behind a link, because it is the question that follows the previous one.

When you share your screen, those images are streamed as video, at around one frame per second. Saelyx also reads the text on your own machine so it arrives exact, but that is an improvement on top and not a replacement: the video travels all the same. Same with the camera. And the audio of the conversation is processed by an external provider with retention switched off.

None of it is kept to train models. The whole detail, including the two times we got it wrong and corrected it, is in what leaves your computer.

06All three are yours, and separately

  • The camera is a separate permission. If you don't grant it, you still get voice and screen exactly the same. And while it is on, the green light on your Mac is on: that LED is driven by the hardware.
  • There is a preview. You see the same frame that is being sent. There is no second image.
  • One exception worth knowing beforehand: there is a setting to send one frame per turn instead of continuous, and in Meeting Mode that setting does not apply. The camera is streamed anyway, even if you chose otherwise.

07Why it was built this way when it needn't have been

It is written in the manifesto and it bears repeating here: all three at once, or none. Doing them in turns was easier and nobody would have noticed the difference in a demo.

You would have noticed, on the first day of real work, without being able to say exactly what was wrong. That is what the short road was buying.

08What it doesn't do

  • It isn't a video call. The camera isn't there to see your face: it is there for you to show it things that don't live inside the computer.
  • It doesn't watch on its own. You open the inputs, with system permissions, and they close when you close them.
  • It is Mac only, macOS 13 or later, and access is by invitation.

09Frequently asked questions

Can I use it without the camera?

Yes. The camera is a separate permission: without granting it, voice and screen work the same.

Does everything have to be open at once?

No. "At once" is what it can do, not what it requires. Each input opens and closes separately.

How do I know the camera is on?

By the green light on your Mac, which the hardware controls and no application switches off, and by the preview, which shows the same frame that is being sent.

Is anything it sees kept?

What you show it is processed at the time and not kept to train models. The only thing that persists is what you choose to keep, and you see that in a list you open and empty.

And in a meeting?

In Meeting Mode the camera is streamed anyway, even if you chose the snapshot setting. We cover it in full in no bot in your meeting.

Invite-only access

Ask for it for your Mac.

Saelyx sees your screen, hears you and acts with you at once, and starts answering in a measured 776 ms. I read every request by hand and I write to you as soon as there is room for you.