← All work
Live demo · runnableAI Systems

Presence

The turn-taking pipeline underneath a conversational avatar. A state machine you can interrupt mid-sentence, with every latency number measured on the page instead of quoted. MORPH came out of this one.

My role
Designed and built it. State machine, provider interface, presence layer, tests.
Stack
TypeScriptReact 19Finite state machineAsyncIterable streamingAbortSignalCanvas 2DVitest

The problem

A conversational presence, whether an avatar, a kiosk or a hologram, fails on turn-taking long before it fails on answer quality. Keep talking after someone starts speaking and it stops being a presence; it becomes a video player. Leave more than about half a second of silence after they finish and they assume they were not heard, so they start again. That produces the overlap it then has to survive. Both failures are state machine problems. Neither one is fixed by a better model.

What I built

Three pieces kept deliberately separate. A pure state machine owns turn-taking (idle, listening, thinking, speaking) with a written-out transition table where illegal events are rejected with a reason instead of silently swallowed. A provider interface streams tokens behind an abort signal, so a response can be abandoned mid-token. A canvas presence layer renders whatever state the machine is in, driven by real output rate. The machine emits effects as data and performs none of them itself, which is what lets all of it be tested without a microphone, a model, or a browser.

The hard part

Barge-in. Interrupting is easy to fake and hard to actually do: the in-flight generation has to stop producing, the audio has to stop, the floor has to return to the visitor, and none of it can leave a half-finished turn behind that corrupts the next one. The abort has to propagate into a stream that may be sitting in a timer or a network read, and the interrupted turn still has to be recorded rather than discarded. An interruption is data about latency, not an error. Getting that right meant making the machine pure so the ordering of those effects is something a test can assert instead of something you check by hand.

Evidence

  • Run it on this pageTake a turn, then interrupt mid-sentence. Every number in the panel is measured from your session.
  • 38 testsLegal and illegal transitions, barge-in before and after the first token, latency accounting, and that an aborted stream stops promptly rather than finishing quietly.
  • No key, no account, no networkThe default provider is scripted so anyone can exercise the real pipeline. The page says so rather than implying a model is answering.

Try it

Take a turn, then talk over it.

Start a conversation, send a turn, then type while it is still replying. That is barge-in: generation aborts mid-token and the floor comes back to you. The panel underneath is measured from your session, not quoted.

idleDormant. Nothing is listening.

Replies are canned. The streaming, the abort path and every number on this page are real.

Last completed turn

budget 500ms to first token

Take a turn and the timings for it appear here. Nothing is quoted — every number is measured from this session.

Architecture

How a request moves through it.

  1. Turn capturein

    An utterance ends, by text or voice. The machine does not know which, so the capture method can change without touching turn-taking.

  2. Presence state machinesvc

    Pure transition table. Emits effects as data; performs none of them. Illegal events are rejected with a reason.

  3. Dispatchsvc

    Mic closes, request goes out, and the turn's timing marks start.

  4. Provider streamext

    Cancellable async token stream behind an AbortSignal. Scripted by default; the hosted adapter is the same interface.

    ↳ Barge-in

    The visitor starts talking. Generation aborts mid-token, audio stops, the floor returns, and the abandoned turn is still recorded. An interruption is a latency measurement, not an error.

  5. Presence layerworker

    Canvas render driven by machine state and real output rate. The visual is a readout, not a loop.

  6. Turn metricsout

    Response gap, time to first token, speaking duration. Recorded for interrupted turns too.

One conversational turn. Barge-in is the path that matters. It can fire at any point after dispatch, and everything downstream has to unwind cleanly.

Technical decisions

What was chosen, and what it cost.

The state machine is pure and performs no effects.

WhyTurn-taking is the part that has to be correct, and purity is what makes it testable. Aborting mid-generation, rejecting an impossible event, the ordering of cancel-then-stop-then-open: all of it is asserted in a test rather than verified by clicking.

CostThe caller has to perform the effects faithfully; a machine that ran them itself would be harder to get wrong at the call site.

Illegal transitions are rejected with a reason, not ignored.

WhyA machine that quietly swallows impossible events hides the bug that produced them. A rejection surfaces in the UI during the demo, which is how you find out the capture layer is firing events out of order.

CostMore states to handle explicitly, and a table that has to be maintained.

The default provider is scripted and requires no key.

WhyA demo behind an API key is a demo nobody runs. Everything that makes this interesting (streaming, cancellation, the latency accounting) is real without one, and the page discloses exactly what is canned.

CostThe replies are not a model's. That is stated on the page rather than implied away.

The interrupted turn is recorded, not discarded.

WhyAn interruption tells you the response gap was too long. Throwing those turns away deletes precisely the data that explains why the presence felt wrong.

CostMetrics have to handle turns where a first token never arrived.

The presence layer is Canvas 2D, not WebGL.

WhyThis page argues the pipeline is the hard part. Shipping a megabyte of 3D to draw a status indicator would undercut that, and a few hundred lines of arithmetic reads the same from across a room.

CostNo depth, no real volumetrics. It evokes a projection rather than simulating one.

Constraints

  • Must run with no API key, no account and no microphone permission.
  • Barge-in has to be possible at any point after dispatch, including before the first token.
  • Every number shown is measured in that session. Nothing is quoted from a benchmark elsewhere.
  • Decorative motion respects reduced-motion; the state is legible without animation.

How it is verified

  • 38 tests across the machine and the provider, run by npm run verify alongside typecheck, lint and build.
  • Transition tests cover every illegal pairing and assert that a rejected event moves nothing and emits nothing.
  • Barge-in is tested both after the first token and before one ever arrives.
  • The abort test asserts the stream stops within a token or two, not after quietly finishing the reply.
  • A turn built from out-of-order timestamps is asserted never to report a negative duration.

Outcome

  • A conversational turn pipeline anyone can run in a browser and interrupt, with its own timings on screen.
  • Turn-taking logic that is independent of capture method. Swapping text for voice activity detection does not touch it.
  • A provider interface a hosted model drops into without the machine or the presence layer changing.

What I would tell someone building this

  • The thing that makes a conversational presence feel alive is almost never the model. It is whether it can stop talking.
  • Purity paid for itself immediately. The ordering of cancel, stop-audio and open-mic is the kind of detail that is invisible by inspection and obvious in a test.
  • Recording interrupted turns instead of dropping them turned interruption from a failure case into the measurement that explains the failure.
  • A demo that needs a key is a demo nobody runs. Making the no-key path the default is the difference between something a stranger can evaluate and something they have to trust you about.

Questions about how this was built, or want something like it?