Architecture¶
This page is the short version. The design record holds the measurements and the alternatives that were rejected.
The pipeline¶
graph LR
A[Hook] --> B[Daemon]
B --> C[Shorten]
C --> D[Speak]
D --> E[Play]
The hook is Claude Code's Stop hook, calling the daemon over a Unix socket. Shortening goes
to Ollama gemma3:4b only for replies over 200 characters once code, tables and paths are
stripped. The player is PowerShell on
Windows, reading the WAV that Kokoro wrote to audio_dir.
- Hook. The
keryx hookentries are async, so Claude Code never waits on them.Stopsends the reply,Notificationsends permission and elicitation prompts, andSessionStartandUserPromptSubmitstart the daemon and preload the summarizer.bin/keryx-replayruns synchronously onUserPromptSubmit, screens the prompt in the shell, and starts Python only for a short prompt that could be "say that again". - Daemon. One daemon per machine listens on a Unix socket, guarded by a lock file so a second daemon exits instead of taking over the socket. One queue serves every session.
- Shorten. The daemon strips code, tables, headings and paths. A reply of 200 characters or fewer is spoken as written. A longer one goes to the Ollama model, which rewrites it as one or two sentences. If Ollama is unreachable or fails, the opening sentences are spoken instead, up to 320 characters.
- Speak. Kokoro-82M, through
kokoro-onnx, turns each sentence into audio. It runs on CUDA when thecudaextra is installed and falls back to the CPU provider otherwise. Synthesis runs up to two sentences ahead of playback. Each sentence is brought to the configured loudness (-16 LUFS by default) and limited so sample peaks stay under -1.5 dBFS. - Play. One long-lived PowerShell process plays each WAV through MCI and reports when the clip stops. The WAV must sit on a Windows drive.
Why playback goes through Windows¶
WSLg's PulseAudio sink suspends when idle and then drops or hangs streams
(microsoft/wslg#1392). PowerShell plays
reliably. A fresh powershell.exe costs 309 to 444 ms, so one process stays open and takes
play, mode, stop, duck and unduck lines on stdin.
Sessions and terminals¶
- A new prompt cancels speech owned by that session or by the same terminal.
- Each request carries a
prompt_id. Hooks run async, so a reply's hook can land after the next prompt's, and the daemon drops a reply from an older prompt than the session's latest. - A terminal is its Claude Code process, found as the hook's nearest ancestor named
claudeorclaude-code. An npm install, which runs asnode .../@anthropic-ai/claude-code/cli.js, counts too. It holds its voice while that process runs, through/clearand resume. - A repo keeps the voice it was first given, stored by label in
voices.json. The repo name is read from.gitdirectly, so a prompt never waits on agitsubprocess. - Headless sessions (
claude -pand the Agent SDK) stay silent. - The daemon exits after 4 hours idle.
Cold start¶
Ollama aborts a model load when its client hangs up. The summarizer therefore loads the model
first under a 120 second timeout, then generates under 30 seconds. The model stays loaded 30
minutes after its last use. SessionStart and each prompt preload it, so the load happens
while you type and while Claude works.
Why "keryx"¶
A keryx (κῆρυξ) was a herald in ancient Greece: the one who carried a message and announced it aloud, briefly, to the people it concerned. keryx does the same for a Claude Code reply.