Speak
Bring intent and context out loud, in the language you already use. The agent captures decisions as you talk.
Voice-first AI workstations
Engineering Capability Map · 2026A spoken AI agent runs the whole session with you — it listens, talks, can be interrupted mid-sentence, and calls tools that actually change what is on screen. You speak; real work is generated, previewed live, and shipped.
Bring intent and context out loud, in the language you already use. The agent captures decisions as you talk.
The agent turns direction into visible decisions, live previews, and real artifacts — you watch them lock in.
Leave with production code, assets, a playable link, or a structured package — ready to use, not a transcript to paste.
Not a chat window. Not a prompt box. A live workstation the agent operates with you.
The same voice-first architecture independently recurs across three built workstations and an in-house voice test lab — each a real-time agent, real barge-in, a tool catalog that mutates a live canvas, and constrained runtime generation.
Design Apple Liquid Glass interfaces by voice or text. Generated native UI renders sub-second on-device, then you hand off SwiftUI code to your IDE.
Speak a world into play. Talk through the hero, obstacles, and world; a hand-built runtime keeps a real playable game on screen — 2D and 3D from one spec.
Talk, direct, refine, export. Shape a brief in conversation and watch logos, app icons, palettes, and backgrounds generate and lock on a live layered canvas.
+ Brain Dump — iOS thought-capture, preparing for the App Store · + an in-house synthetic-customer voice test lab (below)
Full-duplex spoken agents that listen, talk, take tool calls, and can be interrupted mid-sentence — behaving like a senior operator at the desk, not a chatbot in a side panel.
full-duplex · tool-calling · barge-in
While you talk, real artifacts generate and render in front of you — native UI, brand assets, or a playable game — and the preview updates sub-second. Never a mockup.
sub-second render · hot-reload · WYSIWYE
We constrain what the model may emit so every session ends in output that actually renders and runs — and most edits are cheap structural changes, not full re-rolls.
validated spec · guaranteed-usable output
Voice and typed chat are the same agent driving one tool catalog and one pipeline — start talking, finish typing — and the same architecture powers native Mac and the browser.
voice+text parity · macOS + web
A stateful spoken agent made safe against races and stale acknowledgements, that only claims a change that actually happened — and billing that never double-charges.
state-versioned · honest ledger · retry-safe
An in-house synthetic AI "customer" talks to the live product hands-free and grades each session against strict pass/fail gates — catching failures before real users do.
synthetic-customer lab · fail-closed gates
A realtime voice transport and a text transport each contribute only a decoder that produces a common tool-invocation — converging on one dispatcher. Add a surface, write only a decoder. The provider key never leaves the server.
Voice transport
Realtime spoken agent
full-duplex over WebRTC · echo-cancellation barge-in · single-flight tool loop
Text transport
Typed chat agent
function-calling · can barge into a live voice turn · same tool catalog
▼ common tool-invocation ▼
One dispatcher · one state machine
per-stage tool whitelist · versioned state snapshot injected into every turn of both transports · runtime-honesty gate
▼ constrained generation ▼
01 GATE
Kill-switch & wallet reserve
idempotent · release-on-failure
02 GENERATE
Server-authored, key-safe
structured output · strict schema
03 VALIDATE
Local structural & render check
coerce · clamp · retry-without-recharge
04 RENDER
Live preview + handoff
sub-second · export-ready
Server-authoritative thin client. The browser or app never holds a model key — only the connection handshake is relayed through our own backend, which also assembles the instruction set, injects the tool catalog, and meters usage. Config, models, prompts and limits are centralized so behavior changes ship without a rebuild.
On macOS we parse generated interface code into an intermediate representation and render a real native view graph — sub-second, inside a sandbox that forbids running code — with a hand-written GPU-shader approximation for export where the OS won’t composite live glass.
The model fills a validated spec that our hand-written templates render, instead of emitting raw code — so every game is playable and every component renders, and iterating is a cheap structural edit.
You can talk over the agent and it stops and adapts — built on echo-cancellation rather than muting the mic — so interrupting feels like interrupting a person, not toggling a walkie-talkie.
A single-flight tool loop with a priority follow-up queue and de-duplicated cancels means the workstation state never races the model — even under interruption and overlapping requests.
Every action is stamped with a state version, so the agent only confirms a change that matches the current state — no stale or out-of-order acknowledgements — and it only claims what the app actually committed.
A server-side keying-and-tiling pipeline converts a single generated background into layered, seam-scored parallax scenery with automated quality checks — and one deterministic core drives both 2D and 3D runtimes.
An AI "customer" talks to the live product hands-free over a simulated microphone and grades each session against fail-closed gates — including proving the output really reached a working state, and correlating exact per-session cost against the ledger — before real users hit an edge case.
The transferable part
The voice-first spine — realtime agent, constrained generation, live surface, server-authoritative economy, cross-surface parity — is not tied to any one product. It is a repeatable, documented pattern we have already fanned out across design, games, and brand, on both native Apple and the browser. Any workflow that today means a human clicking through a specialist tool can become a conversation with a machine that does the work and shows you the result.
Straight about scope
These are engineering claims, made honestly. One workstation is publicly shipped on the Mac App Store; the browser workstations are live in test. The products are deliberately human-led, agent-assisted — the person makes the decisions; the agent does the heavy lifting and narrates what it did. The guarantee is that generated output is structurally valid and renderable, not that it is flawless. We would rather show you the real system than a highlight reel.
Let's talk
We build focused AI products and partner on voice-led workstation experiences. If your team is thinking about deploying real-time voice agents into a specific workflow, we have already solved a lot of the hard parts.