Voice-first AI workstations

Engineering Capability Map · 2026

We build creative machines you can talk to.

A spoken AI agent runs the whole session with you — it listens, talks, can be interrupted mid-sentence, and calls tools that actually change what is on screen. You speak; real work is generated, previewed live, and shipped.

  • Liquid Glass Lab · macOS · live on the Mac App Store
  • Games Lab · browser · public test
  • Brands Lab · browser · public test
01 — THE INTERACTION MODEL

Conversation is the interface. A working runtime is the result.

01 / SPEAK

Speak

Bring intent and context out loud, in the language you already use. The agent captures decisions as you talk.

02 / SHAPE

Shape

The agent turns direction into visible decisions, live previews, and real artifacts — you watch them lock in.

03 / SHIP

Ship

Leave with production code, assets, a playable link, or a structured package — ready to use, not a transcript to paste.

Not a chat window. Not a prompt box. A live workstation the agent operates with you.

02 — PROOF OF WORK

Concrete products, not a demo

The same voice-first architecture independently recurs across three built workstations and an in-house voice test lab — each a real-time agent, real barge-in, a tool catalog that mutates a live canvas, and constrained runtime generation.

◍ macOS · native

Liquid Glass Lab

Live · Mac App Store

Design Apple Liquid Glass interfaces by voice or text. Generated native UI renders sub-second on-device, then you hand off SwiftUI code to your IDE.

◍ browser · web

Games Lab

Public test

Speak a world into play. Talk through the hero, obstacles, and world; a hand-built runtime keeps a real playable game on screen — 2D and 3D from one spec.

◍ browser · web

Brands Lab

Public test

Talk, direct, refine, export. Shape a brief in conversation and watch logos, app icons, palettes, and backgrounds generate and lock on a live layered canvas.

+ Brain Dump — iOS thought-capture, preparing for the App Store · + an in-house synthetic-customer voice test lab (below)

03 — WHAT WE ARE GOOD AT

Six capabilities, proven in shipping code

P1

Real-time voice agents that operate the workstation

Full-duplex spoken agents that listen, talk, take tool calls, and can be interrupted mid-sentence — behaving like a senior operator at the desk, not a chatbot in a side panel.

full-duplex · tool-calling · barge-in

P2

Live generation into a working preview

While you talk, real artifacts generate and render in front of you — native UI, brand assets, or a playable game — and the preview updates sub-second. Never a mockup.

sub-second render · hot-reload · WYSIWYE

P3

Reliability by construction

We constrain what the model may emit so every session ends in output that actually renders and runs — and most edits are cheap structural changes, not full re-rolls.

validated spec · guaranteed-usable output

P4

One brain, many surfaces

Voice and typed chat are the same agent driving one tool catalog and one pipeline — start talking, finish typing — and the same architecture powers native Mac and the browser.

voice+text parity · macOS + web

P5

Production-grade state, honesty & economy

A stateful spoken agent made safe against races and stale acknowledgements, that only claims a change that actually happened — and billing that never double-charges.

state-versioned · honest ledger · retry-safe

P6

We test voice agents the hard way

An in-house synthetic AI "customer" talks to the live product hands-free and grades each session against strict pass/fail gates — catching failures before real users do.

synthetic-customer lab · fail-closed gates

04 — HOW IT WORKS

Two transports, one brain

A realtime voice transport and a text transport each contribute only a decoder that produces a common tool-invocation — converging on one dispatcher. Add a surface, write only a decoder. The provider key never leaves the server.

Voice transport

Realtime spoken agent

full-duplex over WebRTC · echo-cancellation barge-in · single-flight tool loop

Text transport

Typed chat agent

function-calling · can barge into a live voice turn · same tool catalog

▼   common tool-invocation   ▼

One dispatcher · one state machine

per-stage tool whitelist · versioned state snapshot injected into every turn of both transports · runtime-honesty gate

▼   constrained generation   ▼

01 GATE

Kill-switch & wallet reserve

idempotent · release-on-failure

02 GENERATE

Server-authored, key-safe

structured output · strict schema

03 VALIDATE

Local structural & render check

coerce · clamp · retry-without-recharge

04 RENDER

Live preview + handoff

sub-second · export-ready

Server-authoritative thin client. The browser or app never holds a model key — only the connection handshake is relayed through our own backend, which also assembles the instruction set, injects the tool catalog, and meters usage. Config, models, prompts and limits are centralized so behavior changes ship without a rebuild.

05 — HARD PROBLEMS WE SOLVED

The parts most teams gloss over

01

Rendering AI-generated native UI without compiling it

On macOS we parse generated interface code into an intermediate representation and render a real native view graph — sub-second, inside a sandbox that forbids running code — with a hand-written GPU-shader approximation for export where the OS won’t composite live glass.

02

Constrained generation for guaranteed-usable output

The model fills a validated spec that our hand-written templates render, instead of emitting raw code — so every game is playable and every component renders, and iterating is a cheap structural edit.

03

Full-duplex barge-in that feels human

You can talk over the agent and it stops and adapts — built on echo-cancellation rather than muting the mic — so interrupting feels like interrupting a person, not toggling a walkie-talkie.

04

Serializing an inherently async provider against live state

A single-flight tool loop with a priority follow-up queue and de-duplicated cancels means the workstation state never races the model — even under interruption and overlapping requests.

05

A state-safe, honest spoken agent

Every action is stamped with a state version, so the agent only confirms a change that matches the current state — no stale or out-of-order acknowledgements — and it only claims what the app actually committed.

06

Turning one AI image into a seamless game world

A server-side keying-and-tiling pipeline converts a single generated background into layered, seam-scored parallax scenery with automated quality checks — and one deterministic core drives both 2D and 3D runtimes.

07

Testing the agent with a synthetic customer

An AI "customer" talks to the live product hands-free over a simulated microphone and grades each session against fail-closed gates — including proving the output really reached a working state, and correlating exact per-session cost against the ledger — before real users hit an edge case.

The transferable part

One architecture. Many workstations. Many markets.

The voice-first spine — realtime agent, constrained generation, live surface, server-authoritative economy, cross-surface parity — is not tied to any one product. It is a repeatable, documented pattern we have already fanned out across design, games, and brand, on both native Apple and the browser. Any workflow that today means a human clicking through a specialist tool can become a conversation with a machine that does the work and shows you the result.

  • Design systems
  • Creative & brand
  • Interactive & games
  • Native Apple apps
  • Browser tools
  • Domain workstations

Straight about scope

These are engineering claims, made honestly. One workstation is publicly shipped on the Mac App Store; the browser workstations are live in test. The products are deliberately human-led, agent-assisted — the person makes the decisions; the agent does the heavy lifting and narrates what it did. The guarantee is that generated output is structurally valid and renderable, not that it is flawless. We would rather show you the real system than a highlight reel.

Let's talk

Have a workflow that should feel like a conversation?

We build focused AI products and partner on voice-led workstation experiences. If your team is thinking about deploying real-time voice agents into a specific workflow, we have already solved a lot of the hard parts.

Company

NextSense AI

marketing@nextsense.ai

Product, partnership & platform enquiries · nextsense.ai

Principal

Volkan · Founder

volkan@nextsense.ai

Architecture, deep-dives & direct conversation