React voice UI · local fixture demos

Cartesia Managed Agents React Voice UI

Build a Cartesia Managed Agent voice interface with browser WebSocket sessions, PCM16 audio, captions, client tools, interruption cleanup, and a local simulation.

Requires orb-ui@0.10.0 or later. SDK types, offline fixtures, and browser lifecycle tests have been checked. Live provider credential tests have not been run.

createCartesiaAdapter connects a Managed Agent conversation to orb-ui. It opens Cartesia's agent WebSocket, waits for session_ready, streams mono PCM16 input, plays returned audio, and preserves conversation and client-tool events. It supplies microphone capture and playback; your application supplies its own agent and a fresh browser access token.

Try the local simulation

The preview uses synthetic states and volume levels. It has no Cartesia connection, microphone request, API key, or paid calls. Use it to inspect the interaction before connecting your own agent.

Setup and credential ownership

bash
npm install orb-ui@0.10.0
# Optional, on your trusted backend only:
npm install @cartesia/cartesia-js@4.2.0

The browser adapter uses native WebSocket and Managed Agent API version 2026-08-14. The official server SDK 4.2.0 is used to typecheck token generation. Its TTS socket helpers are separate from this Managed Agent protocol. No Cartesia runtime dependency is bundled into orb-ui.

Use an agent configured in the visitor or developer's own account. Your existing authenticated backend can mint a short-lived token from that account's server credential:

ts
// Trusted server module. Call only after authorizing the caller's own provider account.
import Cartesia from '@cartesia/cartesia-js'

export async function mintCartesiaAgentToken(serverApiKey: string) {
  const client = new Cartesia({ apiKey: serverApiKey })
  return client.accessToken.create({ expires_in: 60, grants: { agent: true } })
}

Connect this function to your application's authenticated /api/my-cartesia-agent-token route. The route returns { token } and should use the caller's authorized provider account. The agent grant is capability-scoped; the current token API does not expose an individual-agent allowlist. Do not turn it into an anonymous owner-funded proxy. Keep API keys server-only; browser tokens stay in memory and must be redacted from WebSocket URL logs.

Copyable React component

tsx
import { useEffect, useMemo, useState } from 'react'
import { Orb } from 'orb-ui'
import { createCartesiaAdapter } from 'orb-ui/adapters'

export function CartesiaVoiceUI({ agentId }: { agentId: string }) {
  const [caption, setCaption] = useState('')
  const adapter = useMemo(
    () =>
      createCartesiaAdapter({
        agentId,
        inputFormat: 'pcm_24000',
        getAccessToken: async (signal) => {
          const response = await fetch('/api/my-cartesia-agent-token', {
            method: 'POST',
            signal,
          })
          if (!response.ok) throw new Error('Could not authorize your Cartesia agent')
          const data: { token: string } = await response.json()
          return data.token
        },
        onEvent: (event) => {
          if (event.type === 'turn_ended') setCaption(event.text ?? '')
        },
      }),
    [agentId],
  )
  useEffect(
    () => () => {
      void adapter.stop().catch(() => undefined)
    },
    [adapter],
  )

  return (
    <section>
      <Orb adapter={adapter} theme="circle" aria-label="Start or stop Cartesia agent" />
      <p aria-live="polite">{caption}</p>
    </section>
  )
}

Every reconnect fetches a fresh token. adapter.callId contains the session's call record ID after readiness; retrieve transcript/recording data only from your trusted backend. The example's caption displays finalized turn_ended.text; build interim assistant captions by appending turn_output_text_delta.text verbatim through onEvent.

Audio and interruption

The adapter sends session_create immediately after WebSocket open and waits for the server to confirm the requested format. Capture uses a local AudioWorklet. Supported formats are pcm_16000, pcm_24000 (default), and pcm_44100: headerless, mono, signed 16-bit little-endian PCM. Input and output use the same rate. Audio is carried as base64 in JSON text frames.

outputDelivery defaults to speaking_pace, including agents with background audio. as_available is available for configurations without background audio; the adapter buffers playback and rejects a backlog above five seconds. Microphone transport backpressure also closes the session rather than allowing an unlimited queue.

TriggerOrb behavior
Microphone permission, token work, socket/session setupconnecting
User turn startslistening; discard previous playback
User turn ends or assistant turn startsthinking
Agent PCM playback starts in a running audio contextspeaking
Playback drains before the assistant turn endsthinking; output volume resets to zero
Assistant turn ends and playback drainslistening
audio_output_clear interruptionStop current/queued output immediately; listening
Stop or normal agent hangupidle and release microphone/audio/socket
Fatal protocol, permission, timeout, or transport failureerror with cleanup

Recoverable error events have fatal: false; they reach onRecoverableError and onEvent without destroying the session. sendToolResult(toolCallId, result, isError) answers a client_tool_call with expects_response: true. Match the supplied tool-call ID, authorize the action in your application, and keep UTF-8 results within 4096 bytes. The adapter does not execute tools on your behalf.

Browser lifecycle and verification

Start from a user gesture on HTTPS or localhost. Microphone permission is acquired before the token/socket work. Stop cancels pending startup, closes the socket normally, stops tracks, disconnects AudioWorklet/analyser nodes, and flushes output. Repeated start/stop calls are idempotent; stale callbacks from old sockets are ignored. Reconnect explicitly after an error; there is no automatic billable retry loop. The default startup deadline is 30 seconds.

Browser audio suspension or interruption flushes playback and reports an error with session cleanup. Resume by starting again from a user gesture. Queued audio in a suspended context never claims a speaking state. When playback drains between generated chunks, the orb shows thinking with zero output volume; active user speech keeps listening priority.

RMS input/output defaults are conservative and have not been calibrated against live Cartesia voices. Supply directional volume calibration for your audio. For a policy that disallows local blob worklet modules, inject createCaptureNode(context) from your existing approved module.

Synthetic tests cover session ordering, audio encoding, playback drain, interruption, recoverable and fatal errors, tool limits, transport backpressure, denied permission, cancellation, and reconnect. Live provider credential tests have not been run. The local simulation validates the visual interaction; confirm your agent behavior and real-device audio before production use.