Every orb-ui theme consumes the same signal: a stable normalized speech envelope from 0 to 1.
The provider adapter is responsible for producing it. Themes do not know whether a raw measurement
came from RMS audio, an SDK gain event, or a provider-specific volume callback.
Input and output have separate profiles because microphone and playback measurements often use
different sources and distributions. Built-in adapters ship with defaults for both directions when
the provider exposes both. Vapi meters input from its SDK-owned local microphone track.
The 0–1 contract
0 means silence or no meaningful activity.
0.5 means ordinary conversational speech.
1 means a strong, uncommon speech peak.
inputVolume is active while the user is being heard.
outputVolume is active while the assistant is speaking.
The distribution matters as much as the endpoints. Normal speech should move through the middle of
the range instead of remaining near zero or saturating at one. This lets a theme or animation preset
work consistently across providers.
Calibration profile
silenceFloor is the raw value at or below which the output is zero. speechReference is the raw
value mapped to normalized 0.5. speechPeak is the raw value mapped to 1. orb-ui derives the
curve between these anchors, so consumers do not have to tune an unexplained exponent.
riseTimeMs and fallTimeMs are the time required for the normalized envelope to move 90% toward
a new value. Shipped profiles currently use 100ms rise and 400ms fall. Processing is based on
elapsed time rather than callback count, so providers with different meter rates still behave
consistently.
These times belong to signal normalization, not theme choreography. They remove provider jitter and
produce a portable speech envelope. A theme can add its own visual response, easing, state
transition, and autonomous motion after normalization.
Built-in adapters keep every available directional envelope warm while a session is active.
Changing between listening, thinking, and speaking selects a direction for rendering; it does
not reset an envelope to an artificial zero. Disconnecting, stopping, or entering an error state
does.
Override a built-in profile
Most applications should keep the shipped defaults. If a custom audio path has a different raw
distribution, pass partial overrides for each affected direction:
Pass a getter to apply a generated profile without restarting the adapter. Diagnostic callbacks
report each stage:
raw is the provider measurement, mapped is the amplitude-mapped value before temporal
processing, and normalized is the stable envelope emitted as outputVolume.
Generate a profile with the guided runner
The provider QA playground generates amplitude anchors instead of asking you to tune sliders:
- Run
pnpm build and pnpm dev:demo.
- Open
/playground and configure the provider.
- Start a real session and select input or output.
- Capture silence, quiet speech, normal speech, and energetic speech.
- Generate the profile and inspect its mapped distribution.
- Repeat for the other direction and compare themes using the same generated profile.
The runner stores generated profiles locally in that browser and applies them to the active session
through calibration getters. It reports mapped quiet, normal, and energetic medians so a maintainer
can verify that normal speech centers around 0.5 and energetic speech uses the upper range. Reset
restores the shipped baseline for the selected direction.
Use the runner when validating a provider default or a materially different audio pipeline. It is
not intended as an installation step for every application.
Repeatable provider QA
Compare normalized activity, not equal raw numbers. ElevenLabs client 1.9.0 averages voice-band
frequency bins. LiveKit client 2.20.0 computes a different frequency-bin measure with the adapter’s
analyser settings. Pipecat’s browser track meters, OpenAI Realtime, and Gemini Live use waveform
RMS. A raw value of 0.2 therefore does not represent the same audio level in every adapter.
For repeatable microphone tests, feed the same recorded speech at quiet, normal, and louder gains
through the browser’s microphone capture path, with silence between passages. Keep the SDK’s
normal echo cancellation, noise suppression, and automatic gain control settings, and measure the
captured track alongside the SDK’s raw level. Browser gain control can boost quiet speech and
compress the difference between source levels. A digital recording tests that processing path;
physical microphones, room noise, and speaking distance still vary.
Record actual provider output as well as raw, mapped, and normalized samples and state changes.
Check silence, ordinary speech, peaks, pauses, interruption, and restart. Replay the same traces
through proposed profiles and themes so audio differences cannot hide a regression. Validate a
candidate with another voice or phrase before changing a shipped default. Clear saved playground
profiles when checking defaults, and keep credentials out of recordings and trace fixtures.
Calibration API
orb-ui/adapters exports the reusable pieces used by the built-in adapters and playground:
Custom adapters can use createVolumeNormalizer to emit the same stable envelope. The calibration
fitter accepts four raw sample arrays and returns both a generated profile and distribution metrics.
What not to calibrate here
Do not tune provider profiles to make one theme larger, faster, or more dramatic. Provider
calibration standardizes measurement semantics. Theme-specific size, deformation, visual response,
easing, and transition behavior belong to theme configuration. Last modified on September 8, 2026