Voice
Voice is how your AI speaks and listens on a real phone call or in a browser tab. The same workflow that runs your SMS and web chats runs your voice calls — but instead of reading text, the caller hears a synthetic voice, and instead of typing, they speak. Gravity Rail's voice stack handles speech recognition, low-latency AI responses, natural barge-in, and call routing end-to-end.
This guide covers everything you need to know as a workspace admin to enable voice, pick the right model and voice, and keep calls flowing cleanly.
Heads up — see also Phone & Voice. Phone & Voice is the quick-start for phone numbers, SMS, and routing rules. This guide is the deep dive on the voice side: models, speech recognition, interruption behavior, and troubleshooting.
What Voice Does
When a caller reaches your number (or clicks the mic on a site), Gravity Rail:
- Answers the call — over the phone network for phone calls, or directly in the browser for web voice.
- Plays a greeting (optional) — a TTS greeting with optional consent acknowledgment.
- Listens — streaming audio to a speech-to-text (STT) model or a native realtime model.
- Thinks — runs the current workflow task through the assigned AI model.
- Speaks — streams synthesized audio back to the caller in near-real time.
- Handles interruption — if the caller starts talking while the AI is speaking, the AI stops and listens.
- Persists the conversation — full transcripts saved to the chat, with a summary generated when the call ends.
The same Workflow assistant can talk on the phone, on a website, or over WhatsApp. Its name, identity, models, and voice are versioned with the Workflow, while the phone number or site decides which Workflow handles the conversation.
Filling Silence: Thinking Sounds & Utterances
When the agent fills out a form, verifies a caller's identity, or looks something up, it's making tool calls — and tool calls take time. On a phone call that processing reads as dead air: the caller hears nothing, wonders if the call dropped, and may hang up. Long enough silences can even trip the carrier's silence timeout and end the call for you.
You can fill that silence two ways, both configured in the Workflow draft's Assistant tab:
1. A sound. Set Thinking Sound to typing or strumming and the caller hears a soft ambient loop (keyboard taps or a gentle strum) while the agent works. It's the audio equivalent of "please hold" — unobtrusive, but unmistakably something is happening.
2. Utterances. Set Thinking Sound to speech and the assistant speaks a short filler phrase instead — "Just bear with me a moment.", "Let me pull that up." — rendered in the configured voice, so it sounds like the same person, not a recording. Provide your own phrases in Thinking Speech Phrases; they rotate round-robin within a session so a caller who triggers several lookups doesn't hear the same line twice in a row.
| Setting | What it does |
|---|---|
Thinking Sound (thinkingSound) | none (default), typing, strumming, or speech. |
Thinking Speech Phrases (thinkingSpeechPhrases) | The phrases spoken when the sound is speech, rotated round-robin per session. |
Thinking Sound Volume (thinkingSoundVolume) | 0–100. Unset means full volume. |
Thinking Sound Initial Delay (thinkingSoundInitialDelay) | Milliseconds of silence before the sound starts (default 500ms) — quick tool calls finish before the sound ever plays, so you don't get a stray half-second of typing on every fast lookup. |
A few things to know:
- Barge-in still works. If the caller speaks while a thinking sound or filler phrase is playing, it stops immediately — the sound never talks over them.
- Utterances match the voice. Filler phrases are synthesized through the same TTS model and voice as the rest of the call, and cached after the first render so they play instantly on later tool calls.
- Tune the delay to your workflow. If your agent's lookups routinely take several seconds (provider searches, EMR queries), keep the default 500ms delay. If callers are noticing the gap before the sound kicks in, lower it.
Test a thinking sound when your Workflow performs lookups. During a synthetic test call, check that callers can hear when the assistant is working and can interrupt it.
Choosing a Voice Pipeline
Every call does three things — Listening, Thinking, Speaking. The Voice pipeline field decides whether one model does all three or three models do one each. It is the first control in the Voice section, and the names on it are the names used below:
| Voice pipeline | How it works | When to use |
|---|---|---|
| Native voice | A single model (OpenAI GPT Realtime, Gemini Live, Grok Voice) listens, thinks and speaks. On the form the three stage labels stay, joined by a brace to that one model. | Lowest latency; natural turn-taking; best for voice-first products. |
| Speech-to-text | Separate STT (Deepgram / OpenAI / xAI), any chat LLM, and separate TTS (Polly / OpenAI / Google / Deepgram / xAI) — one row per stage. | Lets you mix and match — e.g. Claude for reasoning with xAI for voice. More tuning knobs. |
| Off | The agent takes no calls. | Text-only agents. |
Native voice is simpler to configure but offers fewer voice choices. Speech-to-text gives you access to every voice Gravity Rail supports (100+ voices across 5 TTS vendors) with any chat model you like.
Your text model does not run on calls. The Thinking model in the Text section drives web chat, SMS and email only. Under Native voice, the one model you pick there does the thinking on calls — the form says so under the selector.
Switching is safe. Moving between pipelines — or turning voice Off and back on — keeps the settings for the side you left. Pick Native voice, change your mind, and your speech-to-text stack comes back exactly as you built it.
The Options block groups settings by the stage they affect:
| Group | Settings |
|---|---|
| Listening | Language hints, keyterms, end-of-turn thresholds, silence timeout |
| Thinking | Reasoning settings, thinking sound, volume, delay, and speech phrases |
| Speaking | Interruption sound and volume |
| Turn detection | Use Edit to configure turn detection for native OpenAI engines |
Under Native voice there are no stage headings: one model does all three, so the options are one flat list.
Native Grok Voice engines also get a Reasoning option: Default (on), On or Off. Off answers faster but makes tool calls less reliable.
Enabling Voice on a Workspace
Voice is enabled per phone number and per site. There's no workspace-wide "voice on/off" switch — if you have a phone number with Enable Voice on and a workflow connected, voice is live.
Prerequisites
- A Workflow with its Assistant voice configured. You can configure it directly or use Copy from… to start from an Agent.
- Either a Phone Number (for PSTN calls) or a Site with voice enabled (for browser calls).
If your org doesn't have any phone numbers yet, an org owner needs to purchase one first from the Organization's Phone Numbers tab. See Phone & Voice for the walkthrough.
Configuring Inbound Calls
Inbound calls are the common case: someone dials your number and your AI picks up.
1. Connect a phone number to a workflow
- Go to Channels → Phone Numbers.
- Open a phone number (or add one).
- On the number's Settings tab, check Enable voice in Channels.
- In the Voice routing rules, edit the AI answers rule and choose the Workflow whose assistant handles the call.
Save. The first inbound call will be answered by the active Workflow revision's assistant within a few seconds.
2. Decide where calls go
A newly created number has an AI rule for known Members and an anonymous-caller rule that turns unknown callers away with a message. To do anything else — forward after hours, ring your operators, take a voicemail, play a recorded message — add Routing Rules to the number's Voice list on its Settings tab.
Rules run in order and the first match wins. Each one carries a time window (Always, During business hours, After business hours, or a Custom schedule), a caller match (Anonymous, Member, or both), and one action:
| Action | Behavior |
|---|---|
| AI answers | The AI picks up, on the workflow you name or on a plain agent conversation. |
| This number's owner | The number's owner answers: a person is rung in the browser, then takes a voicemail; an Agent answers with realtime voice when available, otherwise the caller goes to voicemail. Offered while the number has an owner. |
| A person answers | Rings an operator group, a named member, the caller's primary contact, or an external phone number. 5–120 s, 30 s by default. |
| Voicemail | Records and transcribes a message after an optional greeting. |
| Play a message | Speaks the message, then hangs up. |
| Ignore | Hangs up silently. Anonymous callers only. |
| Sign up | Invites the caller to join as a Member on a Role you pick. Anonymous callers only. |
Only A person answers can fail to settle the call. If the ring times out or the target can't be reached, the next matching rule runs — that is how you express "ring the desk, and take a voicemail if nobody picks up". Everything else ends the call's trip through the list.
So the old always-on line is a single Always · any caller · AI answers rule, and "AI during business hours, forward after" is two rules: During business hours · any caller · AI answers, then After business hours · any caller · A person answers pointed at the number to forward to.
When the number has an owner and no rule answers a known caller — nothing matches, or every A person answers rule went unanswered — the owner answers. Unknown callers who reach the end of the list are turned away, as on a number with no owner.
Business hours are configured in Settings → Workspace Settings. Rules with a business-hours window use that schedule, in the workspace timezone.
3. Add a greeting (and optional consent)
Some workflows — especially in healthcare or regulated industries — need to announce that the call is recorded or that an AI is answering. Configure this in the number's Channels settings, beside Enable voice:
- Greeting message (
voice_greeting_message) — played to the caller before the AI answers and while a person is being rung. For example: "Thanks for calling Acme Clinic. This call may be recorded and is being answered by our AI assistant." - Require consent (
voice_require_consent) — shown once there is a greeting. When on, the greeting plays inside a gather element. The caller must press1or say "yes" to continue. No response → a polite "No response received. Goodbye." and the call hangs up.
The greeting uses your configured TTS voice, so it sounds like the same assistant that'll take the call.
4. Anonymous callers
A caller Gravity Rail can't match to a workspace member is anonymous, and what happens to them is a routing rule like everything else. Rules that match anonymous callers belong at the top of the list, because whether a caller is admitted is settled before where they go.
- Admit them — edit the AI answers, This number's owner, or A person answers rule to match Anyone or Anonymous, and choose its Role for new callers. An unknown caller is admitted as a new Member on that Role. An A person answers rule that forwards to an external phone number makes no one a Member, so it takes no Role. A rule that admits anonymous callers without a Role shows Role required in the list.
- Invite them to join —
Always · Anonymous · Sign up, with the Role new contacts are created under. On voice this runs the consent flow and texts them a signup link. - Turn them away —
Always · Anonymous · Play a message, with the text the caller hears before the line hangs up. Use Ignore instead to hang up without saying anything. This is the right choice for private, member-only voice lines, and it is what a new number starts with.
Configuring Outbound Calls
Outbound (the AI calls out) is driven by the Phone Call action on a workflow. There is no "dial this number" button — outbound calls are always triggered by automation.
- Create a Phone Call action that places a call from one of your phone numbers.
- Trigger the action from an event rule, a schedule, or a workflow step.
- The outbound call uses the same active Workflow revision and versioned Assistant settings as inbound calls; the only difference is who dialed whom.
Typical triggers
- An appointment is 24 hours away → place a reminder call.
- A form field changes state (e.g. lab result marked abnormal) → call the patient.
- A scheduled campaign runs → dial a list of members sequentially.
Outbound calls are billed per-minute just like inbound. See Analytics & Usage Reports for the per-number and per-agent cost breakdown.
Voice Models & TTS
Voices are picked per Workflow revision. Every voice maps to a TTS model (the engine that synthesizes speech) and a voice name (the specific speaker).
Open the Workflow draft's Assistant tab. You'll see voices compatible with the selected model. Listen to previews and pick one, or use Copy from… to copy an Agent's settings as a starting point.
Voice providers
Gravity Rail supports five TTS providers out of the box. Each has trade-offs:
| Provider | Voices | Strengths | Notes |
|---|---|---|---|
| AWS Polly | ~60 voices across generative and neural engines (Ruth, Matthew, Joanna, etc.) | Reliable AWS infrastructure; predictable pricing; strong multi-language. | Uses polly-neural or polly-generative depending on voice. |
| OpenAI TTS | 11 voices (Alloy, Ash, Ballad, Coral, Echo, Sage, Shimmer, Verse, Marin, Nova, Onyx) | Very natural conversational tone; low latency. | Same voices also available on OpenAI Realtime. |
| Google Chirp | Puck, Kore, Charon, Fenrir, Aoede, Leda, Orus, Zephyr | High-quality; good multilingual support. | Used via Vertex AI or AI Studio. |
| Deepgram Aura 2 | Thalia, Asteria, Luna, Arcas, Perseus, and others | Low-latency; designed for real-time. | Good fit for pipeline mode where STT is also Deepgram. |
| xAI | Eve, Ara, Rex, Sal, Leo | Paired with Grok Voice Think Fast (1 or 2). | Think Fast 2 improves intelligence, phone-line transcription, and time-to-first-audio vs 1. |
The Workflow Assistant editor shows only voices that work with your selected TTS model, so you can't pick an incompatible combination.
Native realtime voices
If the Workflow assistant uses a realtime model (OpenAI GPT Realtime, Gemini Live, Grok Voice Think Fast), the voice is baked into the model configuration, not picked from the TTS catalog:
- OpenAI GPT Realtime: Alloy, Ash, Ballad, Coral, Sage, Verse
- Gemini Live: Puck, Kore, Charon, Fenrir (and more)
- Grok Voice Think Fast 1 / 2: Eve, Ara, Rex, Sal, Leo
Native realtime models emit audio directly, without a separate TTS stage. Choose an available voice in the Workflow's Assistant settings and test it with your call setup.
Switching voices mid-conversation
You generally can't change voices within a single call — the voice is locked to the Workflow revision pinned by the Chat or Assignment when the session starts. For different use cases, configure and publish the appropriate assistant settings on each Workflow.
Speech Recognition (STT)
When you use a pipeline model (anything except a native realtime model), speech recognition happens separately from the LLM. Gravity Rail supports three STT providers:
| Provider | Model | Strengths |
|---|---|---|
| Deepgram | nova-3, flux, Flux Multilingual | Fast streaming; strong on medical/technical vocabulary; keyterm prompting; Flux Multilingual can report the detected language. |
| OpenAI | Live Transcribe | Free-form transcription context and native G.711 input. |
| xAI | xAI STT | Low-cost streaming transcription with automatic recognition across its supported languages. Streaming events do not report which language was detected. |
Deepgram is the default and what we recommend for most workspaces. Its flux model is purpose-built for low-latency streaming with server-side end-of-turn (EOT) detection.
Keyterm prompting
If your workspace handles unusual vocabulary — drug names, procedure codes, product SKUs — you can give Deepgram a list of keyterms to bias transcription toward. Two sources feed the keyterm list:
- Pronunciation terms configured on your workspace (account or org level).
- Manual
keytermsset in the agent's STT config.
Both lists are merged and passed to Deepgram. Keyterm prompting is free — no reason not to use it if you have a known vocabulary.
Never put patient identifiers in this list. It is workspace configuration, is sent to the speech provider on every session, and applies to every caller. Phone numbers, dates, record numbers, and email addresses are rejected on save; everything else still needs a human check before the list goes live.
Language detection
Deepgram Flux Multilingual supports automatic language detection — useful when callers might speak any of its supported languages. Enable languageDetection in the Workflow assistant's STT settings. By default, Gravity Rail detects the first reliable language and locks to it for the rest of the call. This improves recognition stability but does not switch again if the caller changes languages later.
If the same call must support ongoing language changes, also set detectThenLock to false. In that mode, each reliable detection can update the prompt locale used for inference. Use it deliberately: continuous detection is more flexible, but a mistaken detection can also change the language used for the agent's next response. Language detection does not select or change the TTS voice; voice selection is a separate setting or Custom Voice action.
xAI STT also recognizes supported languages automatically, but its streaming events do not report which language it detected. It therefore cannot drive per-turn prompt-locale or voice switching. Use cloud Deepgram Flux Multilingual when that downstream signal is required.
For native realtime models, language detection is handled by the model itself — you just instruct the agent to respond in the caller's language.
Barge-in & Interruption
Natural conversation means people interrupt each other. The voice pipeline is built around barge-in — the moment the caller starts speaking, the AI stops mid-sentence and listens.
How it works
- The AI is speaking — audio is streaming to the caller.
- The caller starts talking, and voice activity detection picks it up.
- The AI stops mid-sentence, and the conversation records exactly how much the caller actually heard — so on the next turn, the AI knows the caller didn't hear the rest and won't act as if they did.
- A new turn starts; the AI listens.
This works the same whether you're on a pipeline model or a native realtime model — each provider detects interruption its own way, but the caller experience is consistent.
When barge-in feels wrong
- AI gets cut off by its own echo (on phone): usually a carrier-side audio issue on the line, not something you can fix in settings. If it happens consistently on calls to your number, contact support.
- AI doesn't stop when the caller speaks: voice detection sensitivity is too low. If you're on native realtime, the
turnDetectionconfig on the Workflow revision's Assistant controls this. For Deepgram, tuneeot_thresholdorvad_silence_threshold. - AI thinks it was interrupted when it wasn't (phantom barge-in): voice detection is too sensitive — background noise is triggering it. Raise the VAD threshold or, on a pipeline model, bump
min_speech_duration_ms.
Web Voice
Voice on a site works like phone voice, but the browser is the carrier. Enable it by:
- Open a site under Channels → Sites.
- Enable voice in the site's settings.
- Attach a Workflow whose revision has a voice-enabled Assistant.
Visitors click a microphone button, grant mic permission, and talk — no phone number needed. Browser voice has no per-minute phone charges, so it's a great fit for self-service portals.
Because the browser carries richer audio than the phone network does, voice quality is usually noticeably better on web than on the phone.
Call Summaries & Transcripts
Every call generates a full transcript — saved alongside the chat in the same conversation view your team already uses for SMS and email. Open the chat, and you'll see:
- Each turn with speaker attribution.
- Tool calls the AI made during the conversation.
- Any data the AI collected in forms.
When a call ends, a short AI-written summary is generated automatically and saved to the chat record. You don't have to wait — the summary shows up a few seconds after the call hangs up.
Audio recording is controlled by a workspace-level setting (enable_audio_recording on Workspace Settings) and is off by default. When enabled, it applies workspace-wide — you can't opt in or out per phone number today. Per-number granularity is on the roadmap. Transcripts are always persisted regardless of the audio-recording toggle.
To turn recording on and play calls back inside chat (including PHI safeguards), see Call Recordings.
Troubleshooting
"The caller said something but the AI didn't respond."
- Check the chat in Gravity Rail. If the transcript is missing or garbled, it's a speech recognition issue — consider switching STT provider or adding keyterms.
- Check the call's status in the chat — if the connection dropped mid-call, the call will show an error status. If you see repeated dropped calls, contact support.
- Voice detection threshold too high: the caller's speech isn't crossing the detection threshold. Lower
vad_threshold(xAI STT) oreot_threshold(Deepgram Flux).
"The AI's voice sounds robotic or choppy."
- Network latency: the most common cause. Pipeline mode is more sensitive than native realtime because it adds STT → LLM → TTS hops. Switch to a native realtime model if you can't fix the network.
- Phone calls specifically sound worse than web: some quality difference is expected — the phone network carries lower-fidelity audio than the browser. But if phone calls sound broken (clicks, garbled stretches), contact support.
- TTS provider delay: compare synthetic calls using another supported voice model. If delays persist, contact support with the affected Chat IDs.
"The AI keeps interrupting itself / talking over the caller."
- Barge-in is too aggressive — see the Barge-in section for tuning.
- On OpenAI Realtime, check
turn_detection.silence_duration_ms— the default is ~500ms, which can misfire in noisy environments. Raise it to 800–1000ms for phone calls.
"The caller heard a short apology and was asked to hang up and try again."
The greeting audio never reached the caller within a few seconds, so the line spoke a fallback instead of sitting in silence. Ask the caller to hang up and call again. If this keeps happening, contact support with the call's chat link.
"The AI answered but hung up immediately."
- Configuration problem: usually an invalid model selection, or the AI answers rule that matched names no Workflow. Double-check the active Workflow revision's Assistant settings and the Workflow on that rule; if both look right, contact support.
- Consent flow timed out: if
voice_require_consentis on and the caller didn't respond, the flow hangs up deliberately with "No response received. Goodbye." - An after-hours rule matched: a Play a message rule matched the call, so the recorded message played and the call ended — this is the intended behavior. Open the number's Voice rules and use Test with the time in question to see which rule ran.
"Calls work on phone but not on the web site."
- Mic permission denied in the browser — visitors need to grant microphone access. The site widget will prompt, but some browsers block it by default.
- Stale page — voice sessions authenticate with a short-lived token; if the page has been open a long time before the mic button is clicked, the connection can be refused. Refresh the page and try again.
- Custom domain proxy/CDN — if your site runs on a custom domain behind a proxy or CDN, it must allow real-time (WebSocket) connections through. Your web team or the proxy's documentation can confirm this.
"The voice changed partway through the call."
- A Custom Voice action changed it — an agent with the Custom Voice ability can change a compatible xAI TTS voice during a call. Other TTS providers apply the new preference on the next call. Language detection does not trigger this action.
- The Chat uses an older Workflow revision — active calls stay pinned to the revision they started with. Publishing new voice settings affects newly started work, not an in-progress call.
"The caller was charged for a call that never actually connected."
This shouldn't happen — Gravity Rail only logs usage once a call actually connects. If you see a billing entry for a call that never connected, contact support with the call's chat link and we'll investigate against the carrier's records.
"Transcripts are missing or partial."
- Call dropped before finalization: if the connection closed abnormally, the last few seconds of audio may not have been transcribed. Check the chat — a message with an error status marks where the call dropped.
- STT provider outage: rare, but if Deepgram, OpenAI, or xAI had a regional incident, transcription may have failed for a window. Check your provider's status page.
Common Setups
An AI answers rule, or an A person answers rule that rings a person, that matches any caller also names its Role for new callers; the setups below leave it out.
Clinic reception line
- Voice rules:
Always · Anonymous · Sign up(as Patient), thenAlways · Member · AI answers - Consent:
voice_require_consent = true, greeting announces recording + AI - Agent: native realtime model (OpenAI GPT Realtime) for lowest latency
- Voice: Coral or Sage (warm, professional)
- Abilities: Calendar Booking, Phone Call Tools (configure forwarding to the reception team)
- Anonymous callers: the signup rule asks for consent and texts them a signup link
Use the clinic's established urgent-care path for urgent requests. This setup handles reception; it does not assess clinical urgency. Test with synthetic data, and confirm PHI authorization before using patient information.
Outbound appointment reminders
- Workflow: reminder flow with Phone Call action
- Schedule: event rule 24h before appointment
- Voice: xAI Leo (warm, familiar)
- STT: Deepgram with keyterms for your clinic's vocabulary
- Outcome tracking: workflow collects confirmation into a form field
Multilingual inbound support
- STT: cloud Deepgram Flux Multilingual with
languageDetection: trueanddetectThenLock: falsefor ongoing language changes - Voice: Google Chirp (strong multilingual)
- Agent instructions: "Respond in the language the caller uses."
- Fallback: because continuous detection is enabled, the caller can use a full phrase such as "Please continue this call in English" to provide a reliable new language sample.
After-hours voicemail
- Voice rules:
During business hours · any caller · AI answers, thenAlways · any caller · Voicemail - Business hours: 9am–5pm weekdays
- Voicemail transcription: auto-transcribed and saved to the chat
- Notification: event rule on new voicemail → Slack notification to on-call staff
Tips
- Always test with a real phone — the browser mic doesn't exhibit the same quirks as a cellular call. Dial your number from a real phone before you go live.
- Pick the shortest plausible greeting — every second of greeting is a second before the caller can talk. Consent prompts in particular should be as short as legally allowed.
- Use pronunciation terms if you have non-standard vocabulary — it's the fastest way to improve transcription quality.
- Keep instruction prompts tight — realtime models are especially sensitive to prompt length; long system prompts increase time-to-first-audio.
- Monitor the Analytics dashboard for call duration, token usage, and failure rate trends. Sudden changes usually signal an upstream provider issue.
Related
- Phone & Voice — Phone number setup, SMS, and routing rules
- Phone Call Tools — AI abilities for hanging up and forwarding calls
- Phone Call Action — Trigger outbound calls from automation
- Sites — Enable browser-based voice on a site
- Agents — Reusable settings you can copy into a Workflow assistant
- Voice Receptionist Tutorial — End-to-end build of a phone-based receptionist
- Analytics & Usage Reports — Per-call and per-number billing data
- Workspace Settings — Business hours for routing-rule time windows