Agents
Voice tab
The Voice tab sets how the agent sounds: the text-to-speech model, the voice, the language it speaks and how fast it talks.
Fields
All fields live under voice in the agent object.
| Field | Type | Default | Allowed | What it does |
|---|---|---|---|---|
provider | string | cartesia | cartesia | The text-to-speech provider. |
model | string | sonic-3.6 | sonic-3.6, sonic-3.5, sonic-3 | The Cartesia Sonic model. |
voice_id | string | "" | a voice id from GET /api/v1/voices | The voice. If empty, the first voice Cartesia lists for language is used. |
voice_name | string | "" | any | A label for the console. It does not change the sound. |
language | string | hi | a language id from GET /api/v1/catalog (hi, en by default) | The language the voice speaks. hi also tells the agent to speak Hinglish in Roman script. |
speed | number | 1.0 | 0.6 to 1.5 | Speaking rate. 1.0 is normal. |
fallbacks | array | [] | up to 3 | Backup voices, tried in order when the voice fails. See Fallback voices. |
{
"voice": {
"provider": "cartesia",
"model": "sonic-3.6",
"voice_id": "a0e99841-438c-4a64-b679-ae501e7d6091",
"voice_name": "Riya",
"language": "hi",
"speed": 1.05
}
}
Finding a voice
GET /api/v1/voices lists Cartesia voices. Filter with language, gender (as Cartesia names it) and q (a search term). Results are cached for 10 minutes. The ids and names below are examples.
curl "https://voxa.abhinavyadav.in/api/v1/voices?language=hi" -H "X-API-Key: $VOXA_API_KEY"
[
{"id": "a0e99841-438c-4a64-b679-ae501e7d6091", "name": "Riya", "description": "Warm, friendly young woman", "gender": "feminine", "language": "hi"}
]
To hear a voice before you pick it, fetch an MP3 preview:
curl -o preview.mp3 \
"https://voxa.abhinavyadav.in/api/v1/voices/a0e99841-438c-4a64-b679-ae501e7d6091/preview?language=hi&model=sonic-3.6&text=Namaste%2C%20main%20Riya%20hoon" \
-H "X-API-Key: $VOXA_API_KEY"
text is up to 300 characters. In the console, the Voice tab shows the voices for the chosen language with a play button on each.
Language and Hinglish
When language is hi:
- the agent’s prompt gets a rule to speak natural Hinglish written in Roman script, never Devanagari, matching how the caller mixes Hindi and English;
- text in Roman script is read with Indian English text normalisation (
en-IN), and text in Devanagari with Hindi normalisation (hi-IN).
Use en for English-only agents. The language list comes from the platform’s settings (languages in GET /api/v1/catalog).
The voice’s
languageand the transcriber’s language are separate settings. For a Hindi agent on Deepgram, set bothvoice.languageandtranscriber.language. See Transcriber tab.
Fallback voices
voice.fallbacks lists up to three backup voices. When the voice fails to connect at the start of a call, or reports an error or drops during the call, the next one takes over for the rest of the call. A reply that was being spoken when the voice failed is spoken again by the backup, so the caller doesn’t miss it.
| Field | Type | Default | What it does |
|---|---|---|---|
provider | string | cartesia | The text-to-speech provider. |
model | string | sonic-3.6 | The Sonic model. |
voice_id | string | required | The backup voice. |
voice_name | string | "" | A label for the console. |
language | string | "" | The language it speaks. Empty means the same as the primary voice. |
A backup uses the primary voice’s speed. The same model and voice can’t appear twice (primary included); saving returns 422.
{
"voice": {
"provider": "cartesia",
"model": "sonic-3.6",
"voice_id": "a0e99841-438c-4a64-b679-ae501e7d6091",
"voice_name": "Riya",
"language": "hi",
"speed": 1.0,
"fallbacks": [
{"provider": "cartesia", "model": "sonic-3", "voice_id": "f9836c6e-a0bd-460e-9d3c-f7299fa60f94", "voice_name": "Meera", "language": ""}
]
}
}
Each switch is a fallback.used event on the call’s timeline, with the reason. Without fallbacks, a failing voice behaves as before: the call shows an error event, and a voice that can’t connect at all fails the call. See Webhook events.
Usage
Every character sent to the voice is counted in the call’s usage.tts_characters, including the welcome message, tool pre_call_messages, the “are you still there?” message and the hang-up message, and replies a fallback voice speaks again.