Agents
Transcriber tab
The Transcriber tab picks the speech-to-text service that turns the caller's speech into text and decides when the caller has finished speaking.
Fields
All fields live under transcriber in the agent object. Some fields apply to one provider only.
| Field | Type | Default | Allowed | Provider | What it does |
|---|---|---|---|---|---|
provider | string | cartesia | cartesia, deepgram | both | The speech-to-text service. |
model | string | ink-2 | ink-2 (Cartesia), nova-3 (Deepgram) | both | The model. If it doesn’t match the provider, Voxa replaces it: ink-2 for Cartesia, nova-3 for Deepgram. |
turn_detection | string | balanced | responsive, balanced, patient | Cartesia | How quickly the agent answers once the caller pauses. |
language | string | hi | hi, multi, en-IN, en | Deepgram | The language the caller speaks. Cartesia ignores it and detects the language itself. |
endpointing_ms | integer | 300 | 10 to 3000 | Deepgram | Milliseconds of silence that end the caller’s turn. |
keyterms | array of strings | [] | up to 100 | both | Words the recogniser should expect, such as doctor or product names. Cartesia uses up to 100, Deepgram the first 50. |
Cartesia example
{
"transcriber": {
"provider": "cartesia",
"model": "ink-2",
"turn_detection": "patient",
"keyterms": ["Dr. Kulkarni", "Aundh", "physiotherapy"]
}
}
Deepgram example
{
"transcriber": {
"provider": "deepgram",
"model": "nova-3",
"language": "multi",
"endpointing_ms": 400,
"keyterms": ["Dr. Kulkarni"]
}
}
Cartesia or Deepgram
Cartesia ink-2 | Deepgram nova-3 | |
|---|---|---|
| Language | Detects it on its own (en, hi, fr, ja, es) | You set it: hi, multi (switches between languages), en-IN, en |
| End of turn | A turn-detection model, tuned with turn_detection | A fixed silence, endpointing_ms |
| Good for | Callers who mix Hindi and English; natural pauses | A known single language; strict control of the pause length |
Your workspace may not be allowed to use both. GET /api/v1/providers shows which providers your workspace can use; saving an agent with a provider it can’t use returns 422.
Turn detection
The agent starts its reply when the caller’s turn ends, so this setting trades speed against cutting the caller off.
turn_detection | Behaviour |
|---|---|
responsive | Answers quickly. Good for short answers (“yes”, “no”, a date). |
balanced | The default. |
patient | Waits longer. Good for slow speakers, elderly callers, or callers who think aloud. |
On Deepgram, raise endpointing_ms (for example to 500) if the agent talks over callers who pause mid-sentence, and lower it if replies feel slow.
If a caller pauses mid-sentence and the turn ends early, Voxa joins the two parts into one message for the LLM when the agent has not spoken in between. The transcript still shows them as two turns.
Usage
The seconds of caller audio sent to the transcriber are counted in the call’s usage.stt_seconds.