Agents
LLM tab
The LLM tab picks the model that decides what the agent says and calls its tools, and sets how varied and how long its replies are.
Fields
All fields live under llm in the agent object.
| Field | Type | Default | Allowed | What it does |
|---|---|---|---|---|
provider | string | gemini | gemini | The LLM provider. Gemini is the only one today. |
model | string | gemini-3.5-flash-lite | a model id from GET /api/v1/catalog | The Gemini model used for replies, the hang-up check and post-call analytics. Not checked when you save, so copy the id exactly. |
temperature | number | 0.4 | 0.0 to 2.0 | Lower is more predictable; higher is more varied. |
max_tokens | integer | 300 | 16 to 4096 | The most tokens one reply can use. A reply that hits the limit is cut short. |
{
"llm": {
"provider": "gemini",
"model": "gemini-3.5-flash-lite",
"temperature": 0.4,
"max_tokens": 300
}
}
Choosing a model
GET /api/v1/catalog returns the models you can pick, in llm_models. The list comes from Gemini and includes its text chat models (image, TTS, embedding and live models are left out).
curl https://voxa.abhinavyadav.in/api/v1/catalog -H "X-API-Key: $VOXA_API_KEY"
{
"llm_models": [{"id": "gemini-3.5-flash-lite", "label": "gemini-3.5-flash-lite"}, "..."],
"tts_models": [{"id": "sonic-3.6", "label": "Sonic 3.6 (latest)"}, "..."],
"stt_models": [{"id": "ink-2", "label": "Ink 2 (en, hi, fr, ja, es; detects language and turn end)", "provider": "cartesia"}, {"id": "nova-3", "label": "Nova 3", "provider": "deepgram"}],
"stt_providers": [{"id": "cartesia", "label": "Cartesia"}, {"id": "deepgram", "label": "Deepgram"}],
"deepgram_languages": [{"id": "hi", "label": "Hindi"}, {"id": "multi", "label": "Multilingual (switches between languages)"}, "..."],
"languages": [{"id": "hi", "label": "Hindi"}, {"id": "en", "label": "English"}]
}
On a phone call, the time from the caller finishing to the agent speaking matters more than anything else. Smaller “flash” models answer faster. Check first_audio_ms on your calls (see The call object) when you compare models.
How the model is used during a call
- Streaming. The reply streams in. Each finished sentence is sent to the voice at once, so the caller hears the first sentence while the model is still writing the rest.
- Tools. The model sees your tools and the built-in
end_call. When it calls a tool, Voxa runs the HTTP request and gives the result back to the model, which then continues. A single caller turn allows up to four model requests, so the agent can chain a few tool calls. - History. The model sees the whole conversation so far. If the caller interrupted, it sees only the part of its reply the caller actually heard, followed by
…. - Usage. Every request’s input and output tokens are added to the call’s
usage(llm_input_tokens,llm_output_tokens,llm_calls), including the hang-up check. Post-call analytics is not counted there.
Tips
- Keep
max_tokenslow (200 to 400). Voice replies should be short, and the voice rules already ask for one to three sentences. - Keep
temperaturebetween0.2and0.6for agents that must follow a script or collect data.