Skip to content
VoxaDocs
Navigation
Open console →

Agents

LLM tab

The LLM tab picks the model that decides what the agent says and calls its tools, and sets how varied and how long its replies are.

Fields

All fields live under llm in the agent object.

FieldTypeDefaultAllowedWhat it does
providerstringgeminigeminiThe LLM provider. Gemini is the only one today.
modelstringgemini-3.5-flash-litea model id from GET /api/v1/catalogThe Gemini model used for replies, the hang-up check and post-call analytics. Not checked when you save, so copy the id exactly.
temperaturenumber0.40.0 to 2.0Lower is more predictable; higher is more varied.
max_tokensinteger30016 to 4096The most tokens one reply can use. A reply that hits the limit is cut short.
{
  "llm": {
    "provider": "gemini",
    "model": "gemini-3.5-flash-lite",
    "temperature": 0.4,
    "max_tokens": 300
  }
}

Choosing a model

GET /api/v1/catalog returns the models you can pick, in llm_models. The list comes from Gemini and includes its text chat models (image, TTS, embedding and live models are left out).

curl https://voxa.abhinavyadav.in/api/v1/catalog -H "X-API-Key: $VOXA_API_KEY"
{
  "llm_models": [{"id": "gemini-3.5-flash-lite", "label": "gemini-3.5-flash-lite"}, "..."],
  "tts_models": [{"id": "sonic-3.6", "label": "Sonic 3.6 (latest)"}, "..."],
  "stt_models": [{"id": "ink-2", "label": "Ink 2 (en, hi, fr, ja, es; detects language and turn end)", "provider": "cartesia"}, {"id": "nova-3", "label": "Nova 3", "provider": "deepgram"}],
  "stt_providers": [{"id": "cartesia", "label": "Cartesia"}, {"id": "deepgram", "label": "Deepgram"}],
  "deepgram_languages": [{"id": "hi", "label": "Hindi"}, {"id": "multi", "label": "Multilingual (switches between languages)"}, "..."],
  "languages": [{"id": "hi", "label": "Hindi"}, {"id": "en", "label": "English"}]
}

On a phone call, the time from the caller finishing to the agent speaking matters more than anything else. Smaller “flash” models answer faster. Check first_audio_ms on your calls (see The call object) when you compare models.

How the model is used during a call

  • Streaming. The reply streams in. Each finished sentence is sent to the voice at once, so the caller hears the first sentence while the model is still writing the rest.
  • Tools. The model sees your tools and the built-in end_call. When it calls a tool, Voxa runs the HTTP request and gives the result back to the model, which then continues. A single caller turn allows up to four model requests, so the agent can chain a few tool calls.
  • History. The model sees the whole conversation so far. If the caller interrupted, it sees only the part of its reply the caller actually heard, followed by ….
  • Usage. Every request’s input and output tokens are added to the call’s usage (llm_input_tokens, llm_output_tokens, llm_calls), including the hang-up check. Post-call analytics is not counted there.

Tips

  • Keep max_tokens low (200 to 400). Voice replies should be short, and the voice rules already ask for one to three sentences.
  • Keep temperature between 0.2 and 0.6 for agents that must follow a script or collect data.
Esc