Calls
Browser calls
A browser call lets you talk to an agent through your microphone, with no phone involved, so you can test a prompt or a tool before the agent calls anyone.
From the console
Open an agent and choose Test in browser. Allow microphone access, fill in values for the agent’s variables if you want to, and start talking. The console shows the live transcript as you speak.
Browser calls are real calls in every other way:
- they get a call record with
channel: "web"anddirection: "web", a transcript, a timeline and a recording; - they hold a call slot while they run;
- they are billed per second like phone calls;
- they fire webhooks.
They differ in that they never wait in the queue (if no slot is free, the call is refused), calling hours don’t apply, and from_number, to_number and {caller_phone} are empty.
From your own code
Any WebSocket client can make a browser call. This is useful for automated tests of an agent.
Connect
wss://voxa.abhinavyadav.in/ws/test/{agent_id}?token=<your API key>
The API key goes in the token query parameter. A key’s workspace is fixed, so no workspace parameter is needed. The key needs calls.place, which every API key has.
The key is in the URL, so it can end up in proxy and server logs. Only connect from code you control, never from a web page you publish.
Protocol
-
Send a start message first, as text, within 10 seconds of connecting.
user_datafills the agent’s variables; send{}if there are none.{"type": "start", "user_data": {"customer_name": "Rohan"}} -
The server answers with the call id:
{"type": "started", "call_id": "4c7d0e5b9a2f4e1c8b3d6a9f0e1d2c3b"} -
Stream audio as binary messages: 16 kHz, 16-bit signed little-endian, mono PCM. Send small chunks (20 to 100 ms) in real time.
-
Receive the agent’s audio as binary messages in the same format, and text messages for the transcript:
Message Meaning {"type": "partial", "text": "..."}What the caller is saying, so far. {"type": "user", "text": "..."}The caller finished a turn. {"type": "assistant", "text": "...", "interrupted": true}The agent finished (or was cut off in) a reply. interruptedappears only when it was cut off.{"type": "clear"}The caller interrupted: stop playing the audio you have buffered. {"type": "error", "text": "..."}Something failed, for example the voice service. -
Hang up by sending
{"type": "hangup"}or closing the socket. The call ends withhangup_by: "caller". When the agent ends the call, the server closes the socket.
After the call, read the record with GET /api/v1/calls/{call_id}.
When the call is refused
The server sends an error message and closes the socket:
| Close code | text |
|---|---|
4401 | sign in to test calls (no valid key or token) |
4402 | your role cannot place calls, this workspace is not active yet, out of credits: add credits in Billing, or all 5 call slots are busy: try again when a call ends |
4403 | <provider> is not available to this workspace |
4404 | agent not found |
Example (Python)
This plays a WAV file (16 kHz, 16-bit, mono) to the agent and prints the transcript. It uses the websockets package.
import asyncio
import json
import os
import wave
import websockets
AGENT_ID = os.environ["AGENT_ID"]
URL = f"wss://voxa.abhinavyadav.in/ws/test/{AGENT_ID}?token={os.environ['VOXA_API_KEY']}"
async def main() -> None:
async with websockets.connect(URL) as ws:
await ws.send(json.dumps({"type": "start", "user_data": {"customer_name": "Rohan"}}))
async def send_audio() -> None:
with wave.open("caller.wav", "rb") as wav:
while chunk := wav.readframes(1600): # 100 ms at 16 kHz
await ws.send(chunk)
await asyncio.sleep(0.1)
await asyncio.sleep(10) # let the agent answer
await ws.send(json.dumps({"type": "hangup"}))
sender = asyncio.create_task(send_audio())
async for message in ws:
if isinstance(message, str):
event = json.loads(message)
if event["type"] in ("started", "user", "assistant", "error"):
print(event)
sender.cancel()
asyncio.run(main())