AI voice agents on live phone calls
Connect a phone number to an AI agent that listens, understands and replies in a natural voice in real time. Answers come from your own documents, and every conversation is transcribed and logged for review.
From "hello" to an answer, in real time
Connect
The caller dials your number. A greeting plays and a two-way audio stream opens over a secure WebSocket.
Listen
Caller audio is buffered in short chunks and sent for speech recognition.
Understand
Whisper turns speech into text.
Reason
Your chosen model streams a reply, using your uploaded documents when they are relevant.
Respond
The reply is spoken in a natural voice and streamed back to the caller.
Record
The conversation, timings and token usage are saved for review and analytics.
Build the voice agent your way
Telephony built in
Twilio Programmable Voice with Media Streams, switched on in code, with nothing to configure in the console.
Choose your AI model
OpenAI GPT-4o or GPT-4o-mini, Anthropic Claude, or AWS Bedrock (Claude, Llama, Mistral), changed with one setting.
Choose your voice
ElevenLabs, AWS Polly neural voices or OpenAI text-to-speech.
Answers from your documents
Upload PDF, Word, TXT or Markdown. Content is indexed with embeddings, and the most relevant passages ground each answer.
Transcripts & analytics
Caller, duration, status, message count and token usage for every call, with each message stored with its timing.
Monitoring dashboard
Browse calls, replay each conversation's messages and see summary statistics in one place.
Web voice chat
A browser voice assistant with spoken replies, live connection status and automatic reconnect.
Streamed replies
The model streams its reply, and audio is sent back to the caller in 20 ms frames.
Swappable providers
Change the speech, model and voice providers independently as your needs change.
Answers from your own documents
Upload the documents your callers ask about. The Voice Connector searches them on every call, so answers reflect your approved content.
- Upload what you already have. PDF, Word, TXT and Markdown files.
- Split and indexed. Content is split into 1,000-character passages and indexed with embeddings for meaning-based search.
- Best passages first. The three most relevant passages ground each answer.
- Graceful fallback. If no document matches, the agent answers from the general model.
Voice AI for regulated teams
Medical information lines
Answer routine product questions from approved documents, around the clock.
Patient support
Consistent, patient answers with a full transcript of every conversation.
Enterprise service desks
Take the first line of calls using your own knowledge base.
The current release handles inbound calls. Call transfer to a person, outbound calling and automatic MICC case creation aren't included yet, so talk to us about your requirements.
Speech, reasoning and voice, in real time
Speech recognition, a language model grounded in your documents and natural text-to-speech work together on every call.
Speech recognition
Caller audio is transcribed with Whisper.
Your choice of model
OpenAI, Anthropic Claude or AWS Bedrock models.
Grounded answers
Your uploaded documents are searched, and the most relevant passages ground each reply.
Natural voices
ElevenLabs, AWS Polly or OpenAI text-to-speech.
Explore the connected suite
Put an AI voice agent on your line
See a live call answered from your own documents, with the transcript on screen as it happens.