October 02, 2026 • 9 min read
AI & Automation Specialist
I design AI-powered communication systems. My work focuses on voice agents, WhatsApp chatbots, AI assistants, and workflow automation built primarily on Twilio, n8n, and modern LLMs like OpenAI and Claude. Over the past 7 years, I've shipped 30+ automation projects handling 250k+ monthly interactions.
If you enjoy the content that I make, you can subscribe and receive insightful information through email. No spam is going to be sent, just updates about interesting posts or specialized content that I talk about.
An AI voice agent can answer every frequently asked question a dealership gets, with the right tone and the right hours, and still lose the caller. Not because the answers are wrong. Because the moment the call needs a person, either the agent does not recognize it or the handoff is so clumsy that the customer hangs up first.
In the 80 to 90 percent of agentic voice flows I have audited for clients, that step is missing or bolted on at the end. This post walks through the minimal version that works in production: a Twilio number, ConversationRelay for the audio loop, an ElevenLabs voice, a small model through OpenRouter, and a handoff that the agent itself triggers. It also gives you the bill, because the bill changes the design.
The demo company is a car dealership. The agent has the company identity, opening hours (Monday to Friday 9 to 19, Saturday 9 to 13), new and used inventory, financing rules, test drive rules, workshop and parts, warranties and FAQs. It also has one explicit instruction: when the request requires a person, for example scheduling a test drive, hand off to the sales team.
That scope is deliberate. When a business is starting with voice agents, you do not give the agent a lot of freedom. It answers FAQs, and what it cannot answer goes to sales. That is a common and correct first deployment, and the handoff is the piece that makes it safe.
A real call from the demo, condensed: I ask if they are open, get the hours. I ask about sedans, get «new sedans from 32 million pesos plus selected used». I ask about test drives, get the requirement (book a slot with an advisor, bring a valid license) and the offer to transfer me. I say yes, and two seconds later the call rings in my browser with a summary of what I asked already on screen.
The layers, in order:
/voice endpoint (running locally through a tunnel during development, a VPS in production)./voice returns TwiML that connects the call to ConversationRelay, with the TTS provider, the voice id and the language./relay endpoint. From here on my server only sees text./relay holds the agent: the company context, the prompt, the tools, the turn log. Responses are streamed back to the phone as they are generated./handoff endpoint persists the reason, sends an SMS summary to the advisor and redirects the call to a phone number or a browser client.
The voice endpoint is short. This is the shape of it:
app.post("/voice", (req, res) => {
const twiml = `
<Response>
<Connect>
<ConversationRelay
url="wss://${process.env.HOST}/relay"
ttsProvider="ElevenLabs"
voice="${process.env.ELEVENLABS_VOICE_ID}"
language="es-AR"
interruptible="true"
reportInputDuringAgentSpeech="true"
/>
</Connect>
</Response>`;
res.type("text/xml").send(twiml);
});
Two attributes matter more than they look. interruptible is what lets the caller talk over the agent and have it stop instantly. I tested it by interrupting mid sentence; it cut immediately. And the provider and voice are configuration, not code: ConversationRelay supports several TTS providers, and swapping ElevenLabs for another one is an attribute change.
If you wait for the model to finish the whole answer before sending text to ConversationRelay, the caller hears seconds of nothing. In a phone call that reads as «the line dropped». The relay streams tokens as they arrive:
for await (const chunk of stream) {
const token = chunk.choices?.[0]?.delta?.content;
if (!token) continue;
ws.send(JSON.stringify({ type: "text", token, last: false }));
}
ws.send(JSON.stringify({ type: "text", token: "", last: true }));
The second latency lever is context loading. Sending the full company context on turn one made the first response noticeably slower, so the project has two strategies behind an env var: full sends everything on the first turn; lazy sends the vital part first and loads the rest while the first answer is being spoken, so the agent is fully loaded by turn two. For a FAQ agent, lazy wins.
The most important file in the repo is the tools definition in the agent folder. The handoff is a tool the model can call, with two triggers stated in its description: the caller explicitly asks for a person, or the request is something the agent cannot do. A second tool ends the conversation when the caller says goodbye.
export const tools = [
{
type: "function",
function: {
name: "request_handoff",
description:
"Transfer the call to a human advisor. Use when the caller asks to speak with a person, or when the request is outside what you can resolve (booking a test drive, complaints, anything requiring action on an account).",
parameters: {
type: "object",
properties: {
reason: { type: "string", description: "Short summary of why the caller needs a human, in the caller's language." },
department: { type: "string", enum: ["sales", "support"] },
},
required: ["reason"],
},
},
},
{
type: "function",
function: {
name: "end_conversation",
description: "End the call when the caller says goodbye or has nothing else to ask.",
parameters: { type: "object", properties: {} },
},
},
];
When request_handoff fires, the relay saves the decision and the reason, texts the advisor a summary, and asks Twilio to redirect the call. The advisor picks up already knowing why the person is calling. In the demo there is a single department; the same structure supports routing to sales, support or the workshop by number.
Adding capabilities to the agent is adding entries to that array. That is the whole extensibility story, and it is enough for most first deployments.
The system prompt states who the agent is, that it must answer only from the context below, how it should talk, when it should end the call, and that it has request_handoff available. It also injects the current date and time on every call:
export function systemPrompt(context: string) {
const now = new Date().toLocaleString("es-AR", { timeZone: process.env.TIMEZONE ?? "America/Argentina/Buenos_Aires" });
return `You answer the phone for ${process.env.COMPANY_NAME}.
Answer only with the information in the context below. If something is not there, say you don't know and offer a human advisor.
Current date and time: ${now}. Use it to answer whether the business is open, how long until it opens, and any relative date.
Speak briefly, in a warm and direct tone, in Rioplatense Spanish.
End the call with end_conversation when the caller says goodbye.
Use request_handoff when the caller asks for a person or asks for something you cannot do.
Context:
${context}`;
}
Passing the date sounds trivial. It is the most common bug I find in voice and messaging agents: a model has no clock, and without one it answers «are you open right now» confidently and wrong. Every relative date question in the demo (open now, how long until Monday) is handled by that one line.
The project talks to models through OpenRouter, so the provider is an env var. I used Claude Haiku 4.5. It is a light model and it is one of the best I have used for real time conversation, precisely because it is light. A model that «thinks» for two extra seconds before speaking makes the agent sound robotic, and a caller waiting on the line hangs up before the smart answer arrives.
One thing that surprised me: Whisper transcribed «tre drives» instead of «test drives». The model still understood and answered about test drives. STT does not have to be perfect for a FAQ agent when the model has enough context to recover.
Two bills.
OpenRouter, Haiku 4.5: about 5 cents of a dollar for roughly 5 minutes of conversation. Call it 1 cent per minute of model.
Twilio ConversationRelay: 9 minutes of calls, 63 cents. That covers the telephony, the speech to text and the text to speech with the ElevenLabs voice. About 7 cents per minute.
So the total lands around 8 cents a minute, and the model is one cent of it. If you are optimizing your voice agent by switching models, you are optimizing the wrong line. Call length is what moves the bill, and the handoff is what shortens the calls the agent cannot resolve.
Infrastructure as bottleneck. Many of the clients I work with do not want to run servers or audio infrastructure. For them the managed path is correct: pay a fee on top of the communication stack you already have, Twilio for the audio, ElevenLabs for the conversational loop, and skip the ops. That is a legitimate deployment and I recommend it often.
Whether you need to tune every step. The moment the flow needs the agent to decide when to listen and when to speak, or to sit in a conference between a customer and a human support person and only intervene when addressed, or to leave the call on a criterion other than «the caller said goodbye», you need your own brain and full control of the turn logic. A telephony vendor should not decide when your agent hears.
The handoff logic is always yours. The billing complaint is the canonical case: an angry caller, an agent with no tools to fix the charge, and an agent that tries anyway. The customer gets more frustrated, now with a machine, and the problem escalates instead of resolving. There has to be a human supervising, or at least available, and the agent needs a clean exit that it can trigger on explicit requests and, when you get there, on sentiment. It sounds basic. In 8 of every 10 flows I review, it is the step that got skipped.
Pick the managed path if you must. Keep the decision of when to hand off in your own code regardless.
-Gonza
Building something like this on Twilio? I help companies design and run these systems in production. Explore my Twilio consulting and development services.
Find out what your communication setup is costing you.
Get the communication audit