AI voice agent human handoff with Twilio ConversationRelay

October 02, 2026 • 9 min read

Home / Blog / AI voice agent human handoff with Twilio ConversationRelay

About the author

Author

Gonzalo Gomez

AI & Automation Specialist

I design AI-powered communication systems. My work focuses on voice agents, WhatsApp chatbots, AI assistants, and workflow automation built primarily on Twilio, n8n, and modern LLMs like OpenAI and Claude. Over the past 7 years, I've shipped 30+ automation projects handling 250k+ monthly interactions.

Subscribe to my newsletter

If you enjoy the content that I make, you can subscribe and receive insightful information through email. No spam is going to be sent, just updates about interesting posts or specialized content that I talk about.

AI voice agent human handoff with Twilio ConversationRelay | Hand off an AI voice agent call to a human with Twilio ConversationRelay, ElevenLabs and OpenRouter: real per-minute costs and the code that decides when

An AI voice agent can answer every frequently asked question a dealership gets, with the right tone and the right hours, and still lose the caller. Not because the answers are wrong. Because the moment the call needs a person, either the agent does not recognize it or the handoff is so clumsy that the customer hangs up first.

 

In the 80 to 90 percent of agentic voice flows I have audited for clients, that step is missing or bolted on at the end. This post walks through the minimal version that works in production: a Twilio number, ConversationRelay for the audio loop, an ElevenLabs voice, a small model through OpenRouter, and a handoff that the agent itself triggers. It also gives you the bill, because the bill changes the design.

 

The setup: a dealership that answers and escalates

 

The demo company is a car dealership. The agent has the company identity, opening hours (Monday to Friday 9 to 19, Saturday 9 to 13), new and used inventory, financing rules, test drive rules, workshop and parts, warranties and FAQs. It also has one explicit instruction: when the request requires a person, for example scheduling a test drive, hand off to the sales team.

 

That scope is deliberate. When a business is starting with voice agents, you do not give the agent a lot of freedom. It answers FAQs, and what it cannot answer goes to sales. That is a common and correct first deployment, and the handoff is the piece that makes it safe.

 

A real call from the demo, condensed: I ask if they are open, get the hours. I ask about sedans, get «new sedans from 32 million pesos plus selected used». I ask about test drives, get the requirement (book a slot with an advisor, bring a valid license) and the offer to transfer me. I say yes, and two seconds later the call rings in my browser with a summary of what I asked already on screen.

 

Architecture: what Twilio owns and what you own

 

The layers, in order:

 

  • The caller dials a Twilio number. Twilio POSTs to my /voice endpoint (running locally through a tunnel during development, a VPS in production).
  • /voice returns TwiML that connects the call to ConversationRelay, with the TTS provider, the voice id and the language.
  • ConversationRelay handles speech to text and text to speech in Twilio's cloud and opens a WebSocket to my /relay endpoint. From here on my server only sees text.
  • /relay holds the agent: the company context, the prompt, the tools, the turn log. Responses are streamed back to the phone as they are generated.
  • When the agent decides to hand off, a separate /handoff endpoint persists the reason, sends an SMS summary to the advisor and redirects the call to a phone number or a browser client.

 

The voice endpoint is short. This is the shape of it:

 

app.post("/voice", (req, res) => {
  const twiml = `
    <Response>
      <Connect>
        <ConversationRelay
          url="wss://${process.env.HOST}/relay"
          ttsProvider="ElevenLabs"
          voice="${process.env.ELEVENLABS_VOICE_ID}"
          language="es-AR"
          interruptible="true"
          reportInputDuringAgentSpeech="true"
        />
      </Connect>
    </Response>`;
  res.type("text/xml").send(twiml);
});

 

Two attributes matter more than they look. interruptible is what lets the caller talk over the agent and have it stop instantly. I tested it by interrupting mid sentence; it cut immediately. And the provider and voice are configuration, not code: ConversationRelay supports several TTS providers, and swapping ElevenLabs for another one is an attribute change.

 

Streaming, because silence is the failure mode

 

If you wait for the model to finish the whole answer before sending text to ConversationRelay, the caller hears seconds of nothing. In a phone call that reads as «the line dropped». The relay streams tokens as they arrive:

 

for await (const chunk of stream) {
  const token = chunk.choices?.[0]?.delta?.content;
  if (!token) continue;
  ws.send(JSON.stringify({ type: "text", token, last: false }));
}
ws.send(JSON.stringify({ type: "text", token: "", last: true }));

 

The second latency lever is context loading. Sending the full company context on turn one made the first response noticeably slower, so the project has two strategies behind an env var: full sends everything on the first turn; lazy sends the vital part first and loads the rest while the first answer is being spoken, so the agent is fully loaded by turn two. For a FAQ agent, lazy wins.

 

The handoff is a tool, not an afterthought

 

The most important file in the repo is the tools definition in the agent folder. The handoff is a tool the model can call, with two triggers stated in its description: the caller explicitly asks for a person, or the request is something the agent cannot do. A second tool ends the conversation when the caller says goodbye.

 

export const tools = [
  {
    type: "function",
    function: {
      name: "request_handoff",
      description:
        "Transfer the call to a human advisor. Use when the caller asks to speak with a person, or when the request is outside what you can resolve (booking a test drive, complaints, anything requiring action on an account).",
      parameters: {
        type: "object",
        properties: {
          reason: { type: "string", description: "Short summary of why the caller needs a human, in the caller's language." },
          department: { type: "string", enum: ["sales", "support"] },
        },
        required: ["reason"],
      },
    },
  },
  {
    type: "function",
    function: {
      name: "end_conversation",
      description: "End the call when the caller says goodbye or has nothing else to ask.",
      parameters: { type: "object", properties: {} },
    },
  },
];

 

When request_handoff fires, the relay saves the decision and the reason, texts the advisor a summary, and asks Twilio to redirect the call. The advisor picks up already knowing why the person is calling. In the demo there is a single department; the same structure supports routing to sales, support or the workshop by number.

 

Adding capabilities to the agent is adding entries to that array. That is the whole extensibility story, and it is enough for most first deployments.

 

The prompt: tell the agent what day it is

 

The system prompt states who the agent is, that it must answer only from the context below, how it should talk, when it should end the call, and that it has request_handoff available. It also injects the current date and time on every call:

 

export function systemPrompt(context: string) {
  const now = new Date().toLocaleString("es-AR", { timeZone: process.env.TIMEZONE ?? "America/Argentina/Buenos_Aires" });
  return `You answer the phone for ${process.env.COMPANY_NAME}.
Answer only with the information in the context below. If something is not there, say you don't know and offer a human advisor.
Current date and time: ${now}. Use it to answer whether the business is open, how long until it opens, and any relative date.
Speak briefly, in a warm and direct tone, in Rioplatense Spanish.
End the call with end_conversation when the caller says goodbye.
Use request_handoff when the caller asks for a person or asks for something you cannot do.

Context:
${context}`;
}

 

Passing the date sounds trivial. It is the most common bug I find in voice and messaging agents: a model has no clock, and without one it answers «are you open right now» confidently and wrong. Every relative date question in the demo (open now, how long until Monday) is handled by that one line.

 

The model: small on purpose

 

The project talks to models through OpenRouter, so the provider is an env var. I used Claude Haiku 4.5. It is a light model and it is one of the best I have used for real time conversation, precisely because it is light. A model that «thinks» for two extra seconds before speaking makes the agent sound robotic, and a caller waiting on the line hangs up before the smart answer arrives.

 

One thing that surprised me: Whisper transcribed «tre drives» instead of «test drives». The model still understood and answered about test drives. STT does not have to be perfect for a FAQ agent when the model has enough context to recover.

 

What it costs per minute

 

Two bills.

 

OpenRouter, Haiku 4.5: about 5 cents of a dollar for roughly 5 minutes of conversation. Call it 1 cent per minute of model.

 

Twilio ConversationRelay: 9 minutes of calls, 63 cents. That covers the telephony, the speech to text and the text to speech with the ElevenLabs voice. About 7 cents per minute.

 

So the total lands around 8 cents a minute, and the model is one cent of it. If you are optimizing your voice agent by switching models, you are optimizing the wrong line. Call length is what moves the bill, and the handoff is what shortens the calls the agent cannot resolve.

 

Three criteria before you let an agent answer your phone

 

Infrastructure as bottleneck. Many of the clients I work with do not want to run servers or audio infrastructure. For them the managed path is correct: pay a fee on top of the communication stack you already have, Twilio for the audio, ElevenLabs for the conversational loop, and skip the ops. That is a legitimate deployment and I recommend it often.

 

Whether you need to tune every step. The moment the flow needs the agent to decide when to listen and when to speak, or to sit in a conference between a customer and a human support person and only intervene when addressed, or to leave the call on a criterion other than «the caller said goodbye», you need your own brain and full control of the turn logic. A telephony vendor should not decide when your agent hears.

 

The handoff logic is always yours. The billing complaint is the canonical case: an angry caller, an agent with no tools to fix the charge, and an agent that tries anyway. The customer gets more frustrated, now with a machine, and the problem escalates instead of resolving. There has to be a human supervising, or at least available, and the agent needs a clean exit that it can trigger on explicit requests and, when you get there, on sentiment. It sounds basic. In 8 of every 10 flows I review, it is the step that got skipped.

 

Pick the managed path if you must. Keep the decision of when to hand off in your own code regardless.

 

-Gonza

7
Twilio,  AI Agents,  OpenRouter,  ConversationRelay,  ElevenLabs,  Voice AI
Published on October 02, 2026

Building something like this on Twilio? I help companies design and run these systems in production. Explore my Twilio consulting and development services.

Find out what your communication setup is costing you.

Get the communication audit

Related posts

Building an AI Outbound Call Sales Assistant with n8n, Twilio, and ElevenLabs

January 30, 2026
IntroductionOutbound sales calls are one of the hardest channels to automate with AI. Latency matters.Costs compound fast.Hallucinations are unacceptable.And voice systems fail loudly when something breaks. In... Read more

How to Run a Free AI Agent on Your VPS with Hermes

August 18, 2026
I have an agent running 24/7 on a VPS I was already paying for and barely using. It sends me an AI and software news... Read more

AI lead recovery system: Voice, WhatsApp, and SMS with N8N

April 27, 2026
AI Lead Recovery System: How I Built a Multi-Channel Outreach Agent with N8N, Twilio, and ElevenLabs IntroductionMost businesses lose leads not because the product is wrong,... Read more

Claude Agent Skills for a WhatsApp AI agent: production setup

September 11, 2026
The first thing my agent did on camera was fail a tool call. It called create_client with arguments my function did not accept, got an... Read more