What Happens in the First Five Seconds of an AI Phone Call

By AI Anyone Team · 2026-08-06 · 6 min read · Technology

Most AI voice products answer with a stored greeting, so every call sounds the same. AI Anyone opens in silence instead, and the character reacts in its own words if you say nothing.

You tap the call button. Then nothing happens.

No chime. No "Hi there, I'm Albert Einstein, how can I help you today?" No cheerful stock sentence that you have heard before, because everyone hears it. The line is just open. The character is listening. It is waiting for you.

That pause is not a bug and it is not the system thinking. It is the design, and it was the single hardest decision in the whole calling feature.

A stored greeting is the moment a product tells you it is a product

Almost every voice product opens the same way. A pre-written greeting fires the moment the call connects. It is the safest possible choice for whoever built it. It proves the audio is working, it fills the dead air, and it reassures you that the thing is on.

It also costs you the illusion in the first second.

A stored greeting is identical on call one and on call forty. It arrives before you have said a word, so it cannot possibly be about you. It has the cadence of an announcement rather than a conversation. Even when the writing is good, the sameness gives it away: you are not talking to a character, you are talking to a menu that happens to have a face.

There is a practical cost too. When a product opens by speaking, the first few seconds belong to it, not to you. Start talking during the greeting and you are talking underneath a recording, so people learn to wait politely for the machine to finish. Nobody does that on a real phone call.

Think about what actually happens when a person picks up. They say something short and unrehearsed, shaped by whatever mood they were in a second ago. And more often than not, the person who placed the call speaks first anyway, because they are the one who wanted something.

You placed this call. The first move is yours.

What the character does with the silence

The call opens quiet and live. The character is listening, so your first word is the one that starts the conversation. Nothing you say gets buried under an announcement, and there is no greeting to sit through before the real call can begin. You say hello, or you say "explain relativity to me like I'm twelve," and the character takes it from there.

Speak first and you never see the rest of this.

About four and a half seconds

Silence forever would just be a broken product. So there is a limit.

If you say nothing for about four and a half seconds, the character speaks first. One short line, in its own words, in character.

That line is generated fresh for that call. It is not written in advance, it is not pulled from a list of five options, and it is not a greeting wearing a costume. Call the same character twice and you get two different lines. It happens at most once per call, so it never turns into a nagging loop, and it never fires at all if you speak first.

The line is a reaction to the situation the character is in: someone called, and now nobody is saying anything. People have opinions about that.

What the shape of that line looks like

To be clear: these are illustrations of the shape such a line takes, not transcripts of anyone's call. None of it is what you will get, because you will get something else.

  • A busy scientist leans toward mild impatience with warmth underneath it. He acknowledges the open line and nudges you toward the point, because he was in the middle of something.
  • A monarch leans toward formality and a little authority. The caller is expected to state their business. The pause is not treated as an awkward moment, it is treated as your turn, and you are wasting it.
  • A philosopher tends to make the silence itself the subject. The pause becomes the interesting thing, and the question comes straight back to you before you have asked one.

Same four and a half seconds of nothing. Three different reactions, because three different people are on the other end. That is the part a stored greeting can never do, no matter how well it is written, because it was written before the silence existed.

Call Einstein and try saying nothing for a few seconds. It is the fastest way to understand what this post is about.

The five seconds actually start a moment earlier

Before any of that, your browser asks for the microphone. That moment is part of the first impression too, and it is where a lot of voice products hand you a generic error and walk away.

There are designed states for the three ways it goes wrong: permission denied, no microphone hardware attached, and the microphone already busy in another app. Each one tells you which of the three happened, because "call failed" is not information, it is a shrug.

The rest of the call follows the same rule

The opening sets a standard, and the rest of the call is built to hold it.

  • You interrupt by tapping the screen. If the character is mid-thought and you already know what you want to say, tap and it stops. No waiting for a paragraph to finish.
  • It answers in the language you speak to it. Start in Spanish and it answers in Spanish. Switch to English halfway through and it switches with you. You do not set anything.
  • You can change the voice mid-call. No reconnecting, no hanging up. The new voice takes over from the next spoken clause. Browsing and previewing voices is free.
  • There are three speaking paces. Quick, Balanced and Unhurried, and all three are free.
  • One tap hides all the on-screen controls if you are recording, and it lasts only for that call.
  • When you hang up, the conversation is written into your text chat with that character. So the next thing it types to you already knows what you said out loud.

That last one matters more than it sounds. The call is not a separate toy that lives in its own box. It is the same relationship, continued in a different medium, and it comes back with you.

Pick someone from 13,000+ characters if Einstein is not the voice you want in your ear.

Calling an AI character is not new, and we did not invent it

Character.AI shipped free two-way character voice calls in June 2024, and named interruption in the same announcement. Calling an AI character is an established category, not a novelty, and anyone telling you otherwise is selling something.

What we did was look at the part everyone treats as solved plumbing, the first five seconds, and refuse to fill it with a recording. A greeting is easier to build and easier to make sound good in a marketing clip. Silence is harder, because it puts the burden on the character to react well to a moment nobody scripted.

What it costs to try

Calling runs in a normal web browser on your phone or your laptop. No app install, no download.

It does require a free account, because a call costs real money to run and a signed-out visitor cannot be metered. Text chat needs that same free account.

Once you are signed in, a free account gets 20 minutes of calling every 30 days. It refreshes, counted from your first call rather than from the first of the calendar month. That is enough for several real conversations, and enough to find out whether the opening does what this post says it does.

Pro is $20 per month and includes 240 minutes of calling per rolling 30 days, counted from your first call rather than from the first of the calendar month. Applying a chosen voice is part of Pro. All three paces are free. The pricing page has the details.

Every one of the 13,000+ characters can be called. Every one of them opens the same way: quietly, listening, waiting to hear what you want. What they do if you make them wait is up to them.

voice callscallingai charactersconversation designproduct design

Try it yourself, out loud

Sign in and you get 20 minutes of calls free every 30 days, no card. It runs in this browser, so there is nothing to install.

Call Einstein and hear him answer See how calling works
More from the AI Anyone blog →
© 2026 AI Anyone. All rights reserved.