FAQ

Voice AI agents — frequently asked questions

The questions buyers and developers ask most about voice agents and the Whissle platform.

What is a voice AI agent?
A voice AI agent is software that holds a real spoken conversation — it listens, understands, responds in a natural voice, and takes actions like booking an appointment or answering a question. Unlike a phone menu, it is open-ended: the caller just talks, and the agent handles the request end to end.
What is Whissle?
Whissle is a voice-AI platform for building meta-aware, multimodal voice agents that answer calls, book appointments, and act on what they hear. It runs its own speech models, publishes its research openly, and ships a gateway you can self-host or use as a cloud API.
How much does a Whissle voice agent cost?
Voice agents are $0.06 per minute, all-in — the price includes speech recognition, the language model and voice synthesis. There is no per-seat license required to reach an agent, so the bill scales with actual usage.
Can I self-host Whissle?
Yes. Whissle's gateway ships as a Docker image that runs the full voice-agent stack — speech recognition, the agent, and voice synthesis — on your own infrastructure or in your VPC, so audio and transcripts stay in your environment. It is also available as a managed cloud API.
What languages does Whissle support?
Whissle supports 23 languages, with its meta-aware speech recognition — emotion, intent and entity detection — built in across them, not bolted on for English only.
What can a Whissle voice agent do?
A Whissle agent can answer inbound and outbound phone calls and web voice sessions, book and manage appointments, answer questions from your own documents (RAG), call tools and integrations to take real actions, and hand off to a human when needed.
Does Whissle work over the phone?
Yes. Whissle agents run over the phone via SIP/Twilio and in the browser over WebRTC, and can be embedded as a widget on your website — the same agent across every channel.
What does “meta-aware” speech mean?
Meta-aware means Whissle's speech models extract more than the words: emotion, intent, age, gender and entities, all in a single forward pass. The agent can react to how something is said and who is saying it, not just the transcript.
How is Whissle different from Vapi, Retell or ElevenLabs?
Those platforms are cloud-only orchestrators that combine third-party speech, language and voice services. Whissle is built around its own meta-aware speech models and can be self-hosted end to end — so it adds the emotion/intent/entity layer those pipelines lack, and lets you keep all audio inside your own environment.
How do I get started with Whissle?
Create a free account at whissle.ai, describe the agent you want in plain language, and connect a phone number or embed it on your site. Developers can follow the build-a-voice-agent guide in the docs to go deeper.
Voice AI agents — frequently asked questions | Whissle · Whissle