The platform
Whissle is the voice AI agents platform
Whissle is a voice-AI platform for building meta-aware, multimodal voice agents — AI that answers calls, books appointments, and acts on what it hears. It runs its own speech models, publishes its research openly, and ships a gateway you can self-host or use as a cloud API.
Who Whissle is for
Teams that put a voice agent in front of customers — receptionists and front desks, appointment booking, customer support, lead qualification, clinical intake — and the developers who build them. It fits companies that need their voice AI self-hosted or on-prem for privacy and compliance, as readily as teams that just want a hosted API.
What makes Whissle different
Meta-aware speech, in one pass
Whissle runs its own speech models that extract emotion, intent, age, gender and entities in a single forward pass — not a chain of third-party services stitched together after the sentence ends. The agent reacts to how something is said, not just the words.
Self-hostable, not cloud-only
The gateway runs as a Docker image you can host on your own infrastructure or on-prem — so audio and transcripts never have to leave your environment. Most voice-agent platforms are cloud-only; Whissle gives you the same product to run yourself, or as a hosted API.
Open models and research
The speech intelligence underneath Whissle is published — models on Hugging Face, papers you can read and cite. You are building on science that is in the open, not a black box.
Transparent, usage-based pricing
Voice agents are $0.06 per minute, all-in — speech, language model and speech synthesis included. No per-seat licensing to reach an agent, and a bill you can work out in ten seconds.
Core capabilities
How Whissle compares
Most voice-agent platforms — Vapi, Retell, Bland, ElevenLabs — are cloud-only orchestrators that stitch together someone else's speech-to-text, language model and voice synthesis. Whissle is built around its own meta-aware speech models and can be self-hosted end-to-end, so you get the emotion / intent / entity layer those pipelines don't have, and the option to keep every second of audio inside your own environment.