Everything we build in the open — SDKs, research, benchmarks, and applications. All repos
Command-line interface for the Voice Agents platform.
Browser JS/TS SDK for embedding a live voice agent in any web app.
Multi-modal Python client for Whissle speech, translation and TTS.
Model Context Protocol server exposing Whissle tools to AI dev ecosystems.
Lightweight JavaScript client for the Whissle API.
All-in-one, prompt-driven speech transcription.
Large-scale entity tagging from real and synthetic speech.
Visual-aware speech recognition.
One-step audio-visual understanding.
Retrieval-augmented generation over speech, text and visual context.
Reference inference for Whissle speech-to-text models.
Segmenting speech data with ASR-driven keyphrase spotting.
Gesture recognition for multimodal interaction.
Tool-agent-user interaction benchmark, adapted for voice agents.
Head-to-head evaluation of ASR and agent models.
Binary deception-detection research on speech.
A Firefox-based browser reimagined around a built-in voice agent.
Native macOS transcription client on the Whissle API.
Personal AI middleware for Claude Code, Cursor and VS Code.
Voice-native personal-assistant experiments.
Avatars that listen back — real-time listening avatars.
Agentic YouTube data downloader built on Google ADK.