Whissle's speech engine returns more than words. With the transcript it returns what it heard about the speaker — emotion, intent, age, gender, dialect — and the entities it tagged inside the sentence, from the same forward pass. There is no second classification call and no LLM in the path.
Two ways in: a WebSocket you stream live audio into, and a POST you hand a file to. Both take the workspace API key you already have.
Every transcript — streaming segment or whole file — can carry these alongside the text:
| metadata.emotion | The emotion the acoustic head heard, e.g. EMOTION_NEUTRAL. |
| metadata.intent | What the utterance is doing, e.g. INTENT_GREETING. |
| metadata.age | Speaker age band, e.g. AGE_18_30. |
| metadata.gender | Speaker gender, e.g. GENDER_MALE. |
| metadata.dialect | Dialect label, where the loaded model has a dialect head. |
| entities | Tagged spans inside the sentence — {type, value, raw} — e.g. type LOCATION, value "San Francisco". |
| metadata_probs | The full distribution behind each label, as [{token, probability}] sorted high to low. Ask for it with metadata_prob. |
Which of these arrive depends on the model this deployment has loaded — the metadata heads are part of the model, not a switch. A model with no emotion head returns no emotion, and the field is simply absent rather than guessed. Ask GET /asr/status what is loaded before you build a UI around a field.
Create a secret key in Settings → API keys (owner/admin). It needs the models:invoke scope, which a secret (wsk_…) key is granted by default. Same key, same workspace wallet as the rest of the API.
On the REST endpoints it is a bearer token, exactly as everywhere else in these docs:
Authorization: Bearer wsk_live_xxxxxxxxxxxxxxxxxxxxOn the WebSocket it rides in the query string instead, as ?token=…. That is not a stylistic choice: the browser WebSocket API gives you no way to set a request header on the handshake, so a header-only scheme would have locked every browser client out. Treat the resulting URL as the secret it contains — do not log it, and do not put a wsk_ key in front-end code. For a browser, mint a short-lived token on your server and hand that to the page.
models:invoke opens /listen, /asr/stream and /asr/transcribe today. /asr/translate and /asr/s2s still require a wh_ device token: they compose speech with a language model, and those legs are not yet metered per second, so a workspace key there would be an unbilled door. Existing wh_ integrations keep working everywhere.WSSwss://aws-gateway-backend.whissle.ai/listen?token=…
WSSwss://aws-gateway-backend.whissle.ai/asr/stream?token=…
The two are the same relay to the same engine — /listen is the name to use when you are embedding live transcription in your own app. Note the host root: these sit beside /bot, not inside it.
The exchange, in order:
{"type":"config"} if the defaults suit you.is_final: false and are superseded; finals arrive with is_final: true and carry the metadata.{"type":"end"} to flush the tail of the audio. The server replies with a final {"type":"end"} once it has emitted everything, and you can close.The config message:
{
"type": "config",
"language": "en",
"sample_rate": 16000,
"use_lm": true,
"metadata_prob": true,
"top_k": 5,
"metadata_tags": ["emotion", "intent", "entity"],
"word_timestamps": false,
"hotwords": ["Acme Corp", "SKU-42"],
"hotword_weight": 10,
"punctuation": true,
"itn": true
}| language | Language code for model and language-model selection. Empty uses the server default. |
| sample_rate | Hz of the PCM you are sending. 16000 (default), or one of 8000, 22050, 44100, 48000 — anything else falls back to 16000. |
| use_lm | Language-model beam search (default true). false is a greedy decode: faster, less accurate. |
| beam_width | Beam width override. Omit for the server default. |
| metadata_prob | Include the probability distributions, not just the winning label. Default true. |
| top_k | How many entries per distribution when metadata_prob is on. Default 5. |
| metadata_tags | Restrict which heads run: emotion, intent, age, gender, dialect, entity, behavior, eval, role. Omit for all of them. |
| word_timestamps | Per-word start/end times, plus pauses and speech rate. Default true on the stream. |
| hotwords | A JSON array of words or phrases to boost during beam search — not a comma-separated string, which is what the file endpoint takes. Up to 200. |
| hotword_weight | How hard to boost them. Default 10. |
| punctuation | Punctuate and capitalise. Default true. |
| itn | Inverse text normalisation — spoken numbers and dates rendered as digits. Default true. |
| speaker_embedding | Attach the speaker embedding vector to each segment. Default false. |
| intent_labels | Your own label set; segments then also carry filtered_intents renormalised within it. |
| model | Pick a specific loaded model by id. Omit for the deployment default. |
A final transcript event:
{
"type": "transcript",
"channel": "microphone",
"text": "I need to move my flight to Friday.",
"audioOffset": 4.216,
"is_final": true,
"utterance_end": true,
"metadata": {
"age": "AGE_18_30",
"gender": "GENDER_MALE",
"emotion": "EMOTION_NEUTRAL",
"intent": "INTENT_REQUEST"
},
"entities": [
{ "type": "DATE", "value": "Friday", "raw": "ENTITY_DATE Friday END" }
],
"confidence": 0.93
}Optional keys appear only when they have something in them: metadata, metadata_probs, entities, words, pauses, speech_rate, speaker_id, speakerEmbedding, filtered_intents, alternatives, uncertain_words. Write your reader so a missing key is normal.
Besides transcript you may see {"type":"error"} (bad JSON, oversized frame) and {"type":"warning"} — the warning means you pushed audio faster than the engine could take it and frames were dropped. Which brings us to the one mistake everybody makes first:
A complete, runnable client:
// npm i ws
import WebSocket from "ws";
import { readFileSync } from "node:fs";
// 16 kHz mono Int16LE PCM, e.g.
// ffmpeg -i call.wav -f s16le -acodec pcm_s16le -ac 1 -ar 16000 call.pcm
const pcm = readFileSync("call.pcm");
const ws = new WebSocket(
`wss://aws-gateway-backend.whissle.ai/listen?token=${process.env.WHISSLE_API_KEY}`,
);
ws.on("open", async () => {
ws.send(JSON.stringify({
type: "config",
language: "en",
sample_rate: 16000,
metadata_prob: true,
top_k: 5,
}));
// 3200 bytes = 100 ms at 16 kHz / 16-bit / mono. Pace the sends: push the whole
// file at once and the server's queue overruns and drops frames.
for (let i = 0; i < pcm.length; i += 3200) {
ws.send(pcm.subarray(i, i + 3200));
await new Promise((r) => setTimeout(r, 100));
}
ws.send(JSON.stringify({ type: "end" }));
});
ws.on("message", (raw) => {
const ev = JSON.parse(raw.toString());
if (ev.type === "end") return ws.close();
if (ev.type === "warning") return console.warn(ev.message);
if (ev.type !== "transcript" || !ev.is_final) return;
console.log(ev.text, ev.metadata ?? {}, ev.entities ?? []);
});
ws.on("close", (code, reason) => {
if (code === 4001) console.error("auth rejected:", reason.toString());
});POSThttps://aws-gateway-backend.whissle.ai/asr/transcribe
POSThttps://aws-gateway-backend.whissle.ai/asr/transcribe/pcm
GEThttps://aws-gateway-backend.whissle.ai/asr/status
multipart/form-data with the audio under file. Use /asr/transcribe for a container the engine can decode (WAV, MP3 and friends) and /asr/transcribe/pcm when you already have raw samples and want to skip header parsing — that one additionally takes sample_rate, channels and bit_depth.
curl -X POST https://aws-gateway-backend.whissle.ai/asr/transcribe \
-H "Authorization: Bearer $WHISSLE_API_KEY" \
-F "file=@call.wav" \
-F "language=en" \
-F "metadata_prob=true" \
-F "top_k=3"{
"transcript": "Another theory states that it originated on the Barbary coast, San Francisco, California.",
"transcript_with_entities": "Another theory states that it originated on the ENTITY_LOCATION Barbary coast END, ENTITY_LOCATION San Francisco END, ENTITY_STATE California END.",
"metadata": {
"age": "AGE_18_30",
"gender": "GENDER_MALE",
"emotion": "EMOTION_NEUTRAL",
"intent": "INTENT_INFORM"
},
"entities": [
{ "type": "LOCATION", "value": "Barbary coast", "raw": "ENTITY_LOCATION Barbary coast END" },
{ "type": "LOCATION", "value": "San Francisco", "raw": "ENTITY_LOCATION San Francisco END" },
{ "type": "STATE", "value": "California", "raw": "ENTITY_STATE California END" }
],
"metadata_probs": {
"emotion": [
{ "token": "EMOTION_NEUTRAL", "probability": 0.999 },
{ "token": "EMOTION_SURPRISE", "probability": 0.0004 },
{ "token": "EMOTION_FEAR", "probability": 0.0002 }
]
},
"inference_time": "0.312s",
"model": "en-in-tech-misc"
}inference_time is a formatted string with its unit attached ("0.312s"), not a number — parse it accordingly.
| file | Required. The audio. |
| language | Language code. Empty uses the server default. |
| metadata_prob | Include the probability distributions. Default false on this endpoint. |
| top_k | Entries per distribution. Default 5. |
| word_timestamps | Per-word timings, pauses and speech rate. Default false on this endpoint. |
| hotwords | A comma-separated string here — the streaming config takes a JSON array instead. Easy to get backwards. |
| hotword_weight | Boost strength. Default 10. |
| use_lm | Language-model beam search. Default true. |
| beam_width | Beam width override. |
| speaker_embedding | Include the speaker embedding vector. Default false. |
| diarize | Separate speakers. Bound it with num_speakers, or min_speakers / max_speakers. |
| neuropsych_mode | A preset domain vocabulary; also turns word_timestamps on. |
Keep uploads under 50 MB — longer recordings belong on the stream, or split at silence. GET /asr/status reports the loaded models, the decoder and the execution provider; it is the honest answer to “will this deployment give me emotion?”.
Labels come back as the model's own tokens — EMOTION_NEUTRAL, AGE_18_30, GENDER_MALE, INTENT_INFORM — not as prose. They are stable strings, so match on them rather than on a display name, and render your own label on top.
The exact vocabulary is a property of the loaded model. Rather than hard-coding a list, ask for metadata_prob once on a sample clip and read the tokens the distribution actually contains: that is the set this deployment can ever return.
Entities are spans the model tagged inside the sentence. Each is { type, value, raw } — the type without its prefix, the text it covers, and the original tagged form. The clean transcript has the tags stripped out; transcript_with_entities keeps them inline if you would rather do your own parsing.
metadata_probs and look at the margin between first and second place before you act on it.On the REST endpoints:
| 401 | Missing, malformed, invalid or revoked key. The body names the header it wanted. |
| 413 | Request body over the gateway cap. |
| 429 | Rate limited. The speech routes carry their own per-caller budget, separate from the rest of the API. |
| 503 | The speech service is unreachable, or its circuit breaker is open — the response carries Retry-After. |
On the WebSocket, failures arrive as a close code:
| 4001 | Authentication failed — no token, wrong prefix, or a token the gateway would not validate. The close reason says which. |
| 1013 | Too many concurrent sessions on that engine. Back off and retry. |
| 1011 | The engine has no model loaded — the deployment is not ready. |
Rate limiting is keyed per caller, so one busy integration does not starve another sharing your egress IP.
POSThttps://aws-gateway-backend.whissle.ai/bot/api/models/transcribe
The platform also has a batch transcription route inside /bot, alongside /api/models/chat and /api/models/tts. It takes the same workspace key and the same models:invoke scope, meters per second of audio into your wallet, and returns an engine-agnostic shape:
curl -X POST https://aws-gateway-backend.whissle.ai/bot/api/models/transcribe \
-H "Authorization: Bearer $WHISSLE_API_KEY" \
-F "file=@call.wav" -F "language=en" -F "diarize=true"
# → { "text": "…", "segments": [ { "start": 0.0, "end": 3.2, "speaker": 0, "text": "…" } ],
# "duration_seconds": 12.4, "diarized": true, "cost_usd": "0.00124" }Two fields there are worth reading before you trust a result. diarized says whether speaker separation actually ran — every segment labelled speaker 0 is both a valid single-speaker answer and what a failed diarizer returns, so the segments alone cannot tell you. warnings lists anything you asked for that this deployment could not deliver, and it exists because the response shape is otherwise identical either way.
Whissle metadata can ride along on this route as an optional metadata object, but it is enrichment and it is conditional: the transcription itself runs on a pre-configured engine chosen by language, and the metadata pass only runs where the deployment has Whissle's own model wired up and for the languages its heads are trained on. When it cannot run, the field is absent and the transcript ships without it.
So: use this route when you want a metered, billed, engine-agnostic transcript and metadata is a bonus. Use /listen or /asr/transcribe when the metadata is the reason you are here.
Full reference for the other à-la-carte models: Models (à la carte) →