DocsSpeech-to-text

Speech-to-text

Whissle's speech engine returns more than words. With the transcript it returns what it heard about the speaker — emotion, intent, age, gender, dialect — and the entities it tagged inside the sentence, from the same forward pass. There is no second classification call and no LLM in the path.

Two ways in: a WebSocket you stream live audio into, and a POST you hand a file to. Both take the workspace API key you already have.

What comes back

Every transcript — streaming segment or whole file — can carry these alongside the text:

metadata.emotionThe emotion the acoustic head heard, e.g. EMOTION_NEUTRAL.
metadata.intentWhat the utterance is doing, e.g. INTENT_GREETING.
metadata.ageSpeaker age band, e.g. AGE_18_30.
metadata.genderSpeaker gender, e.g. GENDER_MALE.
metadata.dialectDialect label, where the loaded model has a dialect head.
entitiesTagged spans inside the sentence — {type, value, raw} — e.g. type LOCATION, value "San Francisco".
metadata_probsThe full distribution behind each label, as [{token, probability}] sorted high to low. Ask for it with metadata_prob.

Which of these arrive depends on the model this deployment has loaded — the metadata heads are part of the model, not a switch. A model with no emotion head returns no emotion, and the field is simply absent rather than guessed. Ask GET /asr/status what is loaded before you build a UI around a field.

Your key, and where it goes

Create a secret key in Settings → API keys (owner/admin). It needs the models:invoke scope, which a secret (wsk_…) key is granted by default. Same key, same workspace wallet as the rest of the API.

On the REST endpoints it is a bearer token, exactly as everywhere else in these docs:

Authorization: Bearer wsk_live_xxxxxxxxxxxxxxxxxxxx

On the WebSocket it rides in the query string instead, as ?token=…. That is not a stylistic choice: the browser WebSocket API gives you no way to set a request header on the handshake, so a header-only scheme would have locked every browser client out. Treat the resulting URL as the secret it contains — do not log it, and do not put a wsk_ key in front-end code. For a browser, mint a short-lived token on your server and hand that to the page.

Availability
A workspace key carrying models:invoke opens /listen, /asr/stream and /asr/transcribe today. /asr/translate and /asr/s2s still require a wh_ device token: they compose speech with a language model, and those legs are not yet metered per second, so a workspace key there would be an unbilled door. Existing wh_ integrations keep working everywhere.

Streaming — WebSocket

WSSwss://aws-gateway-backend.whissle.ai/listen?token=…

WSSwss://aws-gateway-backend.whissle.ai/asr/stream?token=…

The two are the same relay to the same engine — /listen is the name to use when you are embedding live transcription in your own app. Note the host root: these sit beside /bot, not inside it.

The exchange, in order:

  1. Connect with the token in the query string.
  2. Send one JSON config message. Everything in it is optional; send {"type":"config"} if the defaults suit you.
  3. Send binary frames of raw PCM — signed 16-bit little-endian, mono, 16 kHz by default. One frame is at most 1 MB.
  4. Read JSON events as they come. Partial results arrive with is_final: false and are superseded; finals arrive with is_final: true and carry the metadata.
  5. Send {"type":"end"} to flush the tail of the audio. The server replies with a final {"type":"end"} once it has emitted everything, and you can close.

The config message:

{
  "type": "config",
  "language": "en",
  "sample_rate": 16000,
  "use_lm": true,
  "metadata_prob": true,
  "top_k": 5,
  "metadata_tags": ["emotion", "intent", "entity"],
  "word_timestamps": false,
  "hotwords": ["Acme Corp", "SKU-42"],
  "hotword_weight": 10,
  "punctuation": true,
  "itn": true
}
languageLanguage code for model and language-model selection. Empty uses the server default.
sample_rateHz of the PCM you are sending. 16000 (default), or one of 8000, 22050, 44100, 48000 — anything else falls back to 16000.
use_lmLanguage-model beam search (default true). false is a greedy decode: faster, less accurate.
beam_widthBeam width override. Omit for the server default.
metadata_probInclude the probability distributions, not just the winning label. Default true.
top_kHow many entries per distribution when metadata_prob is on. Default 5.
metadata_tagsRestrict which heads run: emotion, intent, age, gender, dialect, entity, behavior, eval, role. Omit for all of them.
word_timestampsPer-word start/end times, plus pauses and speech rate. Default true on the stream.
hotwordsA JSON array of words or phrases to boost during beam search — not a comma-separated string, which is what the file endpoint takes. Up to 200.
hotword_weightHow hard to boost them. Default 10.
punctuationPunctuate and capitalise. Default true.
itnInverse text normalisation — spoken numbers and dates rendered as digits. Default true.
speaker_embeddingAttach the speaker embedding vector to each segment. Default false.
intent_labelsYour own label set; segments then also carry filtered_intents renormalised within it.
modelPick a specific loaded model by id. Omit for the deployment default.

A final transcript event:

{
  "type": "transcript",
  "channel": "microphone",
  "text": "I need to move my flight to Friday.",
  "audioOffset": 4.216,
  "is_final": true,
  "utterance_end": true,
  "metadata": {
    "age": "AGE_18_30",
    "gender": "GENDER_MALE",
    "emotion": "EMOTION_NEUTRAL",
    "intent": "INTENT_REQUEST"
  },
  "entities": [
    { "type": "DATE", "value": "Friday", "raw": "ENTITY_DATE Friday END" }
  ],
  "confidence": 0.93
}

Optional keys appear only when they have something in them: metadata, metadata_probs, entities, words, pauses, speech_rate, speaker_id, speakerEmbedding, filtered_intents, alternatives, uncertain_words. Write your reader so a missing key is normal.

Besides transcript you may see {"type":"error"} (bad JSON, oversized frame) and {"type":"warning"} — the warning means you pushed audio faster than the engine could take it and frames were dropped. Which brings us to the one mistake everybody makes first:

Pace your frames
The server holds a bounded audio queue and drops the oldest frames when you overrun it. Firing a whole file into the socket in a tight loop therefore loses the middle of it — silently, apart from that warning event. Send audio at roughly the rate it was recorded. If you have a file and no live microphone, use the file endpoint instead; it is what it is for.

A complete, runnable client:

// npm i ws
import WebSocket from "ws";
import { readFileSync } from "node:fs";

// 16 kHz mono Int16LE PCM, e.g.
//   ffmpeg -i call.wav -f s16le -acodec pcm_s16le -ac 1 -ar 16000 call.pcm
const pcm = readFileSync("call.pcm");

const ws = new WebSocket(
  `wss://aws-gateway-backend.whissle.ai/listen?token=${process.env.WHISSLE_API_KEY}`,
);

ws.on("open", async () => {
  ws.send(JSON.stringify({
    type: "config",
    language: "en",
    sample_rate: 16000,
    metadata_prob: true,
    top_k: 5,
  }));

  // 3200 bytes = 100 ms at 16 kHz / 16-bit / mono. Pace the sends: push the whole
  // file at once and the server's queue overruns and drops frames.
  for (let i = 0; i < pcm.length; i += 3200) {
    ws.send(pcm.subarray(i, i + 3200));
    await new Promise((r) => setTimeout(r, 100));
  }
  ws.send(JSON.stringify({ type: "end" }));
});

ws.on("message", (raw) => {
  const ev = JSON.parse(raw.toString());
  if (ev.type === "end") return ws.close();
  if (ev.type === "warning") return console.warn(ev.message);
  if (ev.type !== "transcript" || !ev.is_final) return;
  console.log(ev.text, ev.metadata ?? {}, ev.entities ?? []);
});

ws.on("close", (code, reason) => {
  if (code === 4001) console.error("auth rejected:", reason.toString());
});

Batch — one file, one call

POSThttps://aws-gateway-backend.whissle.ai/asr/transcribe

POSThttps://aws-gateway-backend.whissle.ai/asr/transcribe/pcm

GEThttps://aws-gateway-backend.whissle.ai/asr/status

multipart/form-data with the audio under file. Use /asr/transcribe for a container the engine can decode (WAV, MP3 and friends) and /asr/transcribe/pcm when you already have raw samples and want to skip header parsing — that one additionally takes sample_rate, channels and bit_depth.

curl -X POST https://aws-gateway-backend.whissle.ai/asr/transcribe \
  -H "Authorization: Bearer $WHISSLE_API_KEY" \
  -F "file=@call.wav" \
  -F "language=en" \
  -F "metadata_prob=true" \
  -F "top_k=3"
{
  "transcript": "Another theory states that it originated on the Barbary coast, San Francisco, California.",
  "transcript_with_entities": "Another theory states that it originated on the ENTITY_LOCATION Barbary coast END, ENTITY_LOCATION San Francisco END, ENTITY_STATE California END.",
  "metadata": {
    "age": "AGE_18_30",
    "gender": "GENDER_MALE",
    "emotion": "EMOTION_NEUTRAL",
    "intent": "INTENT_INFORM"
  },
  "entities": [
    { "type": "LOCATION", "value": "Barbary coast", "raw": "ENTITY_LOCATION Barbary coast END" },
    { "type": "LOCATION", "value": "San Francisco", "raw": "ENTITY_LOCATION San Francisco END" },
    { "type": "STATE", "value": "California", "raw": "ENTITY_STATE California END" }
  ],
  "metadata_probs": {
    "emotion": [
      { "token": "EMOTION_NEUTRAL", "probability": 0.999 },
      { "token": "EMOTION_SURPRISE", "probability": 0.0004 },
      { "token": "EMOTION_FEAR", "probability": 0.0002 }
    ]
  },
  "inference_time": "0.312s",
  "model": "en-in-tech-misc"
}

inference_time is a formatted string with its unit attached ("0.312s"), not a number — parse it accordingly.

fileRequired. The audio.
languageLanguage code. Empty uses the server default.
metadata_probInclude the probability distributions. Default false on this endpoint.
top_kEntries per distribution. Default 5.
word_timestampsPer-word timings, pauses and speech rate. Default false on this endpoint.
hotwordsA comma-separated string here — the streaming config takes a JSON array instead. Easy to get backwards.
hotword_weightBoost strength. Default 10.
use_lmLanguage-model beam search. Default true.
beam_widthBeam width override.
speaker_embeddingInclude the speaker embedding vector. Default false.
diarizeSeparate speakers. Bound it with num_speakers, or min_speakers / max_speakers.
neuropsych_modeA preset domain vocabulary; also turns word_timestamps on.

Keep uploads under 50 MB — longer recordings belong on the stream, or split at silence. GET /asr/status reports the loaded models, the decoder and the execution provider; it is the honest answer to “will this deployment give me emotion?”.

Metadata reference

Labels come back as the model's own tokens — EMOTION_NEUTRAL, AGE_18_30, GENDER_MALE, INTENT_INFORM — not as prose. They are stable strings, so match on them rather than on a display name, and render your own label on top.

The exact vocabulary is a property of the loaded model. Rather than hard-coding a list, ask for metadata_prob once on a sample clip and read the tokens the distribution actually contains: that is the set this deployment can ever return.

Entities are spans the model tagged inside the sentence. Each is { type, value, raw } — the type without its prefix, the text it covers, and the original tagged form. The clean transcript has the tags stripped out; transcript_with_entities keeps them inline if you would rather do your own parsing.

A distribution is not a certainty
Emotion and intent are acoustic and semantic guesses with real error rates, and a single argmax label hides that. When a decision hangs on one, read metadata_probs and look at the margin between first and second place before you act on it.

Errors, limits and close codes

On the REST endpoints:

401Missing, malformed, invalid or revoked key. The body names the header it wanted.
413Request body over the gateway cap.
429Rate limited. The speech routes carry their own per-caller budget, separate from the rest of the API.
503The speech service is unreachable, or its circuit breaker is open — the response carries Retry-After.

On the WebSocket, failures arrive as a close code:

4001Authentication failed — no token, wrong prefix, or a token the gateway would not validate. The close reason says which.
1013Too many concurrent sessions on that engine. Back off and retry.
1011The engine has no model loaded — the deployment is not ready.

Rate limiting is keyed per caller, so one busy integration does not starve another sharing your egress IP.

The à-la-carte alternative

POSThttps://aws-gateway-backend.whissle.ai/bot/api/models/transcribe

The platform also has a batch transcription route inside /bot, alongside /api/models/chat and /api/models/tts. It takes the same workspace key and the same models:invoke scope, meters per second of audio into your wallet, and returns an engine-agnostic shape:

curl -X POST https://aws-gateway-backend.whissle.ai/bot/api/models/transcribe \
  -H "Authorization: Bearer $WHISSLE_API_KEY" \
  -F "file=@call.wav" -F "language=en" -F "diarize=true"
# → { "text": "…", "segments": [ { "start": 0.0, "end": 3.2, "speaker": 0, "text": "…" } ],
#     "duration_seconds": 12.4, "diarized": true, "cost_usd": "0.00124" }

Two fields there are worth reading before you trust a result. diarized says whether speaker separation actually ran — every segment labelled speaker 0 is both a valid single-speaker answer and what a failed diarizer returns, so the segments alone cannot tell you. warnings lists anything you asked for that this deployment could not deliver, and it exists because the response shape is otherwise identical either way.

Whissle metadata can ride along on this route as an optional metadata object, but it is enrichment and it is conditional: the transcription itself runs on a pre-configured engine chosen by language, and the metadata pass only runs where the deployment has Whissle's own model wired up and for the languages its heads are trained on. When it cannot run, the field is absent and the transcript ships without it.

So: use this route when you want a metered, billed, engine-agnostic transcript and metadata is a bonus. Use /listen or /asr/transcribe when the metadata is the reason you are here.

Full reference for the other à-la-carte models: Models (à la carte) →

Speech-to-text