A5S Async v1
a5s-async-v1 · REST · batch guide
- Recorded audio or video in any common format, from 0.1 s to 10 hours.
- Speaker labels (
speaker_0,speaker_1, …), word timestamps, utterances, SRT/VTT subtitles. - English only:
language_codeacceptsenand regional variants (en-US,en-GB, …); any other value returns400 unsupported_language. - Text is verbatim and lowercase without punctuation; disfluencies (“uh”, “um”) are kept.
- Typical turnaround: well under a minute for most files once processing starts.
A5S v2 Streaming
a5s-v2-streaming · WebSocket · streaming guide
- Live 16 kHz mono PCM16 audio in, partial and final text out as people speak.
- Punctuated, cased text. No speaker labels.
- Optional custom vocabulary to bias recognition toward names and terms.
- English only: the start message’s
languageacceptsenand regional variants (en-US,en-GB, …); any other value is refused withunsupported_language. - Sessions last up to 5 minutes.
Choosing
Languages
Both models support English today. Multilingual support is planned and will be available soon. Accepted codes:en, en-US, en-GB, en-AU, en-CA, en-IN, en-IE, en-NZ, en-ZA (case-insensitive; _ works too).
Access
Accounts can use both models by default. An account may be limited to one model. Requests to a model your account can’t use return403 model_access_denied (streaming: an error message with that code, then close code 4403). GET /v1/account shows enabled for each model. To get access, talk to sales.
GET /v1/models lists the models.