1. Provide the audio
Choose one:
Any audio or video format ffmpeg can read works. The first audio stream is used; channels are mixed down. Uploaded files expire after 24 hours.
2. Create the transcript
metadata (up to 16 KB) is echoed back on the transcript and in webhooks. language_code is optional and accepts English (en, en-US, …) only; multilingual support is planned.
3. Get the result
Wait for a webhook or pollGET /v1/transcripts/{id} every few seconds.
The result
- Times are seconds from the start of the audio.
- Speakers are numbered in order of first speech and are consistent within one transcript only.
- Overlapping speech produces overlapping utterances.
- Audio with no speech completes with empty
text,utterancesandwords.
GET /v1/transcripts/{id}/subtitles?format=srt.
Delete
DELETE /v1/transcripts/{id} cancels a running job or deletes a finished transcript and its audio. Transcripts are otherwise deleted after 30 days.