Skip to main content

1. Provide the audio

Choose one: Any audio or video format ffmpeg can read works. The first audio stream is used; channels are mixed down. Uploaded files expire after 24 hours.

2. Create the transcript

metadata (up to 16 KB) is echoed back on the transcript and in webhooks. language_code is optional and accepts English (en, en-US, …) only; multilingual support is planned.

3. Get the result

Wait for a webhook or poll GET /v1/transcripts/{id} every few seconds.

The result

  • Times are seconds from the start of the audio.
  • Speakers are numbered in order of first speech and are consistent within one transcript only.
  • Overlapping speech produces overlapping utterances.
  • Audio with no speech completes with empty text, utterances and words.
Need captions? GET /v1/transcripts/{id}/subtitles?format=srt.

Delete

DELETE /v1/transcripts/{id} cancels a running job or deletes a finished transcript and its audio. Transcripts are otherwise deleted after 30 days.