Skip to main content
POST
Create a transcript

Authorizations

Authorization
string
header
required

Your API key: Authorization: Bearer aircaps_sk_...

Headers

Idempotency-Key
string

Up to 255 characters. Makes retries safe: the same key and body return the original job.

Maximum string length: 255

Body

application/json

Provide exactly one of audio_url or file_id.

model
enum<string>
default:a5s-async-v1

The model. Any other value returns 400 unsupported_model.

Available options:
a5s-async-v1
audio_url
string<uri>

An http(s) URL to download the audio from (public or presigned), up to 5 GB, resolving to a public address.

Maximum string length: 8192
Example:

"https://example.com/meeting.mp3"

file_id
string
Maximum string length: 64
language_code
string
default:en

Language of the audio. English only today: en or a regional variant (en-US, en-GB, en-AU, en-CA, en-IN, en-IE, en-NZ, en-ZA; case-insensitive, - or _). Any other value returns 400 unsupported_language. Multilingual support is planned.

Maximum string length: 16
Example:

"en-US"

webhook_url
string<uri>

Called with a webhook event when the job finishes.

Maximum string length: 2048
webhook_auth_header_name
string

Name of a header to add to webhook requests. Requires webhook_url.

Maximum string length: 256
Example:

"Authorization"

webhook_auth_header_value
string

Value of that header.

Maximum string length: 4096
metadata
object

Any JSON object up to 16 KB; echoed on the transcript and in webhooks.

Example:

Response

Created. A retried request with the same Idempotency-Key returns the original transcript with 200.

id
string
required

Transcript ID.

Example:

"tr_01k6z3m4q8w7e5r2t1y0u9i8o7"

object
enum<string>
required
Available options:
transcript
Example:

"transcript"

status
enum<string>
required

queued: waiting for a processing slot. processing: being transcribed. completed: result fields present. error: see error.

Available options:
queued,
processing,
completed,
error
created_at
string<date-time>
required

When the transcript was created (UTC).

started_at
string<date-time> | null
required

When processing started.

completed_at
string<date-time> | null
required

When the job finished (completed or error).

audio_url
string | null
required

The submitted audio_url, if any.

file_id
string | null
required

The submitted file_id, if any.

language_code
enum<string>
required

Language of the transcript (normalised).

Available options:
en
model
enum<string>
required
Available options:
a5s-async-v1
audio_duration
number | null
required

Audio duration in seconds, once known.

Example:

3600.512

metadata
object | null
required

Your metadata, echoed back.

webhook_url
string | null
required

Webhook target, if any.

webhook_status_code
integer | null
required

HTTP status of the last webhook delivery.

error
object | null
required

Set when status is error.

text
string

Completed only. Full transcript: verbatim, lowercase, no punctuation.

speakers
object[]

Completed only.

utterances
object[]

Completed only. Speaker turns, in time order; overlapping speech overlaps.

words
object[]

Completed only. Every word, in time order.