> ## Documentation Index
> Fetch the complete documentation index at: https://docs.aircaps.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Batch transcription

> Transcribe recordings with A5S Async v1.

## 1. Provide the audio

Choose one:

| Source | How | Size |
| - | - | - |
| A URL | `audio_url` in the create request. Public or presigned http(s) URL. | ≤ 5 GB |
| Upload URL | [`POST /v1/files/upload-url`](/api-reference/upload-url), then `PUT` the bytes to the returned `upload_url` within an hour. Use the returned `id` as `file_id`. | ≤ 5 GB |
| Direct upload | [`POST /v1/files`](/api-reference/upload-file) with the raw bytes as the body. | ≤ 1 GB |

Any audio or video format ffmpeg can read works. The first audio stream is used; channels are mixed down. Uploaded files expire after 24 hours.

## 2. Create the transcript

```bash theme={null}
curl https://api.aircaps.com/v1/transcripts \
  -H "Authorization: Bearer $AIRCAPS_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: meeting-2026-10-06" \
  -d '{
        "model": "a5s-async-v1",
        "file_id": "file_01k6z...",
        "language_code": "en",
        "webhook_url": "https://example.com/hooks/aircaps",
        "metadata": {"meeting_id": 42}
      }'
```

`metadata` (up to 16 KB) is echoed back on the transcript and in webhooks. `language_code` is optional and accepts English (`en`, `en-US`, …) only; multilingual support is planned.

## 3. Get the result

Wait for a [webhook](/guides/webhooks) or poll [`GET /v1/transcripts/{id}`](/api-reference/get-transcript) every few seconds.

| status | Meaning |
| - | - |
| `queued` | Waiting for one of your account's processing slots |
| `processing` | Being transcribed |
| `completed` | Done; result fields are present |
| `error` | Failed; see `error` |

## The result

```json theme={null}
{
  "id": "tr_01k6z3m4q8w7e5r2t1y0u9i8o7",
  "status": "completed",
  "model": "a5s-async-v1",
  "audio_duration": 3600.512,
  "text": "hey good afternoon everybody ...",
  "speakers": [{"id": "speaker_0", "speaking_time": 1402.11}, {"id": "speaker_1", "speaking_time": 988.4}],
  "utterances": [
    {"speaker": "speaker_0", "start": 0.48, "end": 6.12, "text": "hey good afternoon everybody ...",
     "words": [{"text": "hey", "start": 0.48, "end": 0.71, "speaker": "speaker_0"}]}
  ],
  "words": [{"text": "hey", "start": 0.48, "end": 0.71, "speaker": "speaker_0"}]
}
```

* Times are seconds from the start of the audio.
* Speakers are numbered in order of first speech and are consistent within one transcript only.
* Overlapping speech produces overlapping utterances.
* Audio with no speech completes with empty `text`, `utterances` and `words`.

Need captions? [`GET /v1/transcripts/{id}/subtitles?format=srt`](/api-reference/subtitles).

## Delete

[`DELETE /v1/transcripts/{id}`](/api-reference/delete-transcript) cancels a running job or deletes a finished transcript and its audio. Transcripts are otherwise deleted after 30 days.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.