# Get account
Source: https://docs.aircaps.com/api-reference/account
GET /v1/account
Your usage limit, per-model access (`enabled`) and limits. For higher limits or model access, [talk to sales](https://research.aircaps.com/contact).
# Create a transcript
Source: https://docs.aircaps.com/api-reference/create-transcript
POST /v1/transcripts
Start transcribing audio with A5S Async v1. The job runs asynchronously: poll [Get a transcript](/api-reference/get-transcript) or pass `webhook_url`.
# Delete a transcript
Source: https://docs.aircaps.com/api-reference/delete-transcript
DELETE /v1/transcripts/{id}
Cancels a `queued` or `processing` job (its `error.code` becomes `canceled`), or deletes a finished transcript and its stored data. Afterwards the ID returns `404`.
# Get a transcript
Source: https://docs.aircaps.com/api-reference/get-transcript
GET /v1/transcripts/{id}
Status and, once `completed`, the result.
# List transcripts
Source: https://docs.aircaps.com/api-reference/list-transcripts
GET /v1/transcripts
Your transcripts, newest first. Items omit the result fields (`text`, `speakers`, `utterances`, `words`).
# List models
Source: https://docs.aircaps.com/api-reference/models
GET /v1/models
The available models. No authentication needed.
# Error object
Source: https://docs.aircaps.com/api-reference/objects/error
Every HTTP error uses this envelope. See [Errors](/errors) for all codes.
# File object
Source: https://docs.aircaps.com/api-reference/objects/file
An uploaded file. Use id as file_id for one transcript; the file is deleted when that transcript finishes, or after 24 hours if unused.
# Transcript object
Source: https://docs.aircaps.com/api-reference/objects/transcript
Returned by every transcript endpoint. The result fields (text, speakers, utterances, words) appear only in Get a transcript, once status is completed.
# Utterance object
Source: https://docs.aircaps.com/api-reference/objects/utterance
One speaker turn. Times are seconds from the start of the audio.
# Webhook event
Source: https://docs.aircaps.com/api-reference/objects/webhook-event
The body AirCaps `POST`s to your `webhook_url`. See [Webhooks](/guides/webhooks) for signing and retries.
# Word object
Source: https://docs.aircaps.com/api-reference/objects/word
One word with its timing and speaker.
# Overview
Source: https://docs.aircaps.com/api-reference/overview
Base URL, authentication and conventions.
* **Base URL:** `https://api.aircaps.com/v1`
* **Auth:** `Authorization: Bearer aircaps_sk_...` on every request ([API keys](/authentication))
* **Format:** JSON request and response bodies; timestamps are RFC 3339 UTC; audio times are seconds.
* **IDs:** `tr_...` transcripts, `file_...` uploads, `req_...` request IDs (also in the `x-request-id` header).
* **Errors:** one envelope for all errors ([Errors](/errors)).
* **Idempotency:** send `Idempotency-Key` with `POST /v1/transcripts` to make retries safe.
* **Language:** English only today: `language_code` (batch) and `language` (streaming) accept `en` and regional variants (`en-US`, `en-GB`, …); anything else returns `unsupported_language`. Multilingual support is planned.
* **Pagination:** list endpoints return `has_more` and `next_before`; pass it as `before` for the next page.
| Endpoint | Purpose |
| - | - |
| `POST /v1/transcripts` | Create a transcript |
| `GET /v1/transcripts/{id}` | Get status and result |
| `GET /v1/transcripts` | List transcripts |
| `DELETE /v1/transcripts/{id}` | Cancel or delete |
| `GET /v1/transcripts/{id}/subtitles` | SRT or VTT captions |
| `POST /v1/files/upload-url` | Presigned upload (≤ 5 GB) |
| `POST /v1/files` | Direct upload (≤ 1 GB) |
| `WSS /v1/realtime` | Live streaming |
| `GET /v1/models` | Available models |
| `GET /v1/account` | Account and limits |
| `GET /v1/usage` | Daily usage |
**Objects:** [Transcript](/api-reference/objects/transcript), [Utterance](/api-reference/objects/utterance), [Word](/api-reference/objects/word), [File](/api-reference/objects/file), [Error](/api-reference/objects/error), [Webhook event](/api-reference/objects/webhook-event).
The full OpenAPI 3.1 spec is at [`/openapi.json`](/openapi.json).
# Realtime stream
Source: https://docs.aircaps.com/api-reference/realtime
Live transcription with A5S v2 Streaming over a WebSocket.
**WSS** `wss://api.aircaps.com/v1/realtime`
Requires access to A5S v2 Streaming. For a walkthrough, see [Live streaming](/guides/streaming).
## Connection
`Bearer aircaps_sk_...`, sent with the WebSocket handshake.
The flow: connect, send `start`, wait for `ready`, stream binary audio, send `finish`, then receive `finished`. The server then closes with `1000`.
## Client messages
### start
The first message, as a JSON text frame, within 45 seconds of connecting. Unknown fields are rejected.
`"start"`
Language of the audio. English only today: `en` or a regional variant (`en-US`, `en-GB`, `en-AU`, `en-CA`, `en-IN`, `en-IE`, `en-NZ`, `en-ZA`). Any other value is refused with `unsupported_language`. Multilingual support is planned.
Only `pcm_s16le` (signed 16-bit little-endian PCM).
Only `16000`.
Only `1` (mono).
The streaming engine. Only `"a5sv2"`.
Bias recognition toward `custom_vocabulary`.
Up to 64 terms.
A word or phrase, up to 100 characters.
Bias strength, 0 to 12.
```json Example theme={null}
{"type": "start", "language": "en-US", "encoding": "pcm_s16le", "sample_rate": 16000, "channels": 1,
"custom_vocabulary_enabled": true, "custom_vocabulary": [{"text": "AirCaps", "strength": 6}]}
```
### Audio
Binary frames of raw PCM in the `start` format: non-empty, an even number of bytes, at most 64 KB (larger frames end the session with `frame_too_large`; frames over 1 MB close the connection without a message). Send them after `ready`, in real time; a 2-second burst is allowed. 20 ms frames (640 bytes) are recommended.
### ping
`{"type": "ping"}`. The server answers `pong`. Any message, audio or `ping`, resets the 45-second idle timer.
### finish
`{"type": "finish"}`. The server flushes the final transcript, sends `finished`, and closes.
## Server messages
Each is a JSON text frame with a `type`.
### provider\_status
Startup progress while the session starts.
For example `"Getting things ready, please hold on"`.
### ready
Start sending audio.
Maximum audio for this session: 5 minutes, or less if little usage remains.
The account's usage limit, in milliseconds.
Usage so far, in milliseconds.
`"en-US"`
`16000`
`"pcm_s16le"`
### transcript
The complete transcript so far, punctuated and cased. Replace what you display instead of appending.
`true` when the current segment is settled.
`"a5sv2"`
```json Example theme={null}
{"type": "transcript", "provider": "a5sv2", "text": "Good afternoon, everybody.", "is_final": true}
```
### session\_limit
The session reached 5 minutes or the account's usage limit. The final transcript and `finished` follow, then close code 1000.
Why the session ended.
### finished
Audio received.
Audio counted toward usage.
Account usage after this session.
The account's usage limit.
### provider\_error
Streaming failed during the session. An `error` follows.
What failed.
### error
Sent before the server ends the session abnormally. **Act on `code`**: the WebSocket close code is also set, but proxies may not preserve it.
One of the codes below.
Human-readable detail.
| code | Close | Meaning |
| - | - | - |
| `invalid_start` | 4400 | The first message is missing, late or invalid |
| `unsupported_language` | 4400 | `language` is not English |
| `unknown_provider` | 4400 | `providers` names an unknown engine |
| `audio_before_ready` | 4400 | Audio sent before `ready` |
| `invalid_audio_frame` | 4400 | An empty or odd-length audio frame |
| `invalid_message` | 4400 | A text message other than `ping` or `finish` |
| `unauthorized` | 4401 | Missing or invalid API key |
| `account_disabled` | 4403 | The account is disabled |
| `model_access_denied` | 4403 | No access to A5S v2 Streaming ([talk to sales](https://research.aircaps.com/contact)) |
| `origin_not_allowed` | 4403 | Browser origin not allowed |
| `idle_timeout` | 4408 | No messages for 45 seconds |
| `quota_unavailable` | 4409 | Usage limit used up, or too many live sessions ([talk to sales](https://research.aircaps.com/contact)) |
| `realtime_rate_exceeded` | 4409 | Audio sent faster than real time |
| `frame_too_large` | 4409 | An audio frame over 64 KB |
| `a5sv2_unavailable`, `upstream_unavailable` | 1013 | Temporarily unavailable; retry |
| `internal_error` | 1011 | Server error |
### pong
The reply to `ping`.
## Close codes
| Code | Meaning |
| - | - |
| 1000 | Normal close after `finished` |
| 1011 | Server error |
| 1013 | Temporarily unavailable; retry with backoff |
| 4400 | Invalid message or frame |
| 4401 | Authentication failed |
| 4403 | Forbidden: account disabled, no access to the model, or browser origin not allowed |
| 4408 | Idle timeout |
| 4409 | Limit reached at start (usage, sessions), audio faster than real time, or a frame over 64 KB |
Close codes are informational; use the `error` message's `code`.
# Get subtitles
Source: https://docs.aircaps.com/api-reference/subtitles
GET /v1/transcripts/{id}/subtitles
SRT or WebVTT captions for a completed transcript, with the speaker in each cue.
# Upload a file
Source: https://docs.aircaps.com/api-reference/upload-file
POST /v1/files
Upload up to 1 GB in the request body. For larger files (up to 5 GB), use an upload URL.
# Create an upload URL
Source: https://docs.aircaps.com/api-reference/upload-url
POST /v1/files/upload-url
Get a presigned URL to upload a file of up to 5 GB straight to storage. `PUT` the raw bytes to `upload_url` within an hour, then use `id` as `file_id`.
# Get usage
Source: https://docs.aircaps.com/api-reference/usage
GET /v1/usage
Daily usage of both models.
# API keys
Source: https://docs.aircaps.com/authentication
Create, use and revoke keys.
Every request needs an API key in the `Authorization` header:
```http theme={null}
Authorization: Bearer aircaps_sk_...
```
* **Create keys** in the console: [playground.aircaps.com → API keys](https://playground.aircaps.com/api-keys). A key is shown once; AirCaps stores only a hash.
* **One key, both models.** The same key works for the REST API and the streaming WebSocket.
* **Up to 10 active keys** per account. Use one per service or environment so you can revoke them independently.
* **Revoke** a key in the console. It stops working within 60 seconds; a streaming session already open with it runs until it ends.
* **Keep keys server-side.** Never put a key in a browser, mobile app or public repository.
Requests without a valid key return `401 unauthorized`; a disabled account returns `403 account_disabled`. On the streaming WebSocket these arrive as an `error` message with the same `code`.
Account settings (keys, webhook signing secret, profile) can only be changed from the console, not with an API key.
# Data & privacy
Source: https://docs.aircaps.com/data-privacy
What AirCaps keeps, for how long, and how to delete it.
**We don't use your audio or transcripts to train or improve models.** Your data is used only to process your requests.
## What we keep
| Data | Kept for |
| - | - |
| **Streaming audio and transcripts** (A5S v2 Streaming) | Not stored. Audio is processed in memory and the text is sent only to you. |
| **Uploaded files** (`POST /v1/files` or an upload URL) | Deleted as soon as the transcript that uses them finishes (completed, failed or canceled). Unused uploads are deleted after 24 hours. |
| **Audio from `audio_url`** | Downloaded only while the transcript is processed, then deleted. |
| **Processing copies** of the audio | Deleted when the transcript finishes. |
| **Transcripts** (text, words, speakers) | 30 days after completion, or until you delete them with [`DELETE /v1/transcripts/{id}`](/api-reference/delete-transcript). |
| **Transcript details you send** (`metadata`, `audio_url`, webhook URL) | Removed when the transcript is deleted or expires. |
| **Webhook auth header** (`webhook_auth_header_value`) | Removed as soon as the webhook is delivered or delivery fails for good. |
| **Usage records** (duration, status, timestamps) | Kept for usage reporting and billing; they contain no audio, text or metadata. |
| **Account** (email, name, limits) | For the life of the account. API keys are stored only as SHA-256 hashes. |
| **Logs** | Requests (endpoint, status, timing and IDs). Logs don't contain audio or transcripts. |
## Deleting data
* **A transcript:** `DELETE /v1/transcripts/{id}` deletes it immediately. If it is still processing, it is canceled.
* **Uploads:** deleted automatically, as above.
* **Your account and all its data:** [contact us](https://research.aircaps.com/contact).
## Security
* **In transit:** every connection uses TLS: your client to the API, between our services, to storage and to our database (with certificate verification).
* **At rest:** stored data is encrypted with AES-256.
* **Isolation:** every request is scoped to your account; no account can read another's data.
* **Keys and webhooks:** API keys are stored as hashes and can be revoked at any time; webhooks are signed ([Standard Webhooks](/guides/webhooks)).
## Where data is processed
In the United States, using these infrastructure providers:
| Provider | Used for |
| - | - |
| Modal | Running the API and the models |
| Cloudflare (R2) | Storing uploads and transcripts |
| Supabase | Account and transcript records |
| Render | The public API endpoint (`api.aircaps.com`) and the console |
For a data processing agreement, custom retention or a dedicated deployment, [talk to sales](https://research.aircaps.com/contact).
# Errors
Source: https://docs.aircaps.com/errors
HTTP errors, job errors and retries.
## HTTP errors
Every error uses one shape:
```json theme={null}
{
"error": {
"type": "invalid_request_error",
"code": "unsupported_language",
"message": "language 'de' is not supported; only English (\"en\", \"en-US\", ...) is available today. Multilingual support is planned.",
"param": "language_code",
"request_id": "req_01k6z..."
}
}
```
| HTTP | type | code |
| - | - | - |
| 400 | `invalid_request_error` | `invalid_request`, `unsupported_language`, `unsupported_model` |
| 401 | `authentication_error` | `unauthorized` |
| 403 | `permission_error` | `account_suspended`, `account_disabled`, `session_required`, `usage_limit_reached`, `model_access_denied` |
| 404 | `invalid_request_error` | `not_found`, `file_not_found` |
| 405 | `invalid_request_error` | `method_not_allowed` |
| 409 | `invalid_request_error` | `idempotency_conflict`, `transcript_not_ready` |
| 413 | `invalid_request_error` | `file_too_large` |
| 429 | `rate_limit_error` | `rate_limited`, `queue_full` (honour `Retry-After`) |
| 500 | `api_error` | `internal_error` |
| 502, 504 | `api_error` | `upstream_unavailable`, `upstream_timeout` (`request_id` is `null`) |
Include `request_id` (also in the `x-request-id` header) when you contact support.
* `unsupported_language`: only English is available today; multilingual support is planned.
* `usage_limit_reached`, `model_access_denied`, `queue_full`: account limits. To raise them, [talk to sales](https://research.aircaps.com/contact).
## Job errors
A transcript that fails has `"status": "error"` and an `error` object. None of these count toward your usage.
| code | Meaning |
| - | - |
| `invalid_audio` | Not decodable, no audio stream, empty, or shorter than 0.1 s |
| `audio_too_long` | Longer than 10 hours |
| `file_too_large` | Larger than 5 GB |
| `download_failed` | `audio_url` could not be fetched (HTTP error, timeout, non-public address) |
| `file_not_found` | The uploaded file is missing or expired |
| `usage_limit_reached` | Longer than your remaining usage ([talk to sales](https://research.aircaps.com/contact) to raise it) |
| `internal_error` | Failed on our side after automatic retries. Resubmit; if you used `file_id`, upload the file again (uploads are deleted when their transcript finishes) |
Deleting a queued or processing transcript cancels it: the `DELETE` response shows `error.code` `canceled`, no webhook is sent, and the ID then returns `404`.
## Retries
* Retry `429`, `500`, `502` and `504` with exponential backoff.
* Send an `Idempotency-Key` header with `POST /v1/transcripts` so a retried request never creates a second job. A retry with the same key and body returns the original transcript (`200`), even after it has finished.
# Batch transcription
Source: https://docs.aircaps.com/guides/async-transcription
Transcribe recordings with A5S Async v1.
## 1. Provide the audio
Choose one:
| Source | How | Size |
| - | - | - |
| A URL | `audio_url` in the create request. Public or presigned http(s) URL. | ≤ 5 GB |
| Upload URL | [`POST /v1/files/upload-url`](/api-reference/upload-url), then `PUT` the bytes to the returned `upload_url` within an hour. Use the returned `id` as `file_id`. | ≤ 5 GB |
| Direct upload | [`POST /v1/files`](/api-reference/upload-file) with the raw bytes as the body. | ≤ 1 GB |
Any audio or video format ffmpeg can read works. The first audio stream is used; channels are mixed down.
* **Uploaded files are deleted as soon as the transcript that uses them finishes** (completed, failed or canceled). A `file_id` is for one transcript; to transcribe the same file again, upload it again. Unused uploads are deleted after 24 hours.
* An `audio_url` must be public or presigned, without credentials in the URL. Up to 5 redirects are followed (each must stay on a public address), and the download must finish within 15 minutes.
## 2. Create the transcript
```bash theme={null}
curl https://api.aircaps.com/v1/transcripts \
-H "Authorization: Bearer $AIRCAPS_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: meeting-2026-10-06" \
-d '{
"model": "a5s-async-v1",
"file_id": "file_01k6z...",
"language_code": "en",
"webhook_url": "https://example.com/hooks/aircaps",
"metadata": {"meeting_id": 42}
}'
```
`metadata` (up to 16 KB) is echoed back on the transcript and in webhooks. `language_code` is optional and accepts English (`en`, `en-US`, …) only; multilingual support is planned.
## 3. Get the result
Wait for a [webhook](/guides/webhooks) or poll [`GET /v1/transcripts/{id}`](/api-reference/get-transcript) every few seconds.
| status | Meaning |
| - | - |
| `queued` | Waiting for one of your account's processing slots |
| `processing` | Being transcribed |
| `completed` | Done; result fields are present |
| `error` | Failed; see `error` |
## The result
```json theme={null}
{
"id": "tr_01k6z3m4q8w7e5r2t1y0u9i8o7",
"status": "completed",
"model": "a5s-async-v1",
"audio_duration": 3600.512,
"text": "hey good afternoon everybody ...",
"speakers": [{"id": "speaker_0", "speaking_time": 1402.11}, {"id": "speaker_1", "speaking_time": 988.4}],
"utterances": [
{"speaker": "speaker_0", "start": 0.48, "end": 6.12, "text": "hey good afternoon everybody ...",
"words": [{"text": "hey", "start": 0.48, "end": 0.71, "speaker": "speaker_0"}]}
],
"words": [{"text": "hey", "start": 0.48, "end": 0.71, "speaker": "speaker_0"}]
}
```
* Times are seconds from the start of the audio.
* Speakers are numbered in order of first speech and are consistent within one transcript only.
* Overlapping speech produces overlapping utterances.
* Audio with no speech completes with empty `text`, `utterances` and `words`.
Need captions? [`GET /v1/transcripts/{id}/subtitles?format=srt`](/api-reference/subtitles).
## Delete
[`DELETE /v1/transcripts/{id}`](/api-reference/delete-transcript) cancels a running job or deletes a finished transcript and its stored result. Transcripts are otherwise deleted after 30 days.
# Code samples
Source: https://docs.aircaps.com/guides/samples
Copy-paste examples in Python, Node.js and cURL.
## Transcribe a local file
```python Python theme={null}
# pip install requests
import os, sys, time, requests
API = "https://api.aircaps.com/v1"
HEADERS = {"Authorization": f"Bearer {os.environ['AIRCAPS_API_KEY']}"}
def call(method, path, **kw):
r = requests.request(method, API + path, headers=HEADERS, timeout=60, **kw)
if r.status_code >= 400:
raise SystemExit(r.json()["error"]["message"])
return r.json()
upload = call("POST", "/files/upload-url")
with open(sys.argv[1], "rb") as f:
requests.put(upload["upload_url"], data=f, timeout=600).raise_for_status()
job = call("POST", "/transcripts", json={"model": "a5s-async-v1", "file_id": upload["id"], "language_code": "en"})
while job["status"] in ("queued", "processing"):
time.sleep(3)
job = call("GET", f"/transcripts/{job['id']}")
if job["status"] == "error":
raise SystemExit(job["error"]["message"])
for u in job["utterances"]:
print(f"[{u['start']:8.1f}s] {u['speaker']}: {u['text']}")
```
```javascript Node.js theme={null}
// node transcribe.mjs meeting.mp3 (Node 18+)
import { readFile } from "node:fs/promises";
const API = "https://api.aircaps.com/v1";
const headers = { Authorization: `Bearer ${process.env.AIRCAPS_API_KEY}` };
async function call(method, path, body) {
const r = await fetch(API + path, {
method,
headers: body ? { ...headers, "Content-Type": "application/json" } : headers,
body: body ? JSON.stringify(body) : undefined,
});
const data = await r.json();
if (!r.ok) throw new Error(data.error.message);
return data;
}
const upload = await call("POST", "/files/upload-url");
const put = await fetch(upload.upload_url, { method: "PUT", body: await readFile(process.argv[2]) });
if (!put.ok) throw new Error(`upload failed: ${put.status}`);
let job = await call("POST", "/transcripts", { model: "a5s-async-v1", file_id: upload.id, language_code: "en" });
while (job.status === "queued" || job.status === "processing") {
await new Promise((r) => setTimeout(r, 3000));
job = await call("GET", `/transcripts/${job.id}`);
}
if (job.status === "error") throw new Error(job.error.message);
for (const u of job.utterances) console.log(`[${u.start.toFixed(1)}s] ${u.speaker}: ${u.text}`);
```
```bash cURL theme={null}
# 1. upload URL
curl -X POST https://api.aircaps.com/v1/files/upload-url -H "Authorization: Bearer $AIRCAPS_API_KEY"
# 2. upload the bytes to the returned upload_url
curl -X PUT --upload-file meeting.mp3 ""
# 3. transcribe with the returned id
curl https://api.aircaps.com/v1/transcripts -H "Authorization: Bearer $AIRCAPS_API_KEY" \
-H "Content-Type: application/json" -d '{"model": "a5s-async-v1", "file_id": "file_...", "language_code": "en"}'
```
## Stream a WAV file
```python theme={null}
# pip install websockets==15.0.1
# AIRCAPS_API_KEY=aircaps_sk_... python stream.py audio.wav (mono 16 kHz PCM16 WAV)
import asyncio, json, os, sys, time, wave
import websockets
async def main():
headers = {"Authorization": f"Bearer {os.environ['AIRCAPS_API_KEY']}"}
async with websockets.connect("wss://api.aircaps.com/v1/realtime", additional_headers=headers) as ws:
await ws.send(json.dumps({"type": "start", "providers": ["a5sv2"], "encoding": "pcm_s16le",
"sample_rate": 16000, "channels": 1, "language": "en-US"}))
while (event := json.loads(await ws.recv()))["type"] != "ready":
if event["type"] == "error":
raise RuntimeError(event["message"])
async def receive():
final = ""
async for message in ws:
event = json.loads(message)
if event["type"] == "transcript" and event["is_final"]:
# text is the whole transcript so far: print only the newly settled part
text = event["text"]
print((text[len(final):] if text.startswith(final) else text).strip(), flush=True)
final = text
elif event["type"] in ("session_limit", "error"):
print(event["message"], flush=True)
receiver = asyncio.create_task(receive())
try:
with wave.open(sys.argv[1], "rb") as audio:
start, sent = time.monotonic(), 0
while chunk := audio.readframes(320): # 20 ms
await ws.send(chunk)
sent += len(chunk) // 2
await asyncio.sleep(max(0, start + sent / 16000 - time.monotonic()))
await ws.send(json.dumps({"type": "finish"}))
except websockets.ConnectionClosedOK:
pass # the server ended the session (5-minute limit)
await receiver
asyncio.run(main())
```
Convert any file to the right WAV format with `ffmpeg -i input.mp3 -ac 1 -ar 16000 -sample_fmt s16 audio.wav`.
## Verify a webhook
```python theme={null}
# pip install flask standardwebhooks
import os
from flask import Flask, request
from standardwebhooks.webhooks import Webhook, WebhookVerificationError
app = Flask(__name__)
wh = Webhook(os.environ["AIRCAPS_WEBHOOK_SECRET"])
@app.post("/hooks/aircaps")
def hook():
try:
event = wh.verify(request.get_data(), dict(request.headers))
except WebhookVerificationError:
return "", 400
print(event["event"], event["id"]) # fetch the transcript with GET /v1/transcripts/{id}
return "", 204
```
# Live streaming
Source: https://docs.aircaps.com/guides/streaming
Real-time captions with A5S v2 Streaming over WebSocket.
Endpoint: `wss://api.aircaps.com/v1/realtime`
## Protocol
Open the WebSocket with your key in the handshake header: `Authorization: Bearer aircaps_sk_...`
The first message is JSON. `language` accepts `en` and regional variants such as `en-US`; multilingual support is planned.
```json theme={null}
{
"type": "start",
"providers": ["a5sv2"],
"encoding": "pcm_s16le",
"sample_rate": 16000,
"channels": 1,
"language": "en-US"
}
```
The server may send `provider_status` messages while the session starts, then `ready`. Do not send audio before `ready`.
Send binary frames of mono 16 kHz signed 16-bit little-endian PCM, in real time. 20 ms frames (640 bytes) are recommended; frames may be at most 64 KB. Send `{"type": "ping"}` to keep the connection open while you have no audio to send.
Send `{"type": "finish"}`. The server sends the final transcript, then a `finished` message, and closes.
## Messages from the server
```json theme={null}
{"type": "transcript", "provider": "a5sv2", "text": "The complete transcript so far.", "is_final": false}
```
`text` is always the complete transcript so far; replace what you display rather than appending.
| type | Meaning |
| - | - |
| `provider_status` | Startup progress (`message`) |
| `ready` | Start sending audio. Includes `session_limit_ms` |
| `transcript` | Current text; `is_final` is true when a segment is settled |
| `session_limit` | The 5-minute session or your usage limit was reached; the final transcript and `finished` follow |
| `provider_error` | Streaming failed during the session (`message`); an `error` follows |
| `error` | `code` and `message`; the connection then closes |
| `finished` | `audio_ms` streamed and `billed_ms` counted toward usage |
| `pong` | Reply to `ping` |
## Custom vocabulary
Bias recognition toward names and terms with up to 64 entries. Each needs a `text` (up to 100 characters) and a `strength` from 0 to 12:
```json theme={null}
{
"type": "start",
"providers": ["a5sv2"],
"custom_vocabulary_enabled": true,
"custom_vocabulary": [{"text": "AirCaps", "strength": 6}]
}
```
## Rules
* Send the `start` message within 45 seconds of connecting. It accepts only the documented fields.
* Audio must arrive in real time; sending faster (beyond a 2-second burst) ends the session with `realtime_rate_exceeded`.
* A session ends after 5 minutes (`session_limit`); open a new one to continue.
* No messages (audio or `ping`) for 45 seconds ends the session with `idle_timeout`.
* At most 5 sessions per account at once.
## Errors
Every session that ends abnormally sends an `error` message with a `code` before closing. **Act on that `code`.** The WebSocket close code is also set, but proxies between you and AirCaps may not preserve it.
| code | Close code | Meaning |
| - | - | - |
| `invalid_start`, `unsupported_language`, `unknown_provider` | 4400 | Bad or missing `start` message |
| `audio_before_ready`, `invalid_audio_frame`, `invalid_message` | 4400 | Audio sent before `ready`, an empty or odd-length audio frame, or a text message other than `ping`/`finish` |
| `unauthorized` | 4401 | Missing or invalid API key |
| `account_disabled`, `model_access_denied`, `origin_not_allowed` | 4403 | Account disabled, no access to A5S v2 Streaming, or browser origin not allowed |
| `idle_timeout` | 4408 | No messages for 45 seconds |
| `quota_unavailable`, `realtime_rate_exceeded`, `frame_too_large` | 4409 | No usage left or too many live sessions; audio faster than real time; a frame over 64 KB ([talk to sales](https://research.aircaps.com/contact) to raise limits) |
| `a5sv2_unavailable`, `upstream_unavailable` | 1013 | Temporarily unavailable; retry with backoff |
| `internal_error` | 1011 | Server error |
A normal session ends with `finished` and close code 1000.
See the [Python sample](/guides/samples#stream-a-wav-file).
# Webhooks
Source: https://docs.aircaps.com/guides/webhooks
Get notified when a transcript finishes.
Pass `webhook_url` when you create a transcript. When it finishes, AirCaps sends a `POST`:
```json theme={null}
{"event": "transcript.completed", "id": "tr_01k6z...", "status": "completed", "metadata": {"meeting_id": 42}}
```
`event` is `transcript.completed` or `transcript.failed`. The payload never contains the transcript; fetch it with [`GET /v1/transcripts/{id}`](/api-reference/get-transcript).
* To add your own header to each delivery, set `webhook_auth_header_name` and `webhook_auth_header_value` on the request.
* `webhook_url` must be a public http(s) address. Redirects are not followed.
* Canceled transcripts send no webhook.
## Verify signatures
Deliveries are signed per [Standard Webhooks](https://www.standardwebhooks.com) with your account's webhook secret (`whsec_...`). Find and rotate it in the console under **API keys → Webhook signing secret**.
Headers: `webhook-id`, `webhook-timestamp`, `webhook-signature: v1,`.
```python theme={null}
# pip install standardwebhooks
import os
from standardwebhooks.webhooks import Webhook, WebhookVerificationError
wh = Webhook(os.environ["AIRCAPS_WEBHOOK_SECRET"])
try:
event = wh.verify(request_body_bytes, request_headers) # raw body bytes and the request headers
except WebhookVerificationError:
... # reject with 400
```
Manual check:
1. Base64-decode the secret without its `whsec_` prefix.
2. Compute HMAC-SHA256 with it over `{webhook-id}.{webhook-timestamp}.{raw body}` and base64-encode the result.
3. Compare it in constant time with the value after `v1,` in `webhook-signature`.
4. Reject timestamps more than 5 minutes old.
## Delivery
* Reply with any 2xx within 15 seconds.
* Failures (including timeouts, network errors and 3xx) are retried for about 24 hours: after 30 s, 2 min, 10 min, 30 min, 1 h, 2 h, 4 h, 8 h and 8 h.
* A 4xx other than 408 or 429 stops retries.
* `webhook-id` is stable across retries of one event; use it to deduplicate.
* The transcript shows `webhook_status` (`pending`, `delivered` or `failed`) and the last HTTP status as `webhook_status_code`.
* Rotating the secret applies to pending retries too.
# AirCaps API
Source: https://docs.aircaps.com/index
Speech-to-text for English: live streaming and speaker-labelled batch transcription.
The AirCaps API turns English speech into text with two models:
| Model | ID | Use it for |
| - | - | - |
| **A5S Async v1** | `a5s-async-v1` | Recorded audio and video: speaker labels, word timestamps, subtitles |
| **A5S v2 Streaming** | `a5s-v2-streaming` | Live audio: captions as people speak |
One API key works for both. Base URL: `https://api.aircaps.com/v1`
**Every account includes 5 hours of free usage**, shared by both models. For pricing, higher usage limits or concurrency, or a dedicated deployment, [talk to sales](https://research.aircaps.com/contact).
We don't train on your data; see [Data & privacy](/data-privacy).
Both models transcribe **English** today. Multilingual support is planned and will be available soon.
Transcribe a file in five minutes.
Create and use a key.
Uploads, URLs, polling and webhooks.
The realtime WebSocket protocol.
Pricing, higher limits and concurrency, model access, and dedicated deployments.
Try both models without code in the [Playground](https://playground.aircaps.com/playground).
# Limits
Source: https://docs.aircaps.com/limits
Usage limit, concurrency and rate limits.
## Usage limit
Each account includes **5 hours of free usage**, shared by both models.
* **A5S Async v1** counts the duration of each completed transcript. Failed jobs are not counted.
* **A5S v2 Streaming** counts the audio you stream.
* A file longer than your remaining usage fails with `usage_limit_reached` and is not counted. New requests at a used-up limit return `403 usage_limit_reached`.
* See what's left in the console or with [`GET /v1/account`](/api-reference/account) (`usage_limit`).
## Defaults
| Limit | Default |
| - | - |
| A5S Async v1 jobs processing at once | 5 (more are queued, first in first out) |
| A5S Async v1 jobs waiting in the queue | 2,000 |
| Transcript create requests (`POST /v1/transcripts`) | 600 per minute |
| Maximum file duration / size | 10 hours / 5 GB |
| Direct upload (`POST /v1/files`) | 1 GB (use an upload URL for larger files) |
| Uploaded files | Deleted when the transcript that uses them finishes; unused uploads after 24 hours |
| Transcripts kept (including failed ones) | 30 days, or until you delete them ([Data & privacy](/data-privacy)) |
| A5S v2 Streaming sessions at once | 5 |
| Streaming session length | 5 minutes |
| Active API keys | 10 |
Jobs beyond the processing limit wait as `queued`. Beyond the queue or request-rate limit you get `429` with a `Retry-After` header. Other endpoints are not rate limited.
## Pricing and higher limits
Limits are set per account. For pricing, higher usage limits or concurrency, access to a model, or a dedicated deployment, [talk to sales](https://research.aircaps.com/contact).
research.aircaps.com/contact
# Models
Source: https://docs.aircaps.com/models
What each model does and when to use it.
## A5S Async v1
`a5s-async-v1` · REST · [batch guide](/guides/async-transcription)
* Recorded audio or video in any common format, from 0.1 s to 10 hours.
* Speaker labels (`speaker_0`, `speaker_1`, …), word timestamps, utterances, SRT/VTT subtitles.
* English only: `language_code` accepts `en` and regional variants (`en-US`, `en-GB`, …); any other value returns `400 unsupported_language`.
* Text is verbatim and lowercase without punctuation (apostrophes are kept, as in "let's"); disfluencies ("uh", "um") are kept.
## A5S v2 Streaming
`a5s-v2-streaming` · WebSocket · [streaming guide](/guides/streaming)
* Live 16 kHz mono PCM16 audio in, partial and final text out as people speak.
* Punctuated, cased text. No speaker labels.
* Optional custom vocabulary to bias recognition toward names and terms.
* English only: the start message's `language` accepts `en` and regional variants (`en-US`, `en-GB`, …); any other value is refused with `unsupported_language`.
* Sessions last up to 5 minutes.
## Choosing
| You have | Use |
| - | - |
| Recordings, uploads, meeting files | A5S Async v1 |
| A live microphone or call that needs captions now | A5S v2 Streaming |
| Live audio that later needs speaker labels | Stream for captions, then send the recording to A5S Async v1 |
## Languages
Both models support English today. **Multilingual support is planned and will be available soon.** Accepted codes: `en`, `en-US`, `en-GB`, `en-AU`, `en-CA`, `en-IN`, `en-IE`, `en-NZ`, `en-ZA` (case-insensitive; `_` works too).
## Access
Accounts can use both models by default. An account may be limited to one model. Requests to a model your account can't use return `403 model_access_denied` (streaming: an `error` message with that code, then close code `4403`). `GET /v1/account` shows `enabled` for each model. To get access, [talk to sales](https://research.aircaps.com/contact).
`GET /v1/models` lists the models.
# Quickstart
Source: https://docs.aircaps.com/quickstart
Transcribe a recording with speaker labels.
Sign in at [playground.aircaps.com](https://playground.aircaps.com), open **API keys**, and create a key. It is shown once.
```bash theme={null}
export AIRCAPS_API_KEY="aircaps_sk_..."
```
Pass a URL the API can download (public or presigned):
```bash theme={null}
curl https://api.aircaps.com/v1/transcripts \
-H "Authorization: Bearer $AIRCAPS_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model": "a5s-async-v1", "audio_url": "https://example.com/meeting.mp3", "language_code": "en"}'
```
`language_code` is optional (default `en`).
The response is a transcript object with `"status": "queued"` or `"processing"` and an `id` like `tr_01k6z...`.
```bash theme={null}
curl https://api.aircaps.com/v1/transcripts/tr_01k6z... \
-H "Authorization: Bearer $AIRCAPS_API_KEY"
```
Poll until `status` is `completed` (or `error`), or pass a `webhook_url` when you submit to be notified instead.
## Local files in Python
```python theme={null}
# pip install requests
import os, sys, time, requests
API = "https://api.aircaps.com/v1"
HEADERS = {"Authorization": f"Bearer {os.environ['AIRCAPS_API_KEY']}"}
def call(method, path, **kw):
r = requests.request(method, API + path, headers=HEADERS, timeout=60, **kw)
if r.status_code >= 400:
raise SystemExit(r.json()["error"]["message"])
return r.json()
# 1. Upload straight to storage
upload = call("POST", "/files/upload-url")
with open(sys.argv[1], "rb") as f:
requests.put(upload["upload_url"], data=f, timeout=600).raise_for_status()
# 2. Transcribe
job = call("POST", "/transcripts", json={"model": "a5s-async-v1", "file_id": upload["id"], "language_code": "en"})
# 3. Wait for the result
while job["status"] in ("queued", "processing"):
time.sleep(3)
job = call("GET", f"/transcripts/{job['id']}")
if job["status"] == "error":
raise SystemExit(job["error"]["message"])
for u in job["utterances"]:
print(f"[{u['start']:7.1f}s] {u['speaker']}: {u['text']}")
```
**Language:** both models transcribe English only today. Pass `language_code` (batch) or `language` (streaming) as `en` or a regional variant such as `en-US` or `en-GB`. Any other value returns `unsupported_language`. Multilingual support is planned and will be available soon.
Next: [batch transcription in depth](/guides/async-transcription) or [live streaming](/guides/streaming).