# Get account Source: https://docs.aircaps.com/api-reference/account GET /v1/account Your usage limit, per-model access (`enabled`) and limits. For higher limits or model access, [talk to sales](https://research.aircaps.com/contact). # Create a transcript Source: https://docs.aircaps.com/api-reference/create-transcript POST /v1/transcripts Start transcribing audio with A5S Async v1. The job runs asynchronously: poll [Get a transcript](/api-reference/get-transcript) or pass `webhook_url`. # Delete a transcript Source: https://docs.aircaps.com/api-reference/delete-transcript DELETE /v1/transcripts/{id} Cancels a `queued` or `processing` job (its `error.code` becomes `canceled`), or deletes a finished transcript and its stored data. Afterwards the ID returns `404`. # Get a transcript Source: https://docs.aircaps.com/api-reference/get-transcript GET /v1/transcripts/{id} Status and, once `completed`, the result. # List transcripts Source: https://docs.aircaps.com/api-reference/list-transcripts GET /v1/transcripts Your transcripts, newest first. Items omit the result fields (`text`, `speakers`, `utterances`, `words`). # List models Source: https://docs.aircaps.com/api-reference/models GET /v1/models The available models. No authentication needed. # Error object Source: https://docs.aircaps.com/api-reference/objects/error Every HTTP error uses this envelope. See [Errors](/errors) for all codes. # File object Source: https://docs.aircaps.com/api-reference/objects/file An uploaded file. Use id as file_id for one transcript; the file is deleted when that transcript finishes, or after 24 hours if unused. # Transcript object Source: https://docs.aircaps.com/api-reference/objects/transcript Returned by every transcript endpoint. The result fields (text, speakers, utterances, words) appear only in Get a transcript, once status is completed. # Utterance object Source: https://docs.aircaps.com/api-reference/objects/utterance One speaker turn. Times are seconds from the start of the audio. # Webhook event Source: https://docs.aircaps.com/api-reference/objects/webhook-event The body AirCaps `POST`s to your `webhook_url`. See [Webhooks](/guides/webhooks) for signing and retries. # Word object Source: https://docs.aircaps.com/api-reference/objects/word One word with its timing and speaker. # Overview Source: https://docs.aircaps.com/api-reference/overview Base URL, authentication and conventions. * **Base URL:** `https://api.aircaps.com/v1` * **Auth:** `Authorization: Bearer aircaps_sk_...` on every request ([API keys](/authentication)) * **Format:** JSON request and response bodies; timestamps are RFC 3339 UTC; audio times are seconds. * **IDs:** `tr_...` transcripts, `file_...` uploads, `req_...` request IDs (also in the `x-request-id` header). * **Errors:** one envelope for all errors ([Errors](/errors)). * **Idempotency:** send `Idempotency-Key` with `POST /v1/transcripts` to make retries safe. * **Language:** English only today: `language_code` (batch) and `language` (streaming) accept `en` and regional variants (`en-US`, `en-GB`, …); anything else returns `unsupported_language`. Multilingual support is planned. * **Pagination:** list endpoints return `has_more` and `next_before`; pass it as `before` for the next page. | Endpoint | Purpose | | - | - | | `POST /v1/transcripts` | Create a transcript | | `GET /v1/transcripts/{id}` | Get status and result | | `GET /v1/transcripts` | List transcripts | | `DELETE /v1/transcripts/{id}` | Cancel or delete | | `GET /v1/transcripts/{id}/subtitles` | SRT or VTT captions | | `POST /v1/files/upload-url` | Presigned upload (≤ 5 GB) | | `POST /v1/files` | Direct upload (≤ 1 GB) | | `WSS /v1/realtime` | Live streaming | | `GET /v1/models` | Available models | | `GET /v1/account` | Account and limits | | `GET /v1/usage` | Daily usage | **Objects:** [Transcript](/api-reference/objects/transcript), [Utterance](/api-reference/objects/utterance), [Word](/api-reference/objects/word), [File](/api-reference/objects/file), [Error](/api-reference/objects/error), [Webhook event](/api-reference/objects/webhook-event). The full OpenAPI 3.1 spec is at [`/openapi.json`](/openapi.json). # Realtime stream Source: https://docs.aircaps.com/api-reference/realtime Live transcription with A5S v2 Streaming over a WebSocket. **WSS** `wss://api.aircaps.com/v1/realtime` Requires access to A5S v2 Streaming. For a walkthrough, see [Live streaming](/guides/streaming). ## Connection `Bearer aircaps_sk_...`, sent with the WebSocket handshake. The flow: connect, send `start`, wait for `ready`, stream binary audio, send `finish`, then receive `finished`. The server then closes with `1000`. ## Client messages ### start The first message, as a JSON text frame, within 45 seconds of connecting. Unknown fields are rejected. `"start"` Language of the audio. English only today: `en` or a regional variant (`en-US`, `en-GB`, `en-AU`, `en-CA`, `en-IN`, `en-IE`, `en-NZ`, `en-ZA`). Any other value is refused with `unsupported_language`. Multilingual support is planned. Only `pcm_s16le` (signed 16-bit little-endian PCM). Only `16000`. Only `1` (mono). The streaming engine. Only `"a5sv2"`. Bias recognition toward `custom_vocabulary`. Up to 64 terms. A word or phrase, up to 100 characters. Bias strength, 0 to 12. ```json Example theme={null} {"type": "start", "language": "en-US", "encoding": "pcm_s16le", "sample_rate": 16000, "channels": 1, "custom_vocabulary_enabled": true, "custom_vocabulary": [{"text": "AirCaps", "strength": 6}]} ``` ### Audio Binary frames of raw PCM in the `start` format: non-empty, an even number of bytes, at most 64 KB (larger frames end the session with `frame_too_large`; frames over 1 MB close the connection without a message). Send them after `ready`, in real time; a 2-second burst is allowed. 20 ms frames (640 bytes) are recommended. ### ping `{"type": "ping"}`. The server answers `pong`. Any message, audio or `ping`, resets the 45-second idle timer. ### finish `{"type": "finish"}`. The server flushes the final transcript, sends `finished`, and closes. ## Server messages Each is a JSON text frame with a `type`. ### provider\_status Startup progress while the session starts. For example `"Getting things ready, please hold on"`. ### ready Start sending audio. Maximum audio for this session: 5 minutes, or less if little usage remains. The account's usage limit, in milliseconds. Usage so far, in milliseconds. `"en-US"` `16000` `"pcm_s16le"` ### transcript The complete transcript so far, punctuated and cased. Replace what you display instead of appending. `true` when the current segment is settled. `"a5sv2"` ```json Example theme={null} {"type": "transcript", "provider": "a5sv2", "text": "Good afternoon, everybody.", "is_final": true} ``` ### session\_limit The session reached 5 minutes or the account's usage limit. The final transcript and `finished` follow, then close code 1000. Why the session ended. ### finished Audio received. Audio counted toward usage. Account usage after this session. The account's usage limit. ### provider\_error Streaming failed during the session. An `error` follows. What failed. ### error Sent before the server ends the session abnormally. **Act on `code`**: the WebSocket close code is also set, but proxies may not preserve it. One of the codes below. Human-readable detail. | code | Close | Meaning | | - | - | - | | `invalid_start` | 4400 | The first message is missing, late or invalid | | `unsupported_language` | 4400 | `language` is not English | | `unknown_provider` | 4400 | `providers` names an unknown engine | | `audio_before_ready` | 4400 | Audio sent before `ready` | | `invalid_audio_frame` | 4400 | An empty or odd-length audio frame | | `invalid_message` | 4400 | A text message other than `ping` or `finish` | | `unauthorized` | 4401 | Missing or invalid API key | | `account_disabled` | 4403 | The account is disabled | | `model_access_denied` | 4403 | No access to A5S v2 Streaming ([talk to sales](https://research.aircaps.com/contact)) | | `origin_not_allowed` | 4403 | Browser origin not allowed | | `idle_timeout` | 4408 | No messages for 45 seconds | | `quota_unavailable` | 4409 | Usage limit used up, or too many live sessions ([talk to sales](https://research.aircaps.com/contact)) | | `realtime_rate_exceeded` | 4409 | Audio sent faster than real time | | `frame_too_large` | 4409 | An audio frame over 64 KB | | `a5sv2_unavailable`, `upstream_unavailable` | 1013 | Temporarily unavailable; retry | | `internal_error` | 1011 | Server error | ### pong The reply to `ping`. ## Close codes | Code | Meaning | | - | - | | 1000 | Normal close after `finished` | | 1011 | Server error | | 1013 | Temporarily unavailable; retry with backoff | | 4400 | Invalid message or frame | | 4401 | Authentication failed | | 4403 | Forbidden: account disabled, no access to the model, or browser origin not allowed | | 4408 | Idle timeout | | 4409 | Limit reached at start (usage, sessions), audio faster than real time, or a frame over 64 KB | Close codes are informational; use the `error` message's `code`. # Get subtitles Source: https://docs.aircaps.com/api-reference/subtitles GET /v1/transcripts/{id}/subtitles SRT or WebVTT captions for a completed transcript, with the speaker in each cue. # Upload a file Source: https://docs.aircaps.com/api-reference/upload-file POST /v1/files Upload up to 1 GB in the request body. For larger files (up to 5 GB), use an upload URL. # Create an upload URL Source: https://docs.aircaps.com/api-reference/upload-url POST /v1/files/upload-url Get a presigned URL to upload a file of up to 5 GB straight to storage. `PUT` the raw bytes to `upload_url` within an hour, then use `id` as `file_id`. # Get usage Source: https://docs.aircaps.com/api-reference/usage GET /v1/usage Daily usage of both models. # API keys Source: https://docs.aircaps.com/authentication Create, use and revoke keys. Every request needs an API key in the `Authorization` header: ```http theme={null} Authorization: Bearer aircaps_sk_... ``` * **Create keys** in the console: [playground.aircaps.com → API keys](https://playground.aircaps.com/api-keys). A key is shown once; AirCaps stores only a hash. * **One key, both models.** The same key works for the REST API and the streaming WebSocket. * **Up to 10 active keys** per account. Use one per service or environment so you can revoke them independently. * **Revoke** a key in the console. It stops working within 60 seconds; a streaming session already open with it runs until it ends. * **Keep keys server-side.** Never put a key in a browser, mobile app or public repository. Requests without a valid key return `401 unauthorized`; a disabled account returns `403 account_disabled`. On the streaming WebSocket these arrive as an `error` message with the same `code`. Account settings (keys, webhook signing secret, profile) can only be changed from the console, not with an API key. # Data & privacy Source: https://docs.aircaps.com/data-privacy What AirCaps keeps, for how long, and how to delete it. **We don't use your audio or transcripts to train or improve models.** Your data is used only to process your requests. ## What we keep | Data | Kept for | | - | - | | **Streaming audio and transcripts** (A5S v2 Streaming) | Not stored. Audio is processed in memory and the text is sent only to you. | | **Uploaded files** (`POST /v1/files` or an upload URL) | Deleted as soon as the transcript that uses them finishes (completed, failed or canceled). Unused uploads are deleted after 24 hours. | | **Audio from `audio_url`** | Downloaded only while the transcript is processed, then deleted. | | **Processing copies** of the audio | Deleted when the transcript finishes. | | **Transcripts** (text, words, speakers) | 30 days after completion, or until you delete them with [`DELETE /v1/transcripts/{id}`](/api-reference/delete-transcript). | | **Transcript details you send** (`metadata`, `audio_url`, webhook URL) | Removed when the transcript is deleted or expires. | | **Webhook auth header** (`webhook_auth_header_value`) | Removed as soon as the webhook is delivered or delivery fails for good. | | **Usage records** (duration, status, timestamps) | Kept for usage reporting and billing; they contain no audio, text or metadata. | | **Account** (email, name, limits) | For the life of the account. API keys are stored only as SHA-256 hashes. | | **Logs** | Requests (endpoint, status, timing and IDs). Logs don't contain audio or transcripts. | ## Deleting data * **A transcript:** `DELETE /v1/transcripts/{id}` deletes it immediately. If it is still processing, it is canceled. * **Uploads:** deleted automatically, as above. * **Your account and all its data:** [contact us](https://research.aircaps.com/contact). ## Security * **In transit:** every connection uses TLS: your client to the API, between our services, to storage and to our database (with certificate verification). * **At rest:** stored data is encrypted with AES-256. * **Isolation:** every request is scoped to your account; no account can read another's data. * **Keys and webhooks:** API keys are stored as hashes and can be revoked at any time; webhooks are signed ([Standard Webhooks](/guides/webhooks)). ## Where data is processed In the United States, using these infrastructure providers: | Provider | Used for | | - | - | | Modal | Running the API and the models | | Cloudflare (R2) | Storing uploads and transcripts | | Supabase | Account and transcript records | | Render | The public API endpoint (`api.aircaps.com`) and the console | For a data processing agreement, custom retention or a dedicated deployment, [talk to sales](https://research.aircaps.com/contact). # Errors Source: https://docs.aircaps.com/errors HTTP errors, job errors and retries. ## HTTP errors Every error uses one shape: ```json theme={null} { "error": { "type": "invalid_request_error", "code": "unsupported_language", "message": "language 'de' is not supported; only English (\"en\", \"en-US\", ...) is available today. Multilingual support is planned.", "param": "language_code", "request_id": "req_01k6z..." } } ``` | HTTP | type | code | | - | - | - | | 400 | `invalid_request_error` | `invalid_request`, `unsupported_language`, `unsupported_model` | | 401 | `authentication_error` | `unauthorized` | | 403 | `permission_error` | `account_suspended`, `account_disabled`, `session_required`, `usage_limit_reached`, `model_access_denied` | | 404 | `invalid_request_error` | `not_found`, `file_not_found` | | 405 | `invalid_request_error` | `method_not_allowed` | | 409 | `invalid_request_error` | `idempotency_conflict`, `transcript_not_ready` | | 413 | `invalid_request_error` | `file_too_large` | | 429 | `rate_limit_error` | `rate_limited`, `queue_full` (honour `Retry-After`) | | 500 | `api_error` | `internal_error` | | 502, 504 | `api_error` | `upstream_unavailable`, `upstream_timeout` (`request_id` is `null`) | Include `request_id` (also in the `x-request-id` header) when you contact support. * `unsupported_language`: only English is available today; multilingual support is planned. * `usage_limit_reached`, `model_access_denied`, `queue_full`: account limits. To raise them, [talk to sales](https://research.aircaps.com/contact). ## Job errors A transcript that fails has `"status": "error"` and an `error` object. None of these count toward your usage. | code | Meaning | | - | - | | `invalid_audio` | Not decodable, no audio stream, empty, or shorter than 0.1 s | | `audio_too_long` | Longer than 10 hours | | `file_too_large` | Larger than 5 GB | | `download_failed` | `audio_url` could not be fetched (HTTP error, timeout, non-public address) | | `file_not_found` | The uploaded file is missing or expired | | `usage_limit_reached` | Longer than your remaining usage ([talk to sales](https://research.aircaps.com/contact) to raise it) | | `internal_error` | Failed on our side after automatic retries. Resubmit; if you used `file_id`, upload the file again (uploads are deleted when their transcript finishes) | Deleting a queued or processing transcript cancels it: the `DELETE` response shows `error.code` `canceled`, no webhook is sent, and the ID then returns `404`. ## Retries * Retry `429`, `500`, `502` and `504` with exponential backoff. * Send an `Idempotency-Key` header with `POST /v1/transcripts` so a retried request never creates a second job. A retry with the same key and body returns the original transcript (`200`), even after it has finished. # Batch transcription Source: https://docs.aircaps.com/guides/async-transcription Transcribe recordings with A5S Async v1. ## 1. Provide the audio Choose one: | Source | How | Size | | - | - | - | | A URL | `audio_url` in the create request. Public or presigned http(s) URL. | ≤ 5 GB | | Upload URL | [`POST /v1/files/upload-url`](/api-reference/upload-url), then `PUT` the bytes to the returned `upload_url` within an hour. Use the returned `id` as `file_id`. | ≤ 5 GB | | Direct upload | [`POST /v1/files`](/api-reference/upload-file) with the raw bytes as the body. | ≤ 1 GB | Any audio or video format ffmpeg can read works. The first audio stream is used; channels are mixed down. * **Uploaded files are deleted as soon as the transcript that uses them finishes** (completed, failed or canceled). A `file_id` is for one transcript; to transcribe the same file again, upload it again. Unused uploads are deleted after 24 hours. * An `audio_url` must be public or presigned, without credentials in the URL. Up to 5 redirects are followed (each must stay on a public address), and the download must finish within 15 minutes. ## 2. Create the transcript ```bash theme={null} curl https://api.aircaps.com/v1/transcripts \ -H "Authorization: Bearer $AIRCAPS_API_KEY" \ -H "Content-Type: application/json" \ -H "Idempotency-Key: meeting-2026-10-06" \ -d '{ "model": "a5s-async-v1", "file_id": "file_01k6z...", "language_code": "en", "webhook_url": "https://example.com/hooks/aircaps", "metadata": {"meeting_id": 42} }' ``` `metadata` (up to 16 KB) is echoed back on the transcript and in webhooks. `language_code` is optional and accepts English (`en`, `en-US`, …) only; multilingual support is planned. ## 3. Get the result Wait for a [webhook](/guides/webhooks) or poll [`GET /v1/transcripts/{id}`](/api-reference/get-transcript) every few seconds. | status | Meaning | | - | - | | `queued` | Waiting for one of your account's processing slots | | `processing` | Being transcribed | | `completed` | Done; result fields are present | | `error` | Failed; see `error` | ## The result ```json theme={null} { "id": "tr_01k6z3m4q8w7e5r2t1y0u9i8o7", "status": "completed", "model": "a5s-async-v1", "audio_duration": 3600.512, "text": "hey good afternoon everybody ...", "speakers": [{"id": "speaker_0", "speaking_time": 1402.11}, {"id": "speaker_1", "speaking_time": 988.4}], "utterances": [ {"speaker": "speaker_0", "start": 0.48, "end": 6.12, "text": "hey good afternoon everybody ...", "words": [{"text": "hey", "start": 0.48, "end": 0.71, "speaker": "speaker_0"}]} ], "words": [{"text": "hey", "start": 0.48, "end": 0.71, "speaker": "speaker_0"}] } ``` * Times are seconds from the start of the audio. * Speakers are numbered in order of first speech and are consistent within one transcript only. * Overlapping speech produces overlapping utterances. * Audio with no speech completes with empty `text`, `utterances` and `words`. Need captions? [`GET /v1/transcripts/{id}/subtitles?format=srt`](/api-reference/subtitles). ## Delete [`DELETE /v1/transcripts/{id}`](/api-reference/delete-transcript) cancels a running job or deletes a finished transcript and its stored result. Transcripts are otherwise deleted after 30 days. # Code samples Source: https://docs.aircaps.com/guides/samples Copy-paste examples in Python, Node.js and cURL. ## Transcribe a local file ```python Python theme={null} # pip install requests import os, sys, time, requests API = "https://api.aircaps.com/v1" HEADERS = {"Authorization": f"Bearer {os.environ['AIRCAPS_API_KEY']}"} def call(method, path, **kw): r = requests.request(method, API + path, headers=HEADERS, timeout=60, **kw) if r.status_code >= 400: raise SystemExit(r.json()["error"]["message"]) return r.json() upload = call("POST", "/files/upload-url") with open(sys.argv[1], "rb") as f: requests.put(upload["upload_url"], data=f, timeout=600).raise_for_status() job = call("POST", "/transcripts", json={"model": "a5s-async-v1", "file_id": upload["id"], "language_code": "en"}) while job["status"] in ("queued", "processing"): time.sleep(3) job = call("GET", f"/transcripts/{job['id']}") if job["status"] == "error": raise SystemExit(job["error"]["message"]) for u in job["utterances"]: print(f"[{u['start']:8.1f}s] {u['speaker']}: {u['text']}") ``` ```javascript Node.js theme={null} // node transcribe.mjs meeting.mp3 (Node 18+) import { readFile } from "node:fs/promises"; const API = "https://api.aircaps.com/v1"; const headers = { Authorization: `Bearer ${process.env.AIRCAPS_API_KEY}` }; async function call(method, path, body) { const r = await fetch(API + path, { method, headers: body ? { ...headers, "Content-Type": "application/json" } : headers, body: body ? JSON.stringify(body) : undefined, }); const data = await r.json(); if (!r.ok) throw new Error(data.error.message); return data; } const upload = await call("POST", "/files/upload-url"); const put = await fetch(upload.upload_url, { method: "PUT", body: await readFile(process.argv[2]) }); if (!put.ok) throw new Error(`upload failed: ${put.status}`); let job = await call("POST", "/transcripts", { model: "a5s-async-v1", file_id: upload.id, language_code: "en" }); while (job.status === "queued" || job.status === "processing") { await new Promise((r) => setTimeout(r, 3000)); job = await call("GET", `/transcripts/${job.id}`); } if (job.status === "error") throw new Error(job.error.message); for (const u of job.utterances) console.log(`[${u.start.toFixed(1)}s] ${u.speaker}: ${u.text}`); ``` ```bash cURL theme={null} # 1. upload URL curl -X POST https://api.aircaps.com/v1/files/upload-url -H "Authorization: Bearer $AIRCAPS_API_KEY" # 2. upload the bytes to the returned upload_url curl -X PUT --upload-file meeting.mp3 "" # 3. transcribe with the returned id curl https://api.aircaps.com/v1/transcripts -H "Authorization: Bearer $AIRCAPS_API_KEY" \ -H "Content-Type: application/json" -d '{"model": "a5s-async-v1", "file_id": "file_...", "language_code": "en"}' ``` ## Stream a WAV file ```python theme={null} # pip install websockets==15.0.1 # AIRCAPS_API_KEY=aircaps_sk_... python stream.py audio.wav (mono 16 kHz PCM16 WAV) import asyncio, json, os, sys, time, wave import websockets async def main(): headers = {"Authorization": f"Bearer {os.environ['AIRCAPS_API_KEY']}"} async with websockets.connect("wss://api.aircaps.com/v1/realtime", additional_headers=headers) as ws: await ws.send(json.dumps({"type": "start", "providers": ["a5sv2"], "encoding": "pcm_s16le", "sample_rate": 16000, "channels": 1, "language": "en-US"})) while (event := json.loads(await ws.recv()))["type"] != "ready": if event["type"] == "error": raise RuntimeError(event["message"]) async def receive(): final = "" async for message in ws: event = json.loads(message) if event["type"] == "transcript" and event["is_final"]: # text is the whole transcript so far: print only the newly settled part text = event["text"] print((text[len(final):] if text.startswith(final) else text).strip(), flush=True) final = text elif event["type"] in ("session_limit", "error"): print(event["message"], flush=True) receiver = asyncio.create_task(receive()) try: with wave.open(sys.argv[1], "rb") as audio: start, sent = time.monotonic(), 0 while chunk := audio.readframes(320): # 20 ms await ws.send(chunk) sent += len(chunk) // 2 await asyncio.sleep(max(0, start + sent / 16000 - time.monotonic())) await ws.send(json.dumps({"type": "finish"})) except websockets.ConnectionClosedOK: pass # the server ended the session (5-minute limit) await receiver asyncio.run(main()) ``` Convert any file to the right WAV format with `ffmpeg -i input.mp3 -ac 1 -ar 16000 -sample_fmt s16 audio.wav`. ## Verify a webhook ```python theme={null} # pip install flask standardwebhooks import os from flask import Flask, request from standardwebhooks.webhooks import Webhook, WebhookVerificationError app = Flask(__name__) wh = Webhook(os.environ["AIRCAPS_WEBHOOK_SECRET"]) @app.post("/hooks/aircaps") def hook(): try: event = wh.verify(request.get_data(), dict(request.headers)) except WebhookVerificationError: return "", 400 print(event["event"], event["id"]) # fetch the transcript with GET /v1/transcripts/{id} return "", 204 ``` # Live streaming Source: https://docs.aircaps.com/guides/streaming Real-time captions with A5S v2 Streaming over WebSocket. Endpoint: `wss://api.aircaps.com/v1/realtime` ## Protocol Open the WebSocket with your key in the handshake header: `Authorization: Bearer aircaps_sk_...` The first message is JSON. `language` accepts `en` and regional variants such as `en-US`; multilingual support is planned. ```json theme={null} { "type": "start", "providers": ["a5sv2"], "encoding": "pcm_s16le", "sample_rate": 16000, "channels": 1, "language": "en-US" } ``` The server may send `provider_status` messages while the session starts, then `ready`. Do not send audio before `ready`. Send binary frames of mono 16 kHz signed 16-bit little-endian PCM, in real time. 20 ms frames (640 bytes) are recommended; frames may be at most 64 KB. Send `{"type": "ping"}` to keep the connection open while you have no audio to send. Send `{"type": "finish"}`. The server sends the final transcript, then a `finished` message, and closes. ## Messages from the server ```json theme={null} {"type": "transcript", "provider": "a5sv2", "text": "The complete transcript so far.", "is_final": false} ``` `text` is always the complete transcript so far; replace what you display rather than appending. | type | Meaning | | - | - | | `provider_status` | Startup progress (`message`) | | `ready` | Start sending audio. Includes `session_limit_ms` | | `transcript` | Current text; `is_final` is true when a segment is settled | | `session_limit` | The 5-minute session or your usage limit was reached; the final transcript and `finished` follow | | `provider_error` | Streaming failed during the session (`message`); an `error` follows | | `error` | `code` and `message`; the connection then closes | | `finished` | `audio_ms` streamed and `billed_ms` counted toward usage | | `pong` | Reply to `ping` | ## Custom vocabulary Bias recognition toward names and terms with up to 64 entries. Each needs a `text` (up to 100 characters) and a `strength` from 0 to 12: ```json theme={null} { "type": "start", "providers": ["a5sv2"], "custom_vocabulary_enabled": true, "custom_vocabulary": [{"text": "AirCaps", "strength": 6}] } ``` ## Rules * Send the `start` message within 45 seconds of connecting. It accepts only the documented fields. * Audio must arrive in real time; sending faster (beyond a 2-second burst) ends the session with `realtime_rate_exceeded`. * A session ends after 5 minutes (`session_limit`); open a new one to continue. * No messages (audio or `ping`) for 45 seconds ends the session with `idle_timeout`. * At most 5 sessions per account at once. ## Errors Every session that ends abnormally sends an `error` message with a `code` before closing. **Act on that `code`.** The WebSocket close code is also set, but proxies between you and AirCaps may not preserve it. | code | Close code | Meaning | | - | - | - | | `invalid_start`, `unsupported_language`, `unknown_provider` | 4400 | Bad or missing `start` message | | `audio_before_ready`, `invalid_audio_frame`, `invalid_message` | 4400 | Audio sent before `ready`, an empty or odd-length audio frame, or a text message other than `ping`/`finish` | | `unauthorized` | 4401 | Missing or invalid API key | | `account_disabled`, `model_access_denied`, `origin_not_allowed` | 4403 | Account disabled, no access to A5S v2 Streaming, or browser origin not allowed | | `idle_timeout` | 4408 | No messages for 45 seconds | | `quota_unavailable`, `realtime_rate_exceeded`, `frame_too_large` | 4409 | No usage left or too many live sessions; audio faster than real time; a frame over 64 KB ([talk to sales](https://research.aircaps.com/contact) to raise limits) | | `a5sv2_unavailable`, `upstream_unavailable` | 1013 | Temporarily unavailable; retry with backoff | | `internal_error` | 1011 | Server error | A normal session ends with `finished` and close code 1000. See the [Python sample](/guides/samples#stream-a-wav-file). # Webhooks Source: https://docs.aircaps.com/guides/webhooks Get notified when a transcript finishes. Pass `webhook_url` when you create a transcript. When it finishes, AirCaps sends a `POST`: ```json theme={null} {"event": "transcript.completed", "id": "tr_01k6z...", "status": "completed", "metadata": {"meeting_id": 42}} ``` `event` is `transcript.completed` or `transcript.failed`. The payload never contains the transcript; fetch it with [`GET /v1/transcripts/{id}`](/api-reference/get-transcript). * To add your own header to each delivery, set `webhook_auth_header_name` and `webhook_auth_header_value` on the request. * `webhook_url` must be a public http(s) address. Redirects are not followed. * Canceled transcripts send no webhook. ## Verify signatures Deliveries are signed per [Standard Webhooks](https://www.standardwebhooks.com) with your account's webhook secret (`whsec_...`). Find and rotate it in the console under **API keys → Webhook signing secret**. Headers: `webhook-id`, `webhook-timestamp`, `webhook-signature: v1,`. ```python theme={null} # pip install standardwebhooks import os from standardwebhooks.webhooks import Webhook, WebhookVerificationError wh = Webhook(os.environ["AIRCAPS_WEBHOOK_SECRET"]) try: event = wh.verify(request_body_bytes, request_headers) # raw body bytes and the request headers except WebhookVerificationError: ... # reject with 400 ``` Manual check: 1. Base64-decode the secret without its `whsec_` prefix. 2. Compute HMAC-SHA256 with it over `{webhook-id}.{webhook-timestamp}.{raw body}` and base64-encode the result. 3. Compare it in constant time with the value after `v1,` in `webhook-signature`. 4. Reject timestamps more than 5 minutes old. ## Delivery * Reply with any 2xx within 15 seconds. * Failures (including timeouts, network errors and 3xx) are retried for about 24 hours: after 30 s, 2 min, 10 min, 30 min, 1 h, 2 h, 4 h, 8 h and 8 h. * A 4xx other than 408 or 429 stops retries. * `webhook-id` is stable across retries of one event; use it to deduplicate. * The transcript shows `webhook_status` (`pending`, `delivered` or `failed`) and the last HTTP status as `webhook_status_code`. * Rotating the secret applies to pending retries too. # AirCaps API Source: https://docs.aircaps.com/index Speech-to-text for English: live streaming and speaker-labelled batch transcription. The AirCaps API turns English speech into text with two models: | Model | ID | Use it for | | - | - | - | | **A5S Async v1** | `a5s-async-v1` | Recorded audio and video: speaker labels, word timestamps, subtitles | | **A5S v2 Streaming** | `a5s-v2-streaming` | Live audio: captions as people speak | One API key works for both. Base URL: `https://api.aircaps.com/v1` **Every account includes 5 hours of free usage**, shared by both models. For pricing, higher usage limits or concurrency, or a dedicated deployment, [talk to sales](https://research.aircaps.com/contact). We don't train on your data; see [Data & privacy](/data-privacy). Both models transcribe **English** today. Multilingual support is planned and will be available soon. Transcribe a file in five minutes. Create and use a key. Uploads, URLs, polling and webhooks. The realtime WebSocket protocol. Pricing, higher limits and concurrency, model access, and dedicated deployments. Try both models without code in the [Playground](https://playground.aircaps.com/playground). # Limits Source: https://docs.aircaps.com/limits Usage limit, concurrency and rate limits. ## Usage limit Each account includes **5 hours of free usage**, shared by both models. * **A5S Async v1** counts the duration of each completed transcript. Failed jobs are not counted. * **A5S v2 Streaming** counts the audio you stream. * A file longer than your remaining usage fails with `usage_limit_reached` and is not counted. New requests at a used-up limit return `403 usage_limit_reached`. * See what's left in the console or with [`GET /v1/account`](/api-reference/account) (`usage_limit`). ## Defaults | Limit | Default | | - | - | | A5S Async v1 jobs processing at once | 5 (more are queued, first in first out) | | A5S Async v1 jobs waiting in the queue | 2,000 | | Transcript create requests (`POST /v1/transcripts`) | 600 per minute | | Maximum file duration / size | 10 hours / 5 GB | | Direct upload (`POST /v1/files`) | 1 GB (use an upload URL for larger files) | | Uploaded files | Deleted when the transcript that uses them finishes; unused uploads after 24 hours | | Transcripts kept (including failed ones) | 30 days, or until you delete them ([Data & privacy](/data-privacy)) | | A5S v2 Streaming sessions at once | 5 | | Streaming session length | 5 minutes | | Active API keys | 10 | Jobs beyond the processing limit wait as `queued`. Beyond the queue or request-rate limit you get `429` with a `Retry-After` header. Other endpoints are not rate limited. ## Pricing and higher limits Limits are set per account. For pricing, higher usage limits or concurrency, access to a model, or a dedicated deployment, [talk to sales](https://research.aircaps.com/contact). research.aircaps.com/contact # Models Source: https://docs.aircaps.com/models What each model does and when to use it. ## A5S Async v1 `a5s-async-v1` · REST · [batch guide](/guides/async-transcription) * Recorded audio or video in any common format, from 0.1 s to 10 hours. * Speaker labels (`speaker_0`, `speaker_1`, …), word timestamps, utterances, SRT/VTT subtitles. * English only: `language_code` accepts `en` and regional variants (`en-US`, `en-GB`, …); any other value returns `400 unsupported_language`. * Text is verbatim and lowercase without punctuation (apostrophes are kept, as in "let's"); disfluencies ("uh", "um") are kept. ## A5S v2 Streaming `a5s-v2-streaming` · WebSocket · [streaming guide](/guides/streaming) * Live 16 kHz mono PCM16 audio in, partial and final text out as people speak. * Punctuated, cased text. No speaker labels. * Optional custom vocabulary to bias recognition toward names and terms. * English only: the start message's `language` accepts `en` and regional variants (`en-US`, `en-GB`, …); any other value is refused with `unsupported_language`. * Sessions last up to 5 minutes. ## Choosing | You have | Use | | - | - | | Recordings, uploads, meeting files | A5S Async v1 | | A live microphone or call that needs captions now | A5S v2 Streaming | | Live audio that later needs speaker labels | Stream for captions, then send the recording to A5S Async v1 | ## Languages Both models support English today. **Multilingual support is planned and will be available soon.** Accepted codes: `en`, `en-US`, `en-GB`, `en-AU`, `en-CA`, `en-IN`, `en-IE`, `en-NZ`, `en-ZA` (case-insensitive; `_` works too). ## Access Accounts can use both models by default. An account may be limited to one model. Requests to a model your account can't use return `403 model_access_denied` (streaming: an `error` message with that code, then close code `4403`). `GET /v1/account` shows `enabled` for each model. To get access, [talk to sales](https://research.aircaps.com/contact). `GET /v1/models` lists the models. # Quickstart Source: https://docs.aircaps.com/quickstart Transcribe a recording with speaker labels. Sign in at [playground.aircaps.com](https://playground.aircaps.com), open **API keys**, and create a key. It is shown once. ```bash theme={null} export AIRCAPS_API_KEY="aircaps_sk_..." ``` Pass a URL the API can download (public or presigned): ```bash theme={null} curl https://api.aircaps.com/v1/transcripts \ -H "Authorization: Bearer $AIRCAPS_API_KEY" \ -H "Content-Type: application/json" \ -d '{"model": "a5s-async-v1", "audio_url": "https://example.com/meeting.mp3", "language_code": "en"}' ``` `language_code` is optional (default `en`). The response is a transcript object with `"status": "queued"` or `"processing"` and an `id` like `tr_01k6z...`. ```bash theme={null} curl https://api.aircaps.com/v1/transcripts/tr_01k6z... \ -H "Authorization: Bearer $AIRCAPS_API_KEY" ``` Poll until `status` is `completed` (or `error`), or pass a `webhook_url` when you submit to be notified instead. ## Local files in Python ```python theme={null} # pip install requests import os, sys, time, requests API = "https://api.aircaps.com/v1" HEADERS = {"Authorization": f"Bearer {os.environ['AIRCAPS_API_KEY']}"} def call(method, path, **kw): r = requests.request(method, API + path, headers=HEADERS, timeout=60, **kw) if r.status_code >= 400: raise SystemExit(r.json()["error"]["message"]) return r.json() # 1. Upload straight to storage upload = call("POST", "/files/upload-url") with open(sys.argv[1], "rb") as f: requests.put(upload["upload_url"], data=f, timeout=600).raise_for_status() # 2. Transcribe job = call("POST", "/transcripts", json={"model": "a5s-async-v1", "file_id": upload["id"], "language_code": "en"}) # 3. Wait for the result while job["status"] in ("queued", "processing"): time.sleep(3) job = call("GET", f"/transcripts/{job['id']}") if job["status"] == "error": raise SystemExit(job["error"]["message"]) for u in job["utterances"]: print(f"[{u['start']:7.1f}s] {u['speaker']}: {u['text']}") ``` **Language:** both models transcribe English only today. Pass `language_code` (batch) or `language` (streaming) as `en` or a regional variant such as `en-US` or `en-GB`. Any other value returns `unsupported_language`. Multilingual support is planned and will be available soon. Next: [batch transcription in depth](/guides/async-transcription) or [live streaming](/guides/streaming).