Install
openclaw skills install @gladiaio/gladia-documentation-autoComprehensive Gladia speech-to-text reference auto-synced from docs.gladia.io. Use as a general-purpose fallback when other specialized skills don't match, or when the user needs a broad overview of Gladia capabilities, endpoints, decision guidance, or workflows. Always prefer the official SDK; fall back to raw REST/WebSocket only when SDK cannot satisfy the requirement.
openclaw skills install @gladiaio/gladia-documentation-autoSDK-first: always use the official SDK — see gladia-sdk-integration for policy, setup, and fallback criteria.
Consult these sibling skills as needed:
Gladia is a speech-to-text (STT) API for transcribing audio and video to text. It supports two modes: pre-recorded (async file upload) and live (real-time WebSocket streaming). The API includes audio intelligence features (diarization, translation, sentiment analysis, PII redaction, summarization) and two models: Solaria-3 (highest accuracy on European audio, pre-recorded only, 5 languages) and Solaria-1 (default, 100+ languages, live + async, code switching).
Key files and endpoints:
POST /v2/pre-recorded (create job), GET /v2/pre-recorded/:id (poll result)POST /v2/live (init session), WebSocket connection for streamingx-gladia-key: YOUR_API_KEY@gladiaio/sdk (JavaScript/TypeScript), gladiaio-sdk (Python)gladia transcribe <file> for terminal usePrimary docs: https://docs.gladia.io
Reach for Gladia when:
Do not use Gladia for:
# Set API key in environment
export GLADIA_API_KEY=your_key
# Or pass per request
curl -H "x-gladia-key: your_key" https://api.gladia.io/v2/pre-recorded
const gladia = new GladiaClient({ apiKey: "YOUR_KEY" });
const result = await gladia.preRecorded().transcribe("audio.mp3");
gladia = GladiaClient(api_key="YOUR_KEY").prerecorded()
result = gladia.transcribe("audio.mp3")
const session = gladia.liveV2().startSession({
encoding: "wav/pcm",
sample_rate: 16000,
bit_depth: 16,
channels: 1,
});
session.on("message", (msg) => {
if (msg.type === "transcript" && msg.data.is_final) {
console.log(msg.data.utterance.text);
}
});
session.sendAudio(audioChunk);
session.stopRecording();
gladia auth set your_key
gladia transcribe meeting.wav # text output
gladia transcribe podcast.mp3 -o json # JSON output
gladia transcribe call.wav --diarize -o srt # subtitles with speakers
gladia transcribe mixed.mp3 --code-switching # mixed languages
gladia transcribe audio.mp3 --model solaria-3 --language en
| Model | Best for | Modes | Languages | Code switching |
|---|---|---|---|---|
| Solaria-3 | European real-world audio, highest accuracy | Pre-recorded only | EN, FR, DE, ES, IT | No |
| Solaria-1 | Default, global coverage, live streaming | Pre-recorded + live | 100+ | Yes |
diarization: true)translation: true, set target_languages)sentiment_analysis: true)pii_redaction: true)summarization: true, set type: "general" | "bullet_points" | "concise")named_entity_recognition: true)custom_vocabulary: true, provide vocabulary list)subtitles: true, set formats)| Limit | Value |
|---|---|
| Pre-recorded max duration | 135 minutes (enterprise: 4h15) |
| Live session max duration | 3 hours |
| Pre-recorded concurrency | 25 parallel + 300 queued (paid) |
| Live concurrency | 30 concurrent sessions |
| File size | 1000 MB max |
| Channels (pre-recorded) | 2 (mono/stereo) |
| Channels (live) | 8 |
| Condition | Use Solaria-3 | Use Solaria-1 |
|---|---|---|
| Pre-recorded audio | ✓ | ✓ |
| Live/streaming | ✗ | ✓ |
| European business audio (calls, meetings) | ✓ | — |
| 100+ languages needed | ✗ | ✓ |
| Code switching (mixed languages) | ✗ | ✓ |
| EN, FR, DE, ES, IT only | ✓ | ✓ |
| Clean, formal speech | — | ✓ |
| Scenario | Pre-recorded | Live |
|---|---|---|
| Transcribe uploaded file | ✓ | ✗ |
| Real-time captions | ✗ | ✓ |
| Voice agent / IVR | ✗ | ✓ |
| Meeting recording | ✓ | ✓ (stream during call) |
| Batch processing | ✓ | ✗ |
| Async job with polling | ✓ | ✗ |
| WebSocket streaming | ✗ | ✓ |
| Tool | Best for | Complexity |
|---|---|---|
| SDK (JavaScript/Python) | Application integration, error handling, retries | Low |
| API (REST/WebSocket) | Custom workflows, non-SDK languages | Medium |
| CLI | Terminal, scripts, CI/CD, one-off transcriptions | Very low |
| Method | Use when | Trade-off |
|---|---|---|
Polling (SDK .poll()) | Simple, synchronous flow | Blocks, wastes requests |
| Webhooks | Server-to-server, configured in dashboard | Setup required, less flexible |
| Callbacks | Per-job notification, no polling | Must expose HTTP endpoint |
Get API key from https://app.gladia.io/apikeys and set GLADIA_API_KEY environment variable.
Choose model and features: Decide between Solaria-3 (accuracy, 5 languages) or Solaria-1 (coverage, 100+ languages). List required audio intelligence features (diarization, translation, etc.).
Prepare audio: Ensure file is under 135 minutes, under 1000 MB, and in a supported format (MP3, WAV, M4A, FLAC, OGG, etc.).
Upload and transcribe (SDK):
const result = await gladia.preRecorded().transcribe("audio.mp3", {
model: "solaria-3",
language_config: { languages: ["en"] },
diarization: true,
});
Poll or wait for callback: SDK .transcribe() polls automatically. For raw API, use .poll() or configure a callback URL.
Extract results: Access result.transcription.full_transcript, result.transcription.utterances, result.diarization, result.translation, etc.
Handle errors: Check result.status for "done" or "error". Inspect error details before retrying (transient vs. input issues).
Initialize session (backend):
const response = await fetch("https://api.gladia.io/v2/live", {
method: "POST",
headers: { "x-gladia-key": "YOUR_KEY", "Content-Type": "application/json" },
body: JSON.stringify({
encoding: "wav/pcm",
sample_rate: 16000,
bit_depth: 16,
channels: 1,
language_config: { languages: ["en"] },
}),
});
const { id, url } = await response.json();
Return secure URL to client: Pass url (contains temporary token) to frontend/mobile app. Keep API key on backend.
Client connects to WebSocket and sends audio chunks as they arrive.
Listen for messages: Handle transcript (partial/final), speech_start, speech_end, sentiment_analysis, etc.
Stop recording: Send stop_recording message or close WebSocket with code 1000. Server processes remaining audio and post-processing.
Retrieve final results: Call GET /v2/live/:id to fetch complete transcript, diarization, translation, etc.
Solaria-3 with multiple languages: Solaria-3 does not support code switching. Pass exactly one language in language_config.languages (e.g., ["fr"]), not multiple. Use Solaria-1 for mixed-language audio.
Live session timeout: A single WebSocket session cannot exceed 3 hours. For longer events, start a new session before reaching the limit. Track session duration and reconnect proactively.
Resubmitting the same audio: If a POST already returned 200 or you received a transcription.created webhook, do not resubmit the same audio. Wait on that job ID. Resubmitting creates a new billable job.
Polling without backoff: Polling too aggressively wastes API quota. Use exponential backoff (start at 1s, cap at 10s) or switch to callbacks/webhooks.
Audio format mismatch: Ensure encoding, sample_rate, bit_depth, and channels match your actual audio. Mismatches cause silent failures or garbled output.
Missing language specification: If you know the language, set language_config.languages to skip auto-detection and reduce latency. Auto-detection adds 1–2 seconds.
Callback URL not reachable: If using callbacks, ensure your endpoint is publicly accessible and returns 2xx within a reasonable timeout. Gladia retries failed callbacks.
Concurrency limits: Free tier: 3 pre-recorded concurrent, 1 live. Paid: 25 pre-recorded concurrent, 30 live. Requests beyond the limit queue. Monitor queue time during high load.
Data retention: By default, audio and transcripts are retained. If GDPR/privacy is a concern, enable Zero Data Retention (results delivered only via callbacks, no retrieval).
Partial transcripts accuracy: Partial transcripts use a faster, smaller model than finals. Accuracy degrades with multiple languages or code switching. Use for UX only, not for final output.
Before submitting work:
transcription.full_transcript, transcription.utterances, diarization.speakers, translation.results, etc.Comprehensive page-by-page navigation: https://docs.gladia.io/llms.txt
Critical documentation pages:
For additional documentation and navigation, see: https://docs.gladia.io/llms.txt
This file is auto-synced from https://docs.gladia.io/.well-known/agent-skills/gladia/skill.md Do not edit manually — changes will be overwritten by CI. For additional documentation and navigation, see: https://docs.gladia.io/llms.txt