Skip to main content

Meeting Recorders

Optimize for structured, speaker-attributed meeting notes.
Diarization vs. multi-channel: if each speaker is on a separate audio channel, use the channel field on each utterance — diarization is not needed. See Multiple channels.If all speakers share a single audio channel, enable diarization. See Speaker diarization.

Call Centers

Optimize for speaker identification and accuracy on telephony audio.

Podcasts & Interviews

Optimize for readability and correct speaker attribution.

Subtitles & Captioning

Optimize for readable on-screen captions.
For live captions streamed in real time, use the Realtime API instead — see Live recommended parameters.

Multilingual Content

Optimize for mixed-language audio.
Do not enable code_switching with an empty languages list — the detector would evaluate against 100+ languages and misdetect frequently.

Language Configuration

One of the most common configuration mistakes is misunderstanding how language_config works. Choosing the right setup avoids unnecessary detection overhead and improves accuracy. When to set an explicit language:
  • You know the language of the audio ahead of time.
  • The audio is monolingual (single language throughout).
  • You want the fastest, most accurate results.
When to use auto-detection:
  • You process audio in many different languages and don’t know which one beforehand.
  • You want Gladia to pick the language automatically.
When code_switching is false and no language is set, the language is detected on the first utterance and reused for the rest of the session or file. If the beginning of your audio contains silence, music, or a different language than the main content, this can lead to incorrect detection for the whole transcription.
Even when using auto-detection, pass a small list of likely languages in languages to constrain the search. This improves both accuracy and processing time.

Code Switching

Code switching (language_config.code_switching: true) lets Gladia detect and transcribe multiple languages within the same audio, re-evaluating the language on each utterance. When to enable it:
  • Speakers switch languages mid-conversation (e.g. bilingual meetings, multilingual customer support).
  • You need the detected language returned per utterance.
When NOT to enable it:
  • The audio is in a single language — code switching adds unnecessary processing and can introduce misdetections.
  • You’ve set exactly one language in languages — in that case code_switching is ignored anyway.
Do not enable code_switching with an empty languages list. When no languages are specified, the language detector evaluates every utterance against 100+ supported languages, which leads to frequent misdetections — especially between similar-sounding languages. Always provide a short list of languages you actually expect in the audio.

Custom Vocabulary

Custom vocabulary is a post-transcription replacement based on phoneme similarity. It’s essential for domain-specific terms that speech models frequently mis-transcribe. Best practices:
  • Always provide both the custom_vocabulary flag and a custom_vocabulary_config.
  • Add pronunciations to provide all the close spelling variants. You can use Automatic Phonemic Transcriber (IPA) in order to check if all the different spellings are covered.
  • Keep intensity moderate (0.4-0.6). High values increase false positives where unrelated words get replaced.
  • Set language on individual vocabulary entries when your audio is multilingual and a term is pronounced differently depending on the language.