Live Captions
To make use of Live Captions, it must be explicitly enabled for your organization. Activation may be subject to additional pricing or service terms.
You can verify whether this feature is available by navigating to dashboard.nanostream.cloud/organisation in your dashboard.
In the Enabled Packages section, locate the entry for Live Captions. If it shows Upgrade needed, please contact us.

To activate Live Captions or learn more about available plans, feel free to reach out via nanocosmos.net/contact. We're happy to assist you in finding the best setup for your use case.
Live captions are only available for secure playback. Therefore, by enabling live captions, you will also need to enable the secure feature. Please reach out to our sales team via nanocosmos.net/contact or by email at sales(at)nanocosmos.net if you have any questions.
To learn more about secure playback, visit the dedicated article Secure Playback (H5Live).
Overview
Live Captions convert spoken audio into readable text in real time. This AI-driven feature enhances accessibility and content comprehension across a wide range of live-streaming, especially for:
- Events with spoken content
- Viewers in sound-off environments
- Users who are hard of hearing
- Corporate, educational, or public-facing broadcasts
This provides dynamic, accurate, and easy-to-follow text output that helps viewers stay engaged even without audio. Live captions start automatically as soon as the stream becomes active. The first caption lines typically appear within 5–7 seconds, depending on the selected ASR engine. Captions stop automatically when the stream ends.
To ensure low-latency and reliable delivery, all captions are produced and transmitted through a dedicated real-time output channel, separate from the video stream.
Live Captions and the caption switcher are not included in the default H5Live Player UI. This means: they are not embedded automatically. To allow viewers to enable, disable, or style captions, your playback environment must integrate caption handling explicitly. For implementation guidance or UI integration examples, please contact our support team via nanocosmos.net/support
How It Works
During an active stream, the audio is forwarded to the selected Automatic Speech Recognition (ASR) engine. The engine converts speech into text and outputs a continuous caption stream. The H5Live Player synchronizes to this caption feed and displays it to viewers in real-time.
Some of our ASR services require a 24-hour advance notice. Please contact our sales team via nanocosmos.net/contact to find the best configuration for your business. They will also be happy to give you in-depth advice and recommendations on the ASR types for your use case.
ASR Engines And Langauges
As already explained earlier, Live Captions rely on Automatic Speech Recognition (ASR) engines to convert spoken audio into real-time text. This section explains which ASR providers are available, how they differ, and which languages they support. You will also learn the difference between source and target languages and how they are applied during live caption generation.
Supported ASR Engines
- Deepgram
Deepgram is an enterprise-grade ASR engine designed for high accuracy and very low latency in real-time captioning. It uses neural-network models optimized for live audio and supports multiple languages for both transcription and live translation.
Whisper is an open-source ASR system developed by OpenAI. It offers robust multilingual speech recognition and performs well across diverse audio conditions.
Supported Languages
The source language is the spoken language of the incoming audio. This language is used by the ASR engine to interpret the speech and generate text. The target language defines the output language of the captions. Only engines that support translation can provide multiple output languages.
Some ASR engines offer region-specific language codes (e.g. es-419 for Latin American Spanish). Use these variants when your target audience is primarily from a specific region and you want improved recognition of regional accents, vocabulary, and spelling conventions.
If you do not require a regional focus, the generic language code (e.g. es) is typically sufficient.
| Language | ID | Source | Target | Supported Engines |
|---|---|---|---|---|
| Bulgarian | bg | ✓ | ✓ | Deepgram |
| Catalan | ca | ✓ |