Skip to main content

Overview

Sarvam AI offers two real-time speech recognition services for Indian languages:
  • SarvamSTTService: Uses Sarvam’s transcription WebSocket API with VAD-based segmentation and multiple audio formats (saaras:v4 model by default)
  • SarvamRealtimeSTTService: Uses Sarvam’s realtime WebSocket endpoint (saaras:v3-realtime model) with server-side endpointing, in-band configuration updates, and lower latency

Sarvam STT API Reference

Pipecat’s API methods for Sarvam STT integration

SarvamSTTService Example

Complete example with VAD-based turn detection

SarvamRealtimeSTTService Example

Realtime STT with manual endpointing

Sarvam Documentation

Official Sarvam AI STT documentation and features

Sarvam AI Platform

Access API keys and speech models

Installation

To use Sarvam services, install the required dependency:

Prerequisites

Sarvam AI Account Setup

Before using Sarvam STT services, you need:
  1. Sarvam AI Account: Sign up at Sarvam AI
  2. API Key: Generate an API key from your account dashboard
  3. Model Access:
    • SarvamSTTService: Access to saaras:v3 and saaras:v4 models with support for multiple modes (transcribe, translate, verbatim, translit, codemix)
    • SarvamRealtimeSTTService: Access to the saaras:v3-realtime model

Required Environment Variables

  • SARVAM_API_KEY: Your Sarvam AI API key for authentication

Configuration

SarvamSTTService

str
required
Sarvam API key for authentication.
str
default:"saaras:v4"
deprecated
Sarvam model to use. Allowed values: "saaras:v3", "saaras:v4". Deprecated in v0.0.105. Use settings=SarvamSTTService.Settings(...) instead.
int
default:"None"
Audio sample rate in Hz. Defaults to 16000 if not specified.
Literal['transcribe', 'translate', 'verbatim', 'translit', 'codemix']
default:"None"
Mode of operation. Only applicable to models that support it (e.g., saaras:v3, saaras:v4). Defaults to the model’s default mode.
str
default:"wav"
Audio codec/format of the input file.
SarvamSTTService.InputParams
default:"None"
deprecated
Configuration parameters for Sarvam STT service. Deprecated in v0.0.105. Use settings=SarvamSTTService.Settings(...) instead.
SarvamSTTService.Settings
default:"None"
Runtime-configurable settings for the STT service. See Settings below.
float
default:"None"
Seconds of no audio before sending silence to keep the connection alive. None disables keepalive.
float
default:"SARVAM_TTFS_P99"
P99 latency from speech end to final transcript in seconds. Override for your deployment. See stt-benchmark.
float
default:"5.0"
Seconds between idle checks when keepalive is enabled.

Settings

Runtime-configurable settings passed via the settings constructor argument using SarvamSTTService.Settings(...). These can be updated mid-conversation with STTUpdateSettingsFrame. See Service Settings for details.

SarvamRealtimeSTTService

str
required
Sarvam API key for authentication.
str
default:"wss://api.sarvam.ai/speech-to-text-realtime/ws"
Realtime STT websocket endpoint.
Literal['vad', 'manual']
default:"vad"
Which side detects turn boundaries: vad for Sarvam’s own detection, or manual for the pipeline’s. Decides the turn strategies this service asks the user aggregator to run. Defaults to vad.
int
default:"None"
Declared input audio sample rate, 8000 or 16000. None adopts the pipeline’s input rate.
bool
default:"False"
Whether final transcripts should include segment offsets.
int
default:"None"
Optional VAD prefix padding, used only under endpointing="vad".
SarvamRealtimeSTTService.Settings
default:"None"
Runtime-updatable realtime settings. See Realtime Settings below.
bool
default:"True"
Whether the bot should be interrupted when Sarvam detects user speech. See User Turn Strategies if you pass your own user_turn_strategies.
float
default:"SARVAM_REALTIME_TTFS_P99"
P99 latency from speech end to final transcript in seconds. Override for your deployment.

Realtime Settings

Runtime-configurable settings passed via the settings constructor argument using SarvamRealtimeSTTService.Settings(...). These can be updated mid-conversation via config.update() or STTUpdateSettingsFrame.

Usage

Basic Setup (SarvamSTTService)

With Language and Model Configuration

With Server-Side VAD

Basic Realtime Setup

Realtime with Language and Stream Type

Realtime with Manual Endpointing

Updating Realtime Configuration

Notes

SarvamSTTService

  • Supported models: saaras:v3 and saaras:v4. The default is saaras:v4.
  • Supported languages: Bengali (bn-IN), Gujarati (gu-IN), Hindi (hi-IN), Kannada (kn-IN), Malayalam (ml-IN), Marathi (mr-IN), Tamil (ta-IN), Telugu (te-IN), Punjabi (pa-IN), Odia (od-IN), English (en-IN), and Assamese (as-IN).
  • Fine-grained VAD tuning: Both models support server-side VAD with 10 tuning parameters for speech detection thresholds, frame-count controls, pre-speech padding, interruption sensitivity, and initial-frame skipping.
  • VAD modes: When vad_signals=False (default), the service relies on Pipecat’s local VAD and flushes the server buffer on VADUserStoppedSpeakingFrame. When vad_signals=True, the service uses Sarvam’s server-side VAD and proposes turn starts and stops from the server signals. In that mode it automatically requests ExternalUserTurnStrategies, which resolve those proposals into UserStartedSpeakingFrame and UserStoppedSpeakingFrame, so you don’t need to configure turn strategies manually. Pass your own user_turn_strategies only to override this. This service has no should_interrupt parameter, so the strategies it requests always interrupt the bot when the user starts speaking.

SarvamRealtimeSTTService

  • Supported sample rates: Only 8000 Hz and 16000 Hz are supported. The service validates the sample rate at initialization.
  • Supported languages: auto, Bengali (bn-IN), Gujarati (gu-IN), Hindi (hi-IN), Kannada (kn-IN), Malayalam (ml-IN), Marathi (mr-IN), Tamil (ta-IN), Telugu (te-IN), Punjabi (pa-IN), Odia (or-IN), English (en-IN), Assamese (as-IN), Urdu (ur-IN), Nepali (ne-IN), Konkani (kok-IN), Kashmiri (ks-IN), Sindhi (sd-IN), Sanskrit (sa-IN), Santali (sat-IN), Manipuri (mni-IN), Bodo (brx-IN), Maithili (mai-IN), and Dogri (doi-IN).
  • Endpointing modes:
    • endpointing="vad" (default): Sarvam’s server-side VAD decides turn boundaries. The service proposes turn frames via ProposedUserStartedSpeakingFrame and ProposedUserStoppedSpeakingFrame.
    • endpointing="manual": The pipeline drives turn boundaries via VADUserStartedSpeakingFrame and VADUserStoppedSpeakingFrame. Requires a VAD analyzer on the user aggregator.
  • VAD analyzer required: A VAD analyzer is required in either endpointing mode. Under vad it times transcription latency; under manual it also marks the turn for Sarvam.
  • In-band configuration updates: Settings can be updated mid-conversation via update_config() without reconnecting. Connection-only values (sample_rate, return_timestamps, prefix_padding_ms, endpointing) cannot be updated at runtime.
  • Stream types: fast provides lowest latency, balanced provides moderate latency with better quality, and simulated is for testing. The stream type controls server-side flush cadence; the client always sends audio in 50ms chunks.
The InputParams / params= pattern is deprecated as of v0.0.105. Use Settings / settings= instead. See the Service Settings guide for migration details.

Event Handlers

In addition to the standard service connection events (on_connected, on_disconnected, on_connection_error), Sarvam STT provides: