Sarvam AI has released Saaras V4, a speech-to-text model designed to transcribe all 22 scheduled Indian languages alongside English, including global English accents. The company says the model delivers state-of-the-art accuracy across the Indian-language set. It is available through Sarvam’s API, although its weights have not been publicly released, according to Marktechpost.
Saaras V4 supports transcription, translation, transliteration, code-mixed output and verbatim speech from the same underlying model. Sarvam lists API pricing at ₹30 per hour for real-time, streaming and batch speech recognition, or ₹45 per hour when speaker diarization is included.
A speech model built around Sarvam-3B
Saaras V4 uses an encoder-decoder architecture. An audio encoder turns a recording into embeddings that represent phonetic and acoustic information. A temporal-downsampling adapter then reduces the length of that sequence and maps it into the embedding space used by the language model, helping long recordings fit within the decoder’s context budget.
The decoder is Sarvam-3B, a 3-billion-parameter hybrid state-space language model trained from scratch by Sarvam. It processes audio features together with a text prompt and generates the transcript autoregressively, using each emitted token as input for the next step.
The model is available through Sarvam’s API using the Saaras V4 model identifier. Saaras V3 remains the default model, but Sarvam says V4 uses the same request format, allowing developers to switch models with a one-line change. Self-hosting documentation currently covers Saaras V3 through SageMaker; Sarvam has not released self-hosting support or model weights for V4.
Five output modes and keyterm prompting
Developers can select among five output modes. The default transcription mode returns native script while normalizing numbers and dates. Verbatim mode preserves fillers, spoken numbers and the exact wording of the recording. Codemix keeps English words in English while using the native script for the rest of the utterance.
Translit renders the entire utterance in Latin script, while translate produces an English translation with normalized numbers. Sarvam’s approach places these transformations inside the model rather than relying on separate post-processing steps, which the company says can introduce additional errors.
Saaras V4 also adds keyterm prompting. Developers can provide a JSON list containing up to 50 terms, with each term limited to 64 characters, to bias recognition toward names, brands or other important vocabulary. The feature influences recognition but does not guarantee that the requested terms will appear. Sarvam recommends Codemix mode when a brand such as PhonePe should remain in Latin script.
On the IndicContextEval benchmark, Sarvam reports a 16.03% word error rate in the L5 keyword-prompting setting, which it identifies as the lowest score on that benchmark.
Reported benchmark results
For English, Sarvam evaluated seven datasets, including AMI, GigaSpeech, LibriSpeech, SPGISpeech, VoxPopuli and AI4Bharat’s Indian-accented Svarah dataset. Using the normalization procedure from Hugging Face’s Open ASR Leaderboard, the company reports that Saaras V4 achieved the lowest average word error rate among the models it benchmarked.
On AI4Bharat’s Vistaar benchmark, Sarvam measured performance across 10 Indian languages using both conventional word error rate and LLM-WER. The latter adds a semantic check intended to distinguish meaning-changing errors from harmless spelling or formatting differences in Indic scripts.
Sarvam also reports lower error rates on noisy audio. On the Kathbath Noisy dataset, which includes compressed, clipped and background-heavy recordings, the company says Saaras V4’s LLM-WER was less than half that of Deepgram Nova-3 and GPT-4o Transcribe. Its reported language-identification error was 2.9% across the top 10 Indian languages and 5.22% across all 22.
These figures remain vendor-reported, and independent reproduction has not been published.
Streaming, batch transcription and pricing
Saaras V4 supports WebSocket streaming with partial results and a vendor-reported time to first token below 150 milliseconds. Its synchronous REST endpoint handles clips of up to 30 seconds, while asynchronous batch jobs support files of up to two hours and can include speaker diarization.
Sarvam provides Python and Node.js SDKs and lists integrations with LiveKit Agents, Pipecat and the Vercel AI SDK. The company positions the model against services including Deepgram Nova-3, ElevenLabs Scribe v2 and OpenAI GPT-4o Transcribe, particularly on coverage of Indian languages and built-in output formats.
The immediate limitation is deployment flexibility: Saaras V4 is API-only, while the company’s available SageMaker self-hosting documentation remains for V3. For developers building applications that need broad Indian-language coverage, code-mixed speech or multiple transcript formats, the API release makes those capabilities available without requiring them to manage the model infrastructure.