Indian artificial intelligence company sarvam ai has unveiled saaras v4, its latest automatic speech recognition model designed to understand multilingual conversations and challenging real-world audio. the new ai speech model is built to handle multiple languages, code-mixed speech, background noise and different english accents, expanding the capabilities of speech technology for indian and international users.
What Is Sarvam AI Saaras V4?
Saaras v4 is the latest generation of sarvam ai's speech recognition technology. according to the company, the model combines an audio encoder with a 3-billion-parameter hybrid state-space language model trained in-house. this architecture is designed to convert spoken audio into useful text while supporting different output formats.
The model supports five output modes: transcription, translation, verbatim transcription, transliteration and code-mixed text. this gives developers greater flexibility when building applications that need to process speech from multilingual users.
Support for 22 Indian Languages
One of the biggest features of saaras v4 is its focus on india's diverse linguistic environment. sarvam says the model delivers state-of-the-art performance across 22 indian languages and also supports english.
The supported indian languages include hindi, bengali, tamil, telugu, marathi, gujarati, kannada, malayalam, punjabi, odia, assamese, urdu, nepali, konkani, kashmiri, sindhi, sanskrit, santali, manipuri, bodo, maithili and dogri.
Designed for Real-World Audio
Real-world conversations can include background noise, multiple speakers, accents and people switching between languages. saaras v4 is designed to work with these conditions rather than relying only on clean studio-quality recordings.
Sarvam says the model is robust to noisy audio, code mixing and dialect variations. its language-identification system reports a 2.9% error rate across the top 10 indian languages and 5.22% across all 22 indian languages in the company's evaluation.
New Features for Developers
Saaras v4 is available through sarvam's speech-to-text apis. developers can use rest, batch and websocket-based interfaces depending on their application requirements.
Another useful feature is keyterm prompting. developers can provide important names, brands, locations or technical terms to help the model recognize domain-specific vocabulary more accurately. sarvam's documentation says saaras v4 supports up to 50 keyterms for this purpose on supported api endpoints.
Potential Uses of Saaras V4
The new speech model could be used for ai voice agents, customer-service systems, call analytics, meeting transcription, education tools, accessibility services and multilingual applications.
Sarvam has also made saaras v4 generally available for its voice agents platform, where the model can automatically handle multilingual and mixed-accent conversations.
Sarvam AI's Growing Speech Technology Push
With saaras v4, sarvam ai is expanding its focus on speech technology designed for india's multilingual environment while adding support for global english. the combination of multiple output modes, indian-language coverage, code-mixed speech and real-world audio processing could make the model useful across a wide range of voice-based ai applications.
As voice interfaces become increasingly important in artificial intelligence, saaras v4 represents sarvam's latest effort to make speech recognition more capable across languages, accents and everyday communication environments.