Sorika Labs Releases Vaani Speech AI Series: Zero-to-SOTA on Community GPUs

- Zero-to-SOTA on Community GPUs: Converged entirely on free/community tiers of Google Colab and Kaggle (NVIDIA T4 / P100 GPUs) without expensive H100 clusters
- Mathematically Sound Stabilization: High-Rank LoRA (r=64/128, α=128/256), gradient accumulation, custom loss masking (labels = -100), and frozen audio encoder shielding
- Zero Adapter Overhead: 100% permanently fused weights (model.safetensors) for instantaneous plug-and-play production inference
- Vaani-2 Flash (809M Turbo): #1 Overall English Accuracy (5.90% Avg WER across 7 benchmarks) with an insane 21.5x to 30.3x Real-Time inference speed (RTF)
- Vaani-1.1 Lite (~600M Conformer-LLM): #1 SOTA on Indian Speech (11.78% Hindi Avg WER), excelling in high-noise acoustic environments (3.60% LibriSpeech Other)
- Vaani-1 Lite (244M): Slashes baseline Whisper-small Hindi WER from 47.77% down to 19.32% (>2.5x error reduction) with CTranslate2 edge/CPU support
- Hugging Face Hub Release: Complete open safetensors weights and collections available under sorika-labs
Sorika Labs announces the official public release of the Vaani Speech AI Series — a breakthrough suite of sovereign, high-efficiency speech recognition and multimodal foundation models. Engineered and converged entirely on community hardware (Google Colab and Kaggle community GPUs), Vaani proves that sovereign, frontier-grade speech intelligence does not require massive compute clusters burning hundreds of thousands of dollars. Through rigorous architectural discipline, High-Rank LoRA adaptation, and frozen encoder shielding, Vaani establishes new State-Of-The-Art benchmarks across 10 evaluation datasets.
The Zero-to-SOTA Engineering Feat
Training frontier foundation models under strictly limited community compute.
Unlike large tech conglomerates burning hundreds of thousands of dollars on massive H100 clusters, Sorika Labs engineered, stabilized, and converged the entire Vaani speech series on free and community tiers of Google Colab and Kaggle (NVIDIA T4 and P100 GPUs).
This engineering feat underscores our foundational thesis: intelligence is an architectural and mathematical discipline, not merely a function of brute-force electrical expenditure. By pairing lightweight backbones with high-rank low-rank adaptation, we achieved State-Of-The-Art speech recognition that rivals or exceeds models trained with orders of magnitude more capital.
“True foundational research is demonstrated when you produce state-of-the-art weights with architectural elegance under extreme resource constraints. Vaani is our proof that sovereign AI can be engineered by solo architects and small labs on community GPUs.”
— Darsh Yadav — Founder & Chief AI Architect, Sorika LabsMathematical Foundations & Catastrophic Forgetting Shielding
High-Rank LoRA, custom loss masking, and NFC Unicode normalization.
Fine-tuning foundation speech models on resource-constrained GPUs usually triggers catastrophic forgetting or acoustic degradation. We resolved this through four deliberate mathematical techniques:
1. High-Rank LoRA (r=64/128, α=128/256): Provides sufficient representational capacity for complex multilingual phonetics while keeping trainable parameters minimal.
2. Frozen Audio Encoder Shielding: The acoustic feature extractors remain entirely frozen, preserving universal auditory features while adapting target cross-attention and projection matrices.
3. Custom Loss Masking: Implementing strict labels = -100 masking on padding and prefix tokens ensures clean loss backpropagation without noisy gradient artifacts.
4. NFC Unicode Text Normalization: Native Devanagari and Indic scripts undergo rigorous NFC normalization, preventing composite character fragmentation and erratic tokenization splits.
5. 100% Permanently Fused Weights: Every model is released with permanently fused weights into clean model.safetensors. Developers experience zero PEFT adapter loading overhead and immediate plug-and-play execution.
Product Architecture Matrix: Lite Series vs. Flash Series
Two distinct product tiers designed for edge sovereignty and high-volume enterprise APIs.
The Vaani family is organized into two distinct series to match specific real-world deployment profiles:
Lite Series (Ultra-Lightweight, Edge & Mobile Friendly):
• Vaani-1 Lite ASR (244M): Built on Whisper-small. Solves the severe baseline breakdown of Whisper on Indian speech, slashing Hindi WER from 47.77% down to 19.32% (>2.5x error reduction) while retaining full 99+ global language capability. Runs natively on CPUs, laptops, and mobile devices. Includes Faster-Whisper CTranslate2 release (sorika-labs/vaani1-lite-asr-turbo).
• Vaani-1.1 Lite ASR (~600M): Built on Qwen3-ASR-0.6B-hf (Conformer audio encoder + Qwen3 LLM reasoning decoder). Captures #1 SOTA on Indian Speech with an 11.78% Hindi Average WER. Excels in noisy acoustic environments, achieving 12.36% on Kathbath Hindi and an extraordinary 3.60% on LibriSpeech Other.
Flash Series (High-Throughput APIs & Real-Time Streaming):
• Vaani-2 Flash ASR (809M): Built on Whisper-large-v3-turbo with a 4-layer fast decoder. Achieves #1 Overall English Accuracy (5.90% Avg WER across 7 public benchmarks) while delivering an insane 21.5x to 30.3x Real-Time inference speed (RTF). It processes 30 seconds of speech in under 1 second on standard GPU.
• Vaani-2.1 Flash (In Active Development): Next-generation roadmap target focused on sub-10% Hindi WER and native Hinglish code-switching without losing the 25x+ RTF velocity.
The Grand Benchmark Leaderboard
| Benchmark Dataset | Domain | Whisper-Small 244M | Vaani-1 Lite 244M | Vaani-2 Flash 809M Turbo | Qwen3-Base 600M | Vaani-1.1 Lite 600M LLM | Flash Speed |
|---|---|---|---|---|---|---|---|
| AMI-Cleaned | English (Meetings) | 12.74% | 12.99% | 10.41% | 9.05% | 15.13% | 29.4x RTF ⚡ |
| Earnings22 | English (Financial Calls) | 8.62% | 9.90% | 8.62% 🥇 | 10.99% | 10.50% | 30.3x RTF ⚡ |
| GigaSpeech | English (Audiobooks/Pods) | 8.90% | 8.43% | 7.28% | 6.77% | 7.58% | 19.1x RTF ⚡ |
| LibriSpeech Clean | English (Clean Read) | 5.63% | 9.67% | 4.84% | 4.49% | 4.75% | 15.9x RTF ⚡ |
| LibriSpeech Other | English (Noisy / Accents) | 11.30% | 8.70% | 4.35% | 3.93% | 3.60% 🥇 | 15.3x RTF ⚡ |
| SPGISpeech | English (Financial) | 3.89% | 5.41% | 3.55% 🥇 | 4.74% | 5.58% | 18.9x RTF ⚡ |
| VoxPopuli | English (Parliamentary) | 5.29% | 9.20% | 2.22% 🥇 | 3.38% | 8.99% | 28.1x RTF ⚡ |
| Kathbath (Hindi) | Hindi (Indic Crowdsource) | 50.00% | 16.29% | 22.47% | 14.61% | 12.36% 🥇 | 5.9x RTF |
| FLEURS (Hindi) | Hindi (Conversational STT) | 48.88% | 25.56% | 16.38% | 14.39% | 14.64% | 8.5x RTF |
| Vaani Private Test | Hindi (Real Conversations) | 44.44% | 16.11% | 14.44% | 7.78% | 8.33% | 6.8x RTF |
| English Average WER | 7 English Benchmarks | 8.05% | 9.18% | 5.90% 🏆 | 6.19% | 8.02% | 21.5x Avg 🚀 |
| Hindi Average WER | 3 Indic Benchmarks | 47.77% | 19.32% | 17.76% | 12.26% | 11.78% 🏆 | 7.1x Avg |
| Overall Average WER | All 10 Benchmarks | 19.97% | 12.23% | 9.46% | 8.01% | 9.15% | 17.2x Avg |
*Note: Lower Word Error Rate (WER) represents superior transcription accuracy. Evaluated on identical audio splits using the standard Whisper Text Normalizer. 🥇 indicates benchmark winner; 🏆 indicates grand category champion.
Interactive Inference Snippets
from transformers import pipeline
# Load Sorika Labs Vaani-2 Flash (Fused SafeTensors, Zero Adapter Overhead)
pipe = pipeline("automatic-speech-recognition", model="sorika-labs/vaani2-flash", device="cuda")
# 21.5x - 30.3x Real-Time Inference Speed (30s audio in < 1 second)
result = pipe("speech.wav")
print(result["text"])Deploy Open-Source Vaani from Hugging Face Hub
All foundation checkpoints are released as 100% permanently fused model.safetensors with zero adapter loading overhead. Available under permissible open-source terms at huggingface.co/sorika-labs.
sorika-labs
Official Sorika Labs hub hosting all open-source foundation models, safetensors checkpoints, and benchmarks.
Vaani Flash Series
High-throughput streaming models including Vaani-2 Flash (809M Turbo) with #1 English accuracy.
Vaani Lite Series
Mobile and edge foundation models including Vaani-1.1 Lite (600M LLM) and Vaani-1 Lite (244M).