ProductOct 5, 2026 • 7 min read

Sorika Labs Releases Vaani Speech AI Series: Zero-to-SOTA on Community GPUs

Darsh YadavFounder & Chief AI Architect • Sorika AI Labs
Sorika Labs Releases Vaani Speech AI Series: Zero-to-SOTA on Community GPUs
Announcement Highlights
  • Zero-to-SOTA on Community GPUs: Converged entirely on free/community tiers of Google Colab and Kaggle (NVIDIA T4 / P100 GPUs) without expensive H100 clusters
  • Mathematically Sound Stabilization: High-Rank LoRA (r=64/128, α=128/256), gradient accumulation, custom loss masking (labels = -100), and frozen audio encoder shielding
  • Zero Adapter Overhead: 100% permanently fused weights (model.safetensors) for instantaneous plug-and-play production inference
  • Vaani-2 Flash (809M Turbo): #1 Overall English Accuracy (5.90% Avg WER across 7 benchmarks) with an insane 21.5x to 30.3x Real-Time inference speed (RTF)
  • Vaani-1.1 Lite (~600M Conformer-LLM): #1 SOTA on Indian Speech (11.78% Hindi Avg WER), excelling in high-noise acoustic environments (3.60% LibriSpeech Other)
  • Vaani-1 Lite (244M): Slashes baseline Whisper-small Hindi WER from 47.77% down to 19.32% (>2.5x error reduction) with CTranslate2 edge/CPU support
  • Hugging Face Hub Release: Complete open safetensors weights and collections available under sorika-labs

Sorika Labs announces the official public release of the Vaani Speech AI Series — a breakthrough suite of sovereign, high-efficiency speech recognition and multimodal foundation models. Engineered and converged entirely on community hardware (Google Colab and Kaggle community GPUs), Vaani proves that sovereign, frontier-grade speech intelligence does not require massive compute clusters burning hundreds of thousands of dollars. Through rigorous architectural discipline, High-Rank LoRA adaptation, and frozen encoder shielding, Vaani establishes new State-Of-The-Art benchmarks across 10 evaluation datasets.

The Zero-to-SOTA Engineering Feat

Training frontier foundation models under strictly limited community compute.

Unlike large tech conglomerates burning hundreds of thousands of dollars on massive H100 clusters, Sorika Labs engineered, stabilized, and converged the entire Vaani speech series on free and community tiers of Google Colab and Kaggle (NVIDIA T4 and P100 GPUs).

This engineering feat underscores our foundational thesis: intelligence is an architectural and mathematical discipline, not merely a function of brute-force electrical expenditure. By pairing lightweight backbones with high-rank low-rank adaptation, we achieved State-Of-The-Art speech recognition that rivals or exceeds models trained with orders of magnitude more capital.

“True foundational research is demonstrated when you produce state-of-the-art weights with architectural elegance under extreme resource constraints. Vaani is our proof that sovereign AI can be engineered by solo architects and small labs on community GPUs.”

— Darsh Yadav — Founder & Chief AI Architect, Sorika Labs

Mathematical Foundations & Catastrophic Forgetting Shielding

High-Rank LoRA, custom loss masking, and NFC Unicode normalization.

Fine-tuning foundation speech models on resource-constrained GPUs usually triggers catastrophic forgetting or acoustic degradation. We resolved this through four deliberate mathematical techniques:

1. High-Rank LoRA (r=64/128, α=128/256): Provides sufficient representational capacity for complex multilingual phonetics while keeping trainable parameters minimal.

2. Frozen Audio Encoder Shielding: The acoustic feature extractors remain entirely frozen, preserving universal auditory features while adapting target cross-attention and projection matrices.

3. Custom Loss Masking: Implementing strict labels = -100 masking on padding and prefix tokens ensures clean loss backpropagation without noisy gradient artifacts.

4. NFC Unicode Text Normalization: Native Devanagari and Indic scripts undergo rigorous NFC normalization, preventing composite character fragmentation and erratic tokenization splits.

5. 100% Permanently Fused Weights: Every model is released with permanently fused weights into clean model.safetensors. Developers experience zero PEFT adapter loading overhead and immediate plug-and-play execution.

Hardware SubstrateGoogle Colab T4 & Kaggle Dual-GPU
Adaptation GeometryHigh-Rank LoRA (r=64/128, α=128/256)
Weight Format100% Permanently Fused Safetensors
Top English WER5.90% Avg (Vaani-2 Flash)
Top Hindi WER11.78% Avg (Vaani-1.1 Lite)
Inference Velocity21.5x – 30.3x Real-Time (RTF)

Product Architecture Matrix: Lite Series vs. Flash Series

Two distinct product tiers designed for edge sovereignty and high-volume enterprise APIs.

The Vaani family is organized into two distinct series to match specific real-world deployment profiles:

Lite Series (Ultra-Lightweight, Edge & Mobile Friendly):

• Vaani-1 Lite ASR (244M): Built on Whisper-small. Solves the severe baseline breakdown of Whisper on Indian speech, slashing Hindi WER from 47.77% down to 19.32% (>2.5x error reduction) while retaining full 99+ global language capability. Runs natively on CPUs, laptops, and mobile devices. Includes Faster-Whisper CTranslate2 release (sorika-labs/vaani1-lite-asr-turbo).

• Vaani-1.1 Lite ASR (~600M): Built on Qwen3-ASR-0.6B-hf (Conformer audio encoder + Qwen3 LLM reasoning decoder). Captures #1 SOTA on Indian Speech with an 11.78% Hindi Average WER. Excels in noisy acoustic environments, achieving 12.36% on Kathbath Hindi and an extraordinary 3.60% on LibriSpeech Other.

Flash Series (High-Throughput APIs & Real-Time Streaming):

• Vaani-2 Flash ASR (809M): Built on Whisper-large-v3-turbo with a 4-layer fast decoder. Achieves #1 Overall English Accuracy (5.90% Avg WER across 7 public benchmarks) while delivering an insane 21.5x to 30.3x Real-Time inference speed (RTF). It processes 30 seconds of speech in under 1 second on standard GPU.

• Vaani-2.1 Flash (In Active Development): Next-generation roadmap target focused on sub-10% Hindi WER and native Hinglish code-switching without losing the 25x+ RTF velocity.

EVALUATION MATRIX · 10 BENCHMARKS

The Grand Benchmark Leaderboard

Benchmark DatasetDomainWhisper-Small
244M
Vaani-1 Lite
244M
Vaani-2 Flash
809M Turbo
Qwen3-Base
600M
Vaani-1.1 Lite
600M LLM
Flash Speed
AMI-CleanedEnglish (Meetings)12.74%12.99%10.41%9.05%15.13%29.4x RTF ⚡
Earnings22English (Financial Calls)8.62%9.90%8.62% 🥇10.99%10.50%30.3x RTF ⚡
GigaSpeechEnglish (Audiobooks/Pods)8.90%8.43%7.28%6.77%7.58%19.1x RTF ⚡
LibriSpeech CleanEnglish (Clean Read)5.63%9.67%4.84%4.49%4.75%15.9x RTF ⚡
LibriSpeech OtherEnglish (Noisy / Accents)11.30%8.70%4.35%3.93%3.60% 🥇15.3x RTF ⚡
SPGISpeechEnglish (Financial)3.89%5.41%3.55% 🥇4.74%5.58%18.9x RTF ⚡
VoxPopuliEnglish (Parliamentary)5.29%9.20%2.22% 🥇3.38%8.99%28.1x RTF ⚡
Kathbath (Hindi)Hindi (Indic Crowdsource)50.00%16.29%22.47%14.61%12.36% 🥇5.9x RTF
FLEURS (Hindi)Hindi (Conversational STT)48.88%25.56%16.38%14.39%14.64%8.5x RTF
Vaani Private TestHindi (Real Conversations)44.44%16.11%14.44%7.78%8.33%6.8x RTF
English Average WER7 English Benchmarks8.05%9.18%5.90% 🏆6.19%8.02%21.5x Avg 🚀
Hindi Average WER3 Indic Benchmarks47.77%19.32%17.76%12.26%11.78% 🏆7.1x Avg
Overall Average WERAll 10 Benchmarks19.97%12.23%9.46%8.01%9.15%17.2x Avg

*Note: Lower Word Error Rate (WER) represents superior transcription accuracy. Evaluated on identical audio splits using the standard Whisper Text Normalizer. 🥇 indicates benchmark winner; 🏆 indicates grand category champion.

DEVELOPER INTEGRATION

Interactive Inference Snippets

sorika-labs/vaani2-flashPython 3.10+
from transformers import pipeline

# Load Sorika Labs Vaani-2 Flash (Fused SafeTensors, Zero Adapter Overhead)
pipe = pipeline("automatic-speech-recognition", model="sorika-labs/vaani2-flash", device="cuda")

# 21.5x - 30.3x Real-Time Inference Speed (30s audio in < 1 second)
result = pipe("speech.wav")
print(result["text"])
OFFICIAL OPEN-SOURCE WEIGHTS & REPOSITORIES100% OPEN SOURCE

Deploy Open-Source Vaani from Hugging Face Hub

All foundation checkpoints are released as 100% permanently fused model.safetensors with zero adapter loading overhead. Available under permissible open-source terms at huggingface.co/sorika-labs.

Direct Model Repositories: