SORIKA
RESEARCH

Sorika Labs is an artificial intelligence research and product studio founded by Darsh Yadav. We are building a future where neural intelligence is lightweight, sovereign, deterministic, and deeply aligned with human thought.

While AI capabilities have advanced dramatically, key gaps remain. Massive discrete foundation models increasingly suffer from non-deterministic hallucinations, severe serial latency bottlenecks, and quadratically growing memory consumption. Knowledge of how these architectures operate remains concentrated, limiting public scientific discourse and production reliability.

To bridge these gaps, we research and engineer non-autoregressive neural architectures that achieve broadcast fidelity with strictly constant memory complexity and sub-20 millisecond execution.

Learning by doing

Research and product co-design. Products enable iterative learning through real-world deployment, while foundational research strengthens products. We operate as an agile studio where theoretical mathematics directly becomes shipping software.

Non-autoregressive reliability. The most critical systems require 100% deterministic execution. Rather than relying on unconstrained stochastic next-token codebooks, our acoustic and perception engines eliminate word skips and hallucinations by design.

Measure what truly matters. We focus on practical efficiency metrics: time-to-first-audio below 20ms, constant $O(1)$ memory footprints, and multi-resolution spectral fidelity that runs on consumer hardware.

Publications & Foundation Models

Open technical reports, architecture specifications, and empirical benchmarks.

PRIMARY RESEARCH REPORT #002HYBRID CAUSAL CONVOLUTIONS & GQASEPTEMBER 2026

Swen 1.1: Efficient Edge Intelligence via Hybrid Double-Gated Convolutions and Test-Time Metacognitive Reasoning

Darsh Yadav (Founder & Chief AI Architect, Sorika Labs) with the Sorika Labs Research Team

We address the quadratic complexity and KV cache memory wall of traditional Transformers on edge devices. By interleaving 10 double-gated causal convolution layers with 6 grouped-query attention layers, Swen 1.1 delivers 128k context with 10.6× less cache memory footprint while outperforming models 2×–3× its size across HumanEval, GSM8K, and MultiArith.

Pretraining28T Tokens
RAM Footprint10.6× Reduction
HumanEval Pass@166.0% 🥇
MultiArith85.0% 🥇
PAPER #001 • SPEECH INTELLIGENCE

Suzune S1: Non-AR TTS

Parameters80M Non-AR
Architecture12-L PL-BERT + iSTFTNet
Audio Fidelity24kHz Broadcast Quality
Inference RTF0.018x GPU / 0.087x CPU
Memory Footprint~320 MB constant O(1)

Ultra-fast non-autoregressive speech synthesis engine with 12-layer PL-BERT, continuous AdaLN prosody, and NSF harmonic vocoding across English and Hindi.

TECHNICAL NOTE • PERCEPTION

Kaori: Spatial Layout OCR

Parameters120M Spatial Backbone
ArchitectureMulti-Scale Spatial Transformer
Input ModalityHigh-Res Multi-Column PDFs
Layout Accuracy99.4% Structure Recovery
LanguagesLatin, Devanagari, Historical

Layout-guided attention for zero-shot extraction of complex tables, mathematical formulas, and historical Indian manuscripts.

Under Development