Swen 1.1 Architecture Matrix
Comprehensive architectural specifications, empirical benchmarks, and deployment trade-offs across the three official Swen models architected by Darsh Yadav.
FLAGSHIP GENERALIST
Swen-1.1-Instruct
High-efficiency language model for code generation, agentic tool use & workflows.
Primary BenchmarkHumanEval: 66.0% 🥇
Scale / Weights1.17 Billion (~2.34 GB BF16)
Context Window128,000 Tokens
Best for: Real-time coding assistants, tool calling, agentic loops, long-context parsing
View Model Suite PURE MATH SPECIALIST
Swen-1-Math
Ultra-compact mathematical engine with autonomous symbolic scratchpads.
Primary BenchmarkMultiArith: 85.0% 🥇
Scale / Weights350 Million (~708 MB BF16 / FP32)
Context Window64,000 Tokens
Best for: On-device arithmetic, Olympiad math, financial calculation, pure CPU execution
View Model Suite METACOGNITIVE REASONER
Swen-1.1-Thinking
Test-time deliberative reasoning engine for self-reflection & backward induction.
Primary BenchmarkGPQA Diamond: 39.4% 🥇
Scale / Weights1.17 Billion (~2.34 GB BF16)
Context Window128,000 Tokens
Best for: Complex combinatorial game theory, multi-step logic proofs, scientific deduction
View Model Suite Scroll horizontally to inspect benchmark data
| Feature / Evaluation Metric | ● Swen-1.1-Instruct | ● Swen-1-Math | ● Swen-1.1-Thinking |
|---|---|---|---|
| Empirical Reasoning Benchmarks | |||
| HumanEval (Pass@1)Zero-shot Python code synthesis benchmark | 66.0% 🥇 | 28.4% | 62.5% |
| GSM8K (Grade School Math)Multi-step word mathematical problems | 54.0% | 44.0% (at 350M) | 56.8% 🥇 |
| MATH-500 (Olympiad-Level)Challenging competition-level high school math | 24.2% | 35.0% 🥇 | 38.1% 🥇 |
| MultiArith AccuracyDeterministic multi-step arithmetic correctness | 68.0% | 85.0% 🥇 | 78.5% |
| GPQA DiamondGraduate-level scientific and mathematical reasoning | 38.0% | 22.4% | 39.4% 🥇 |
| Game Theory & CombinatoricsBackward induction & impartial games (Nim, etc.) | 52.0% | 64.0% | 92.4% 🥇 |
| Architecture & Formulation | |||
| Parameter ScaleTotal non-embedding trainable weights | 1.17 Billion | 350 Million | 1.17 Billion |
| Context WindowNative uncompressed token processing length | 128,000 Tokens | 64,000 Tokens | 128,000 Tokens |
| Neural BackboneCore attention and convolution formulation | Hybrid Double-Gated Conv + Linear GQA | Linear Causal Convolutions | Deliberative Hybrid Conv + GQA |
| Scratchpad ProtocolIntermediate cognitive deliberation format | <|tool_call_start|> Agent Protocol | <|cot_start|> Algebraic Scratchpad | <think> ... </think> Metacognitive CoT |
| KV-Cache Memory ComplexityScaling behavior across long document prompts | O(1) Constant bounded buffer | O(1) Strict linear streaming | O(1) Bounded reflection cache |
| Weights FootprintDisk & RAM consumption in BF16 precision | ~2.34 GB | ~708 MB | ~2.34 GB |
| Hardware & Execution Profile | |||
| Time-to-First-Token (TTFT)Initial prompt processing and prefill response latency | < 15 ms | < 8 ms (Ultra-fast) | ~ 20 ms (Deliberation phase) |
| Decode ThroughputSustained generation speed (RTX 6000 Ada / M3 Max) | 80–90 tok/s | 140+ tok/s | 95 tok/s |
| Minimum System RAMRequired system memory for edge inference | 4 GB RAM | 1 GB RAM (Runs in L3 Cache) | 4 GB RAM |
| Optimal Silicon TargetRecommended deployment environment | MacBook, Laptops, Edge AI PCs | Pure CPU, Embedded Chips, Raspberry Pi | Workstation GPUs, RTX, Cloud Instances |
SORIKA RESEARCH PUBLICATION #002
Read the Full Swen 1.1 Technical Report
Authored by Darsh Yadav. Contains complete ablation studies on hybrid double-gated causal convolutions, test-time metacognitive reasoning, and exact benchmark reproductions.