Zosma AI · Research Lab

Technical Research &
Benchmark Studies

Open research from the Zosma AI hardware lab. Model performance benchmarks, inference optimization studies, and quantization experiments — all run on real hardware, all reproducible.

3
Studies Published
1
Research Areas
Filter
⚡ BenchmarkModel Research
HF Model Benchmark: DavidAU Fable-Fusion-711 MTP Q6_K · q8_0 KV · 200K ctx
Peak Decode Speed
103.8 t/s
MTP @ 2K context
Decode @ Full Depth
61.9 t/s
@ 180K context (−40%)
Peak Prefill
2,254 t/s
@ 2K context (cold)
119 t/s
0
1%25%50%75%90%
ZOSMA AI · RESEARCH LABRTX 5090 · vLLM · compressed-tensors
ResearchModel Research
Aug 4, 2026·5 min read

HF Model Benchmark: DavidAU Fable-Fusion-711 MTP Q6_K · q8_0 KV · 200K ctx

Full benchmark of DavidAU/Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-NEO-MAX-MTP-GGUF (Q6_K MTP) with q8_0 KV cache on a single RTX 5090 via llama.cpp. 200K context window. Decode 62-104 t/s across 2K-180K context, prefill 1.2-2.3K t/s, MTP speculative decoding.

NVIDIA RTX 5090 (32GB VRAM)DavidAU/Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-NEO-MAX-MTP-GGUF (Q6_K MTP, q8_0 KV)
200,000-token context on single RTX 5090 with q8_0 KV cache — stable through 90% depth sweep at 97.7% VRAM
Decode 61.9-103.8 t/s across context: 103.8 t/s @ 2K → 61.9 t/s @ 180K (−40%) with MTP speculative decoding
Prefill 1,184-2,254 t/s; peaks at 2K depth (2,254 t/s), TTFT 1.05s @ 2K → 152.2s @ 180K
by Arjun NayakRead
⚡ BenchmarkModel Research
HF Model Benchmark: Myric Laguna S 2.1 APEX i-compact (Q8, 120K ctx, q4_0 KV)
Peak Decode Speed
91.2 t/s
@ 2K context
Max Context
122,880
120K window, 3-GPU
Peak Prefill
1,741 t/s
@ 40K context (25%)
105 t/s
0
1%25%50%75%
ZOSMA AI · RESEARCH LABRTX 5090 · vLLM · compressed-tensors
Research02
Model Research

HF Model Benchmark: Myric Laguna S 2.1 APEX i-compact (Q8, 120K ctx, q4_0 KV)

Full benchmark of Myric/Laguna-S-2.1-APEX-GGUF (Laguna-S-2.1-APEX-i-compact.gguf) on a 3-GPU rig (RTX 5090 + 2x RTX 5060 Ti). Q8_0 weights, q4_0 KV cache, 122880-token context. Prefill 1402-1741 t/s, decode 43-91 t/s across the context sweep.

RTX 5090 (32GB) + 2x RTX 5060 Ti (16GB), 4:2:2 tensor splitMyric/Laguna-S-2.1-APEX-GGUF (Laguna-S-2.1-APEX-i-compact.gguf, Q8_0)
122,880-token (120K) context on 3 GPUs with q4_0 KV cache — stable through 75% depth sweep
Decode 43-91 t/s across context: 91.2 t/s @ 2K → 43.2 t/s @ 120K (−53%)
Aug 3, 2026·6 min read
Read
⚡ BenchmarkModel Research
HF Model Benchmark: Trithemius Fable-Fusion 711 PrismAURA 5.5-bit
Peak Decode Speed
103.0 t/s
MTP @ 161K context
Max Context Window
262,144
No-MTP (+47K vs MTP)
Peak Prefill
14,352 t/s
No-MTP @ 2K context (cold)
MTP (215K)No-MTP (262K)
118 t/s
0
1%25%50%75%100%
ZOSMA AI · RESEARCH LABRTX 5090 · vLLM · compressed-tensors
Research03
Model Research

HF Model Benchmark: Trithemius Fable-Fusion 711 PrismAURA 5.5-bit

Full benchmark of trithemius/Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-MTP-PrismAura-5.5bit from HuggingFace. RTX 5090 performance study: 2x MTP decode speedup (100.6 vs 52.2 t/s), 262K dense / 215K MTP context, vLLM compressed-tensors.

NVIDIA RTX 5090 (32GB VRAM)trithemius/Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-MTP-PrismAura-5.5bit
2x decode speedup from MTP speculative decoding: 100.6 t/s vs 52.2 t/s @ 2K context
262K context window on dense (No-MTP) config, 215K with MTP heads on single RTX 5090
Aug 2, 2026·6 min read
Read