ResearchModel Research
Aug 2, 2026·6 min read
HF Model Benchmark: Trithemius Fable-Fusion 711 PrismAURA 5.5-bit
Full benchmark of trithemius/Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-MTP-PrismAura-5.5bit from HuggingFace. RTX 5090 performance study: 2x MTP decode speedup (100.6 vs 52.2 t/s), 262K dense / 215K MTP context, vLLM compressed-tensors.
NVIDIA RTX 5090 (32GB VRAM)trithemius/Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-MTP-PrismAura-5.5bit
2x decode speedup from MTP speculative decoding: 100.6 t/s vs 52.2 t/s @ 2K context
262K context window on dense (No-MTP) config, 215K with MTP heads on single RTX 5090
600W power cap buys nothing: decode flat (MTP ±5% noise), prefill +3-4% only — 500W cap stays
by Arjun NayakRead