HF Model Benchmark: DavidAU Fable-Fusion-711 MTP Q6_K · q8_0 KV · 200K ctx
Full benchmark of DavidAU/Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-NEO-MAX-MTP-GGUF (Q6_K MTP) with q8_0 KV cache on a single RTX 5090 via llama.cpp. 200K context window. Decode 62-104 t/s across 2K-180K context, prefill 1.2-2.3K t/s, MTP speculative decoding.