bunch of perf numbers@penny i actually got some diff results too, direct comparison between gemma4:26b & QAT counterpart with 256k context window was like 20t/s vs 25t/s, 128k was 23t/s and 30t/s for QAT.
then i tried against qwen3.6 MTP vs gemma4 QAT on smth with actual tokens, 256k ctx for both. qwen = 25t/s and gemma was 35 t/s
That's all using ollama w/ rocm, gonna try ollama vulkan then llama.cpp tomorrow