User avatar
le gros tung tung a l'air @penny@social.xenofem.me
2mo
new gemma is cool
1
0
1
0
User avatar
le gros tung tung a l'air @penny@social.xenofem.me
2mo
got a solid 4 t/s boost on a 3060 with the boosted vram
2
0
1
0
User avatar
Two Hollywood Phonies @autumn@cafe.autumn.town
2mo
@penny been using it too its awesome, 35 t/s with the moe model, full 256k ctx enabled (but not filled)
1
0
1
0

User avatar
le gros tung tung a l'air @penny@social.xenofem.me
2mo
@autumn thats the 26b one right
1
0
0
0
User avatar
Two Hollywood Phonies @autumn@cafe.autumn.town
2mo
@penny yeah yeah its been great
1
0
1
0
User avatar
Two Hollywood Phonies @autumn@cafe.autumn.town
2mo
@penny i used the regular one and now the qat and theres a big diff in token speed, i was messing around with it today lol
2
0
1
0
User avatar
Two Hollywood Phonies @autumn@cafe.autumn.town
2mo
bunch of perf numbers @penny i actually got some diff results too, direct comparison between gemma4:26b & QAT counterpart with 256k context window was like 20t/s vs 25t/s, 128k was 23t/s and 30t/s for QAT.

then i tried against qwen3.6 MTP vs gemma4 QAT on smth with actual tokens, 256k ctx for both. qwen = 25t/s and gemma was 35 t/s

That's all using ollama w/ rocm, gonna try ollama vulkan then llama.cpp tomorrow
0
0
1
0
User avatar
le gros tung tung a l'air @penny@social.xenofem.me
2mo
@autumn fuckkkk ive been stuck with like 20 t/s
1
0
0
0
User avatar
Two Hollywood Phonies @autumn@cafe.autumn.town
2mo
@penny ive got an rx 6700xt with 12gb of VRAM so thats probably the difference lmao
1
0
1
0
User avatar
le gros tung tung a l'air @penny@social.xenofem.me
2mo
@autumn ive got a 3060 with the same vram but the clock speeds are total shit
0
0
1
0