User avatar
kopper colon_three @kopper@not-brain.d.on-t.work
3mo
huh the llama.cpp vulkan branch seems to be able to run qwen3.5 9b at speeds id expect from rocm, on my rx580. may be because i downloaded q4_0 instead of the q4_k quant (this is documented to be slower on vk). i genuinely thought the difference was vk compute was lower perf but maybe rocm literally has no reason to exist other than making me get pissed off

gonna make it wipe my hard drive with the exec shell command tool
6
0
1
0
User avatar
kopper colon_three @kopper@not-brain.d.on-t.work
3mo
ok this is getting absurd now how am i htiting 24t/s on a 9b model that just isn't supposed to happen. it's not natural. is the rx580 this goated.
2
0
2
0
User avatar
kopper colon_three @kopper@not-brain.d.on-t.work
3mo
i used to be getting barely 15t/s with pretty much this exact setup a few months ago. this is worrying me. i may actually start using the damn thing for real. this is worryingly usable (beyond the fact that it doesn't know shit, because it's a 9b model)
1
0
1
0
User avatar
kopper colon_three @kopper@not-brain.d.on-t.work
3mo
yeah sure i'm running a 26b-a3b model with the same speed i used to run a 8b model ok fine that tracks. card dropped by amd due to being old btw.
1
0
1
0
User avatar
kopper colon_three @kopper@not-brain.d.on-t.work
3mo
i want to see what the deal with "vibe coding" is. surely a model this small isn't gonna do any good but at least i'm not paying anyone but the electricity guys
1
0
1
0
User avatar
kopper colon_three @kopper@not-brain.d.on-t.work
3mo
(none of this will end up online btw)
1
0
1
0
User avatar
kopper colon_three @kopper@not-brain.d.on-t.work
3mo
thanks amd
[60679] radv/amdgpu: The CS has been cancelled because the context is lost. This context is guilty of a soft recovery.
[60679] terminate called after throwing an instance of 'vk::DeviceLostError'
[60679]   what():  vk::Queue::submit: ErrorDeviceLost
1
0
1
0
User avatar
kopper colon_three @kopper@not-brain.d.on-t.work
3mo
this is the worst error code for "it timed out 4head" github.com/ggml-org/llama.cpp/issues/21724
1
0
1
0
User avatar
kopper colon_three @kopper@not-brain.d.on-t.work
3mo
its going great
1
0
1
0
User avatar
kopper colon_three @kopper@not-brain.d.on-t.work
3mo
it removed the file it just wrote i wonder how long it's gonna take till it notices
1
0
0
0
User avatar
kopper colon_three @kopper@not-brain.d.on-t.work
3mo
it completely ignored it and decided to re-write the file as part of something else instead of my system prompt to edit the file in place (in a vain attempt to save tokens, which i still need to do despite it being local because it takes so fucking long)
1
0
0
0
User avatar
kopper colon_three @kopper@not-brain.d.on-t.work
3mo
i restarted like 10 times i can feel my iq going so low i'm starting to believe iq is real. i should take a break
1
0
1
0
User avatar
kopper colon_three @kopper@not-brain.d.on-t.work
3mo
lobotomized, just like me <3
2
0
2
0
User avatar
Two Hollywood Phonies @autumn@cafe.autumn.town
3mo
@kopper what ui is that?
1
0
1
0

User avatar
kopper colon_three @kopper@not-brain.d.on-t.work
3mo
@autumn the one built into llama.cpp lmao
1
0
1
0
User avatar
kopper colon_three @kopper@not-brain.d.on-t.work
3mo
@autumn i do Not trust literally Any llm tooling. i feel uneasy using the modelcontextprovider/node typescript sdk but i only convinced myself to do so because it has like two dependencies
2
0
1
0
User avatar
kopper colon_three @kopper@not-brain.d.on-t.work
3mo
@autumn and before that i spent literally the entire day trying to get the models to write the MCP protocol from scratch. would not do. i gave them the documentation and they said fuck off and made some shit up
2
0
1
0
User avatar
Two Hollywood Phonies @autumn@cafe.autumn.town
3mo
@kopper sounds about right.
0
0
1
0
User avatar
Xbox 370 @eblu@wetdry.world
3mo
@kopper @autumn model context protocol protocol
0
0
1
0
User avatar
Two Hollywood Phonies @autumn@cafe.autumn.town
3mo
@kopper oh right lmao it seems more comprehensive than openwebui
0
0
1
0