huh the llama.cpp vulkan branch seems to be able to run qwen3.5 9b at speeds id expect from rocm, on my rx580. may be because i downloaded q4_0 instead of the q4_k quant (this is documented to be slower on vk). i genuinely thought the difference was vk compute was lower perf but maybe rocm literally has no reason to exist other than making me get pissed off
gonna make it wipe my hard drive with the exec shell command tool
i used to be getting barely 15t/s with pretty much this exact setup a few months ago. this is worrying me. i may actually start using the damn thing for real. this is worryingly usable (beyond the fact that it doesn't know shit, because it's a 9b model)
i want to see what the deal with "vibe coding" is. surely a model this small isn't gonna do any good but at least i'm not paying anyone but the electricity guys
[60679] radv/amdgpu: The CS has been cancelled because the context is lost. This context is guilty of a soft recovery.
[60679] terminate called after throwing an instance of 'vk::DeviceLostError'
[60679] what(): vk::Queue::submit: ErrorDeviceLost
it completely ignored it and decided to re-write the file as part of something else instead of my system prompt to edit the file in place (in a vain attempt to save tokens, which i still need to do despite it being local because it takes so fucking long)
@autumn i do Not trust literally Any llm tooling. i feel uneasy using the modelcontextprovider/node typescript sdk but i only convinced myself to do so because it has like two dependencies