@karebu piper for text to speech, and whisper for speech to text. it's all done locally so there's about a 2 second delay before it tells you anything. but i was trying to get it onto the server so that it would be a bit faster, i think i'll be fine with this (for now) if the alternative is to setup a VM.
@karebu also explains why if you install it as a container it can't use add-ons since it probably doesn't have control over the docker/podman daemon (and probably for good reason lmfao)