The short answer is no.
I have played around with ollama and whisper. It’s just too slow to be practical. The cost of the hardware is preclusive.
That said, I do selfhost openwebui and use inference end points from huggingface and ovh.
I’ve never used chatgpt or claude and I have to wonder whether those alternatives are really as terrible as the models available on huggingface. The output is always super plausible but usually just plain wrong.


Kagi.
Searxng too much effort.