Last year I got into local LLMs, changed my graphics card, and had a lot of fun trying out various things,
but once I tried something a bit bigger, I found 20-30B class models to be far too inefficient.
It's not that it's impossible, it's just too slow, and the results were also lackluster.
On top of that, with Chinese LLMs led by DeepSeek recently becoming very cheap in price,
the financial burden has also become much lighter, so
I removed several things I'd installed to run local LLMs on my system.
Originally I merged 2 PCs into one to run local LLMs.
2 GPUs, 2x memory.. it was fun, but there wasn't really a specific use case for it.
After thinking it over carefully, I figured just paying for a service and using it comfortably would be better in many ways,
so I went back to using 2 separate PCs.
On top of that, I used to have ComfyUI installed for generating images,
but I switched to just paying FAL.ai to generate images instead.
Using paid services for everything turned out to be better for my mental health.
I use the hermes agent, and this hermes agent offers a free model.
These days they're offering Step3.7-flash for free. Still, I use Deepseek-v4 though...
Models like Gemma-4-26B-A4B or Gemma-4-31B are
offered for free on Google AI Studio.
In many ways, in the field where I'm just messing around for fun, there's no need to use a local LLM anymore,
so I decided to boldly clean it all up this time.