I've been playing around with testing the Qwen3.x series on a 5060 Ti 16GB and an M1 Max 32GB using Ollama, and my initial impression from a few test runs was "It's doing better than I thought?"
However, when I actually used it for longer tasks like real work, the story changed.
Even though it seemed to finish normally a few times, if the workload increased even slightly to resemble a real work environment, failures would occur. Even when it said the task was completed, I sometimes wondered if it was truly finished correctly. There were also instances where it got stuck in the middle and I had to cancel and restart it.
My conclusion was: "It's fun for testing, but not yet ready for real use."
Then I tried running Qwen3 27B EXL3 on my RTX 5060 Ti with ExLlama, and the feeling was quite different.
The most noticeable difference was that it ran consistently. It wasn't just running well once or twice; it kept running even when I repeated similar tasks. For the first time, I thought, "This is good enough to actually use?"
Especially for someone like me who does more than just simple chatting, such as code analysis, modification, and multiple steps of continuous work, this difference is quite significant.
ExLlama definitely feels easier to use for real applications compared to Ollama at the moment.
Of course, there are drawbacks. ExLlama doesn't allow you to use any model as is; you have to use models that have been converted to a supported format like EXL3.
Still, it's quite interesting that even with relatively limited VRAM of 16GB on an RTX 5060 Ti, we can run a 27B model this stably.
It's not perfect yet,
"we've moved from the stage of testing local LLMs to the stage where they can be used for real work."
I plan to continue using Ollama and ExLlama alternately to see the differences based on actual work in the near future. 😄