I saw this thing called colibri come out, tried it, and it worked so effortlessly well
that I just let out a hollow laugh...
Sure, the speed is brutally slow, but the fact that it "works" at all is a huge deal on its own.
And no complicated setup needed, no messing around configuring an NVIDIA CUDA environment or anything like that, dead simple
I feel like I could just kick something off in the background for a long time and have it handle stuff like this.
Do I need to cancel my opencode go subscription now? lol
Ref
Setup
# Download and set up the colibri repo
# No external dependencies, so it finishes instantly
cd ~/github
git clone <https://github.com/JustVugg/colibri> && cd colibri/c && ./setup.sh
# Download the dedicated GLM 5.2 model
# It's around 350-370GB so it takes forever
hf download mateogrgic/GLM-5.2-colibri-int4-with-int8-mtp --local-dir ~/glm52
# Run a test
# Running it on pure raw CPU, no GPU
COLI_MODEL=~/glm52 ./coli chat
# If the test works fine, time to serve it!
COLI_MODEL=~/glm52 ./coli serve --host 127.0.0.1 --port 8008 --model-id glm-5.2-colibri