I finished fine-tuning two stock models with Spark, and I was wondering how far I could push this thing??? So I found one with almost 128 billion parameters.

The model itself is about 80GB, but when I run it, the KV cache keeps sucking up RAM.
I uploaded it to Olllama and checked with --verbose, and it came out to about 28.8 tokens.
It's a 122 billion model, but it's moe act10b, so I wonder if this much token generation is normal.

So I tried asking Claude Code.
It was really slow at first, but I connected searxng to enable web search and gave it all the permissions, and it started working well on its own.
I think I can run it locally and do other things at the same time.
Honestly, I don't have high expectations.
Claude Code is a tool tailored for sonnets or operas, and this one is from a different company, so I lowered my expectations a bit.
Still, it's better than just using it raw. It has some harnessing, so it seems better than connecting to OpenCLOE.

I told it to look for an open-source MMORPG.
Maybe because of the model difference, I don't think it can use permissions freely.
Later, if the price of Spark drops or I find a good used one, I want to try using the DeepSpeed 4Flash version.