I'm trying out Qwen-3.5:122B with the dgx spark.

208

I finished fine-tuning two stock models with Spark, and I was wondering how far I could push this thing??? So I found one with almost 128 billion parameters.

The model itself is about 80GB, but when I run it, the KV cache keeps sucking up RAM.

I uploaded it to Olllama and checked with --verbose, and it came out to about 28.8 tokens.

It's a 122 billion model, but it's moe act10b, so I wonder if this much token generation is normal.

So I tried asking Claude Code.

It was really slow at first, but I connected searxng to enable web search and gave it all the permissions, and it started working well on its own.

I think I can run it locally and do other things at the same time.

Honestly, I don't have high expectations.

Claude Code is a tool tailored for sonnets or operas, and this one is from a different company, so I lowered my expectations a bit.

Still, it's better than just using it raw. It has some harnessing, so it seems better than connecting to OpenCLOE.

I told it to look for an open-source MMORPG.

Maybe because of the model difference, I don't think it can use permissions freely.

Later, if the price of Spark drops or I find a good used one, I want to try using the DeepSpeed 4Flash version.

로그인한 회원만 댓글 등록이 가능합니다.

개발한당

KR | ID | EN
  • IDR
  • KOR
7.52 ▼ -0.01

2026.10.09 KEB 하나은행 고시회차 2840회

다가오는 한인 행사일정

  • 등록 된 일정이 없어요!