I tried using the newly released llm tool with qwen3.8-flash-next spark alone.

59

https://github.com/HawkBearPig/dgpp

I saw this tool mentioned on the NVIDIA forum in the DGX Spark board and decided to try it out.

It seems like it only supports a few models right now, but luckily qwen3.8-flash-next supports both 1 spark and 2 spark, so I set it up with 1 spark.

This is roughly the TPS I got.

It gets around 40 TPS alone, so I expect it to get around 20~26 for actual use.

When I used vllm before,

1 spark got around 20 TPS.

2 spark worked a little better than the new tool.

Theoretically, if I use 2 spark with this tool now, I could get around 70~80????

I'm not sure yet.

로그인한 회원만 댓글 등록이 가능합니다.

개발한당

KR | ID | EN
  • IDR
  • KOR
7.52 ▼ -0.01

2026.10.09 KEB 하나은행 고시회차 1861회

다가오는 한인 행사일정

  • 등록 된 일정이 없어요!