I reset 2spark with qwen3.8-flash-next sglang.

58.0.***.***
13

When I first did it with 1 spark, the bench came out around 24~26, and in actual use it came out around 16~20.

And

When I did it with 2 sparks, it went up almost twice as much.

However, in actual use, it came out to about 25~33 tokens, so there was a difference from the bench.

And today I tried setting it up with sglang.

There is a difference between vLLM and sglang.

It seems that sglang is currently making a difference of about 10~20%.

However, in actual use, DeepSeek feels faster. The difference is bigger than the number of tokens.

로그인한 회원만 댓글 등록이 가능합니다.

개발한당

KR | ID | EN
  • IDR
  • KOR
7.66 0.01

2026.09.04 KEB 하나은행 고시회차 955회

다가오는 한인 행사일정

  • 등록 된 일정이 없어요!