
I tested it on dgx spark.
Since HuggingFace uses vllm, I set it up and used it, and it came out to 19~20 tokens.
It also threw errors.
Today, when I went to the llama homepage, it said it was uploaded yesterday, so I downloaded it and did a functional test.
It took some time to upload to spark's ram, and since the initial token count was 9, I thought I couldn't use it, but I tried several times just in case, and the token generation speed gradually increased.
Seeing that about 20~22 are generated, if vllm or other optimization tools come out later, it seems like it will generate more than 30.
If it's 20, it's the perfect generation speed to entrust work before bed and receive the results the next day.
I hope for the development of llama, as qwen3.6-27b also came out to about 22 when it was optimized.????