https://x.com/bleysg/status/2083448647908569321
I tried "give this post to the agent".
Theoretically, it's a maximum of 59 tokens, but the agent said that when multiple requests come in, it shows the maximum number.
I installed and connected it to open source just in case.

This is the TPS token generation speed when you first set it up and connect.
These days, I think Thinking High has the best cost-performance among LLMs, so I switched to High and measured it.

Token generation has increased a bit.
So, I thought that if I used NonThinking, more tokens would be generated, so I tested it.

The agent said that this is the maximum due to bottlenecks in Spark, but Thinking High seems to be the best.
It's too slow for conversation, and it might be okay to leave it running, but I don't know about the quality. They say that the 0731 weight came out very well, so even if you quantize it with 2 bits, it seems like it will be better than the existing DeepSpeed4Flash.
I need one more Spark ㅠㅠ