I googled and asked it to write a prompt for testing agents
[Role]
You are a professional AI agent that assists with data analysis and schedule management.
[Objective]
Analyze the user's request and, if necessary, use the provided tools to gather accurate information, then answer clearly.
[Rules & Constraints]
1. Do not make up information you do not know; answer honestly that you do not know.
2. Always write answers in Korean, stating the key conclusion first and adding supplementary explanation afterwards.
3. Data containing numbers must be organized in table or list form.
[Workflow]
Step 1: Identify the intent of the user's input.
Step 2: Determine whether additional information (weather, calculation and so on) is needed to answer the question.
Step 3: Output the final answer to the user in a readable form.
[Test input scenario]
"Check the current weather in Seoul and judge whether it is good for planning a day's schedule. And also calculate the result of 3 times 15."
I gave thinking max to gpt5.6 luna, which is said to have cut its price sharply recently

I said hello first and then entered the prompt above, and it used 20,105 tokens
Next is deepseek4 pro. Its thinking maxes out at high

It clocks in at 22,267 tokens
Next is the latest deepseek4 flash 0731 with thinking max

It used 15,464 tokens. That is a token diet that stands apart from the other models. It is thinking max, but I cannot tell the difference in quality
Next is hy3, which was announced as being on a similar tier?? or a slightly higher tier than the existing deepseek4 flash. Its thinking maxes out at high

It used 22,528 tokens
As a bonus I tested qwen3.5-122b. This model cannot turn thinking on. When I measured tps 10 times at 2048 tokens, it averaged about 34 tokens. It is a so-so, usable level

It used 14,288 tokens, making it the model that went on the biggest token diet????? Partly because it has no thinking and partly because the model is small, but doing a good token diet is an ability too
Judging by current evaluations and benchmarks, deepseek4 flash 0731 offers the best value-for-money performance among the models I tested. gpt5.6 luna has also discounted a lot, so for people who feel resistance toward Chinese products or who would rather not have information go to China, that is good value too
It is an AI era where something becomes old in a month or two, so the latest is good, but China is too scary
The US too, rather than top performance at top price, will probably need to hold the high end firmly among low, mid and high tiers, and survive with specialized pricing plans for low and mid-tier products by company
In China the government pushes support such as electricity discounts and various other things, so aggressive pricing is possible. Labor costs also differ enormously from the US, but I hear that saving on server maintenance costs is where the money is
I hope Korea too puts out a decent model at a reasonable price. If the price is American and the performance is a Chinese model from two years ago, I think I would be very disappointed