
Existing other cloud models were structured to reduce usage while running on their own servers, but kimi-k3 is just paid. Is it because the parameters are too large to run...? Other companies have been doing things like model usage 2x events these days, but there's nothing like that at Olllama Cloud, which is a bit disappointing.
As models get bigger and token usage increases, I wonder if I should keep using Olllama Cloud...
But Olllama is convenient...