openmic.social is an uncensored community. You may encounter strong language, controversial opinions, and mature or NSFW material. You must be 18+ to browse. Illegal content is prohibited and removed on sight — please report it. By continuing, you accept that you may see content you personally disagree with.
You’re talking about the cost to the consumer, where I was talking about the energy cost of an individual unit of work by a given model, but they’re both relevant. The underlying energy cost of generating a token at a given level of capability has been falling as hardware and inference become more efficient. Those efficiency gains can feed through into lower token prices for users, but also a more capable model can often complete the same task with fewer tokens, fewer retries and less prompting.
So even if the headline price of a new frontier model looks similar or higher, the actual cost, both in energy and money, of getting a given piece of work done can still fall substantially. It’ll keep doing so too, its still in its infancy.