Can you cite your source on the claim that “inference is currently insanely profitable”? Everything I read suggests that openai and anthropic lose money on their plans.
My caveats were clearly stated… After capital expenditure, it’s just operational costs, where electricity & cooling are the big ones.
At that point, it is insanely profitable to serve. The cheap API prices on open weights models hints at the profit margins involved in the US (the frontier labs and hyperscalers don’t open their books for us), unsurprisingly)
Therefore, the longer they can serve existing and lower cost models at the current rates, the better for their bottom line. It’s just common sense in business.
It doesn’t mean the company as a whole is profitable. I expect we’ll see turmoil in the coming months and years, and the prize will be compute capacity, with electricity & cooling options.
You can’t just write off capital expenditure though. The hardware, even for “effecient” MOE inference is still very expensive to buy, house, run, and cool. Even assuming open-weight model serving at $0 r&d for the models themselves, mixing high-prefill workloads doesn’t batch well with decode heavy concurrency (or other prefill-heavy jobs). The moment you do anything nontrivial you start running into very complicated architectural problems to efficiently solve at scale.
Hardware that is useful for 5-10 years at most, plus development and support for the inference workflows, doesn’t leave a lot of margin on the table.
My gut, along with basically everything I read, suggests that not most (even pure inference) shops are not profitable and are still floating on loans or vc money.
Can you cite your source on the claim that “inference is currently insanely profitable”? Everything I read suggests that openai and anthropic lose money on their plans.
My caveats were clearly stated… After capital expenditure, it’s just operational costs, where electricity & cooling are the big ones.
At that point, it is insanely profitable to serve. The cheap API prices on open weights models hints at the profit margins involved in the US (the frontier labs and hyperscalers don’t open their books for us), unsurprisingly)
Therefore, the longer they can serve existing and lower cost models at the current rates, the better for their bottom line. It’s just common sense in business.
It doesn’t mean the company as a whole is profitable. I expect we’ll see turmoil in the coming months and years, and the prize will be compute capacity, with electricity & cooling options.
You can’t just write off capital expenditure though. The hardware, even for “effecient” MOE inference is still very expensive to buy, house, run, and cool. Even assuming open-weight model serving at $0 r&d for the models themselves, mixing high-prefill workloads doesn’t batch well with decode heavy concurrency (or other prefill-heavy jobs). The moment you do anything nontrivial you start running into very complicated architectural problems to efficiently solve at scale.
Hardware that is useful for 5-10 years at most, plus development and support for the inference workflows, doesn’t leave a lot of margin on the table.
My gut, along with basically everything I read, suggests that not most (even pure inference) shops are not profitable and are still floating on loans or vc money.
If you assume they are unprofitable, the Q only becomes whether they are more or less unprofitable by serving the older models for longer.