OpenAI cuts the price of its small AI by 80% three weeks after launch

OpenAI cuts the price of its small AI model by 80%, three weeks after its launch


OpenAI launched its family of models GPT-5.6 on July 9. On July 30, the company reduced the price of the smallest of the three by five times. Three weeks between the two. Just enough time to finish a half-eaten pot of yogurt.

Producing the equivalent of seven novels with this model used to cost 6 dollars. Now, it costs 1.20 dollars. And this drop is not a holiday gift.

What has changed, in numbers

Let’s start with a word that keeps coming up: the token. It’s a piece of a word that the model ingests and spits out. One million tokens represents, very roughly, 700,000 words. That’s seven or eight novels. You pay separately for what you give the model to read, the input, and what it writes, the output. The output always costs much more.

Bar chart comparing output price per million tokens before and after July 30, 2026 for GPT-5.6 Luna, Terra, and Sol

The small model got hit by the sales. The high-end model, however, hasn’t budged a cent.

Here are the details of the three ranges:

  • Luna, the small fast model: from 1 dollar to 20 cents for input, and from 6 dollars to 1.20 dollars for output. An 80% drop.
  • Terra, the mid-range model: from 2.50 dollars to 2 dollars for input, and from 15 dollars to 12 dollars for output. A 20% drop.
  • Sol, the high-end model: still 5 dollars for input and 30 dollars for output. Nothing. Not a cent.

This last point says more than the other two. The luxury segment doesn’t go on sale as long as no one offers the same thing cheaper right next door.

The official explanation, and the one we can guess

OpenAI explains this drop by its technical progress: a 20% reduction in total operating costs, and an efficiency gain of over 15% in token production. This is probably true. But it doesn’t explain why the smallest model accounts for almost the entire drop.

The other explanation can be summed up in one name: DeepSeek. Its V4 Flash version costs 14 cents for input and 28 cents for output. At this level, a model charged 6 dollars per million tokens for output can no longer defend itself with a nice speech. It must revise its pricing. The price gap in the market has become frankly absurd. The most expensive cutting-edge model costs 180 dollars per million tokens produced, compared to 28 cents for DeepSeek V4 Flash. That’s 643 times more expensive. Yet, in both cases, the end customer mainly sees text coming out of a box.

Personally, I find this healthy. For two years, we were told that artificial intelligence was too expensive to be sold off. It only took one serious competitor for the price to collapse overnight.

The real good deal is hidden in the footnote

There’s something better than the announced drop, and almost no one talks about it. GPT-5.6 charges 90% less for a text it has recently read. This is called caching, a temporary memory that prevents the model from rereading the same thing multiple times at full price.

Illustration in two halves comparing the complete rereading of a thick book at each question and the same reading with bookmarks, with the bill shrinking

On the left, the model rereads your file from the first page at each question. On the right, it has placed bookmarks. The difference is especially noticeable at the bottom of the bill.

Specifically, when you chat with an assistant, the entire history of the conversation is sent back to it with each new message. Your previous question, its answer, the document you pasted, everything starts from the beginning. Without caching, you pay this rereading at full price, again and again. With caching, the part that hasn’t changed goes from 20 cents to 2 cents per million tokens with Luna.

There is an honest trade-off to note. The first caching of a text costs 1.25 times the normal rate. The operation only becomes profitable starting from the second reading. For a conversation or an agent, that is to say a program that works autonomously on the same file for an hour, this second reading comes almost immediately.

Specifically, what does it change for you

You may not be paying OpenAI directly by the million tokens. These rates may seem very far from your daily life. Yet, they determine the functions that applications can offer you, and at what price.

Two panels, on the left a single phone with a padlock and a price tag, on the right a whole house where the television, fridge, car, and phone each have a small spark

When the raw material costs almost nothing, we put it everywhere. Including where no one asked for it.

1. Paid functions become included functions

A small business selling an application for 3 euros per month cannot easily add an AI that costs 40 cents per user per month. It reserves this function for a premium subscription, or it does without it. Divide this cost by five and the calculation changes completely. The function can then enter the basic offer.

This is what could happen to the summary of your emails, to the translation in your recipe application, to photo retouching, or to sorting your documents. Nothing spectacular, taken separately. But dozens of small functions can stop being paid options.

2. AI descends into inexpensive objects

As long as a request is expensive, we reserve AI for products for which the customer pays a lot: enterprise software, high-end cars, or professional subscriptions. At 20 cents per million tokens, we can start putting it in an internet box, an intercom, a supermarket checkout, or a toy.

It's the same story as the electronic chip. In the 70s, a microprocessor cost the price of a used car and lived in an air-conditioned room. Today, there are about thirty in your kitchen. Most perform missions as glorious as making the oven clock blink. When the price collapses, the product does not just become cheaper. It appears in places where no one would have thought to put it before.

3. Rising bills will stop rising

This is the most down-to-earth consequence. If you pay a subscription that uses AI, your provider may have just seen their own bill shrink. They probably won't give you the difference back; no one does that. But they can stop raising their prices and compete on the functions included in their offer.

Let's keep a reserve, because we are not selling dreams. The decrease concerns the small model. The really difficult tasks use the high-end, whose price has not changed. So we shouldn't expect a miracle everywhere. The fairly simple tasks will rather gradually shift to the small model. And as today's small model quickly becomes better than last year's big one, this eventually concerns many tasks.

What I think about it

What strikes me is not the amount of the decrease. It's the timing. Just three weeks between the launch of a model and the division of its price by five. In this sector, a pricing grid now has the lifespan of a baguette.

For those building a product with these tools, the lesson is very concrete: never build a business plan on today's rate, and never bet on a single supplier. The price will continue to drop. As for the model you are using today, it could end up in the clearance bin before Christmas.

For everyone else, there is good news hidden in there. Artificial intelligence, once costly to integrate into your phone services, is becoming one of the cheapest elements. What will remain costly is knowing what to do with it.

Join the conversation

You need an account to comment on this article. Creating one is free and takes under a minute.

  • The XMLTV file, free to download every day
  • Comment on articles and reply to other readers
  • Get an e-mail when an article you follow is updated

No comments yet.

Une erreur s'est produite. Cette application peut ne plus répondre jusqu'à ce qu'elle soit rechargée.Veuillez contacter l'auteur. Reload 🗙