Is It Cheaper to Use the API or the Subscription? The August 2026 Math

The Briefing, Aug. 19, 2026: the API or the subscription, which one is actually cheaper

Model prices fell again this summer. Your monthly bill didn’t. On August 10 Anthropic confirmed that Claude Sonnet 5’s introductory rate — $2 per million input tokens, $10 per million output — is now the standard price, and that the increase to $3/$15 scheduled for September 1 will not happen. OpenAI’s cheapest current model, GPT-5.6 Luna, sits at $0.20 in and $1.20 out per million tokens. Meanwhile the consumer plan you actually pay for is still $20 a month, the same number it has been for years.

That gap is why one search phrase keeps climbing: “is it cheaper to use the Claude API or the subscription.” It is a fair question, and most of the answers floating around get one thing badly wrong. Here is the arithmetic, with every price read off the vendor’s own page on August 19, 2026 — and an honest account of what the cheaper route takes away from you.

The short version. If you have fewer than about five long conversations a day, paying per token costs less than $20 a month — often far less. If you use voice, images, memory, or the phone app, the plan is buying you things the meter does not sell. And for a great many people the honest answer is a third one: the free tier, which now covers more than it used to.

What the two prices actually are, today

Start with the shelf prices. These are per user, per month.

Plan Price What the vendor’s page says
ChatGPT Free $0 “Unlimited text chats with GPT-5.6 Luna,” subject to abuse guardrails
ChatGPT Go $8 More messages, uploads, images, voice — “This plan may include ads”
ChatGPT Plus $20 Advanced reasoning with GPT-5.6, projects, scheduled tasks
ChatGPT Pro from $100 5× or 20× more usage, Sol Pro reasoning
Claude Free $0 A real free tier, not a timed trial
Claude Pro $17 annual · $20 monthly $200 billed up front on the annual plan
Claude Max from $100 Higher usage ceilings
Read off chatgpt.com/pricing and claude.com/pricing on August 19, 2026.

Now the meter. This is what the same companies charge when you pay for exactly what you use, priced per million tokens — where one token is roughly four characters, or about three-quarters of an English word.

Model Input Output Cached input
Claude Haiku 4.5 $1 $5 $0.10
Claude Sonnet 5 $2 $10 $0.20
Claude Opus 5 $5 $25 $0.50
GPT-5.6 Luna $0.20 $1.20 $0.02
GPT-5.6 Terra $2 $12 $0.20
GPT-5.6 Sol $5 $30 $0.50
Per million tokens, standard processing. From platform.claude.com and openai.com, checked August 19, 2026.

Put plainly: a million words of output from Claude Sonnet 5 costs about $13. A million words is roughly thirteen novels. The monthly plan costs $20.

A household electricity meter with its display reading zero kilowatt-hours
This is the shape of the choice: a flat monthly plan, or a meter that starts at zero and counts. The difference is that with electricity you know roughly what a month looks like. With tokens, almost nobody does. Photo: RobbieIanMorrison, Wikimedia Commons — CC BY 4.0.

Is it cheaper to use the API or the subscription?

Here is the calculation, with the assumptions stated so you can disagree with them. Say a “chat” means ten back-and-forth turns: you write about 100 words each time, the model answers with about 400. That is a substantial working session — drafting something, thinking a problem through — not a one-line question.

On Claude Sonnet 5, one such chat costs about 12 cents. Which means:

Bar chart of monthly pay-per-token cost at 20, 60, 100, 164 and 200 long chats against the $20 plan

Twenty long chats a month — one most weekdays — runs about $2.44. Sixty runs $7.32. You have to reach roughly 164 long chats a month, about five every single day, before the meter costs more than the $20 plan.

At the cheap end the gap is almost comic. The same twenty chats on GPT-5.6 Luna cost 27 cents; you would need something like 1,500 chats a month to spend $20. At the expensive end it flips fast: on Claude Opus 5 the break-even falls to about 66 chats a month, roughly two a day.

These are estimates built on one usage pattern, not quotes. But the direction is not in doubt, and it is where every honest calculation lands: most people paying $20 are buying headroom they never use.

The part most comparisons get wrong

Nearly every “API vs subscription” post multiplies your message count by the output price and stops there. That undercounts badly, and here is why: the model has no memory of your conversation between turns. Every time you press send, the entire thread — every question and every answer so far — is sent again as input.

So the cost of a conversation grows with the square of its length, not in a straight line. Watch what that does:

Conversation length Input tokens Output tokens Input vs output
3 turns 3,300 1,600 2.1×
5 turns 8,833 2,667 3.3×
10 turns 34,333 5,333 6.4×
20 turns 135,333 10,667 12.7×
Same assumptions as above: 100-word questions, 400-word answers, plus a short system preamble.

In a twenty-turn thread you are billed for nearly thirteen times more input than output. It feels like you are paying for the answers. You are mostly paying to re-read the conversation.

That has a practical consequence no pricing page mentions: starting a fresh chat for a new subject is a cost decision, not just a tidiness habit. Dragging an unrelated question into hour three of an existing thread can cost ten times what asking it in a new one would. Caching softens this — both vendors bill repeated context at roughly a tenth of the input rate — but only when the same material is reused quickly, and it is not something a casual user controls.

The rear of a server rack in a data center, dense with cabling
What a token actually rents: a slice of time on hardware like this. The per-token price is the only place that cost is ever shown to you directly. Photo: Derrick Coetzee, NERSC, Wikimedia Commons — CC0.

So what is the $20 actually buying?

If the meter is that cheap, why does anyone pay the plan? Because the plan is not selling you tokens. It is selling you everything wrapped around them, and the vendors’ own feature lists say so:

  • The app itself. Web, iPhone, Android. Metered access is an interface built for developers; using it as a person means running somebody else’s client.
  • Voice. Talking to it instead of typing — the single feature that matters most to the older readers we hear from.
  • Memory and projects. It remembers you between conversations. On the meter, memory is something you would have to build.
  • Images, uploads, deep research. Priced separately, or not available at all, on the metered side.
  • A ceiling. $20 is $20. A meter has no upper bound, which is exactly how people end up with a surprising bill.

That last one deserves weight. The subscription is partly an insurance policy against your own future usage — and against a script left running by mistake. If a variable bill would make you anxious, that peace of mind is worth $20 a month, and you can stop reading here with a clear conscience.

For most people, the answer is neither

Here is the finding that surprised us most while checking these pages. OpenAI’s pricing page now describes its free tier as “unlimited text chats with GPT-5.6 Luna” — subject to abuse guardrails, in its own words — with limited uploads, images, voice and memory on top. Claude’s free tier remains a real one.

If what you do is ask questions and read answers, the free tier is no longer the crippled demo it was two years ago. The upgrade you are weighing may be buying you image generation and longer memory rather than better thinking. That is a perfectly good reason to pay. It is a poor reason to pay by accident.

Three clerks working at automatic accounting machines in an office in 1960
Metered billing is not new — this is a company’s accounting machine room in 1960. What is new is that the meter now runs inside your conversation, and nothing on screen shows you the dial. Photo: unknown author, 1960, Wikimedia Commons — Public Domain.

What to do with this

Four steps to work out whether your AI subscription is worth its price
  1. Count, do not guess. Open your chat history and count real conversations in a normal week — not messages. Multiply by four.
  2. Note how long they run. Three quick turns is a different animal from twenty. Long threads are where the money is.
  3. List what else you use. Voice, images, uploads, memory, the phone app. Every item on that list is an argument for the plan.
  4. Compare against $240 a year, not $20 a month. It is the same number, and it makes the decision easier to see.
  5. Try the free tier for one week before your next renewal. If you never hit a wall, you have your answer.
  6. If you do move to metered access, set a spending limit in the provider’s console on day one, before you make a single call.

Honest limits on the numbers above. Our per-chat costs are estimates from a stated usage pattern (100-word questions, 400-word answers, ten turns), not a bill anyone received. Real threads vary enormously. Token counts also differ between models — Anthropic notes that its newer models use a tokenizer producing roughly 30% more tokens for the same text, which changes the arithmetic even at an identical headline rate.

Consumer plans and metered access do not route to identical models either, so this compares like classes, not one product with itself. And prices move: the figures here were read off vendor pages on August 19, 2026, and are worth re-checking before you act on them. This is general information about consumer pricing, not financial advice.

If you want to go deeper

  • “Are prices jumping on September 1?” Not for Claude Sonnet 5 — Anthropic’s own pricing documentation states the scheduled increase to $3/$15 will not occur. No consumer plan change has been announced either. We keep the running list on our AI Price Tracker.
  • “Which assistant should I be paying for at all?” Price is the last question, not the first. We compared the three main ones on real writing work in ChatGPT vs Claude vs Gemini.
  • “Can I avoid the ads without paying?” Partly, and the free controls do more than most people realize — we walked through all four options yesterday.
A short, plain explanation of what a token is and how it gets billed. Cribl, March 2026 · 2.1K subscribers · about 1,200 views (checked Aug. 19, 2026).
A longer beginner walkthrough of token pricing. Eamonn Cottrell, June 2025 · 12.4K subscribers · about 4,200 views (checked Aug. 19, 2026). The rates shown on screen are from 2025 and have since fallen — the method still holds.

Frequently asked questions

Is it cheaper to use the Claude API or the subscription?

For light and moderate use, the API. At Claude Sonnet 5’s published rate of $2 per million input tokens and $10 per million output, a ten-turn working conversation costs roughly 12 cents, so twenty such chats a month comes to about $2.44 against $20 for Claude Pro. The break-even sits near 164 long chats a month. On Claude Opus 5 it drops to about 66. The subscription still buys the app, voice, memory and a fixed ceiling, none of which the API includes.

Is it cheaper to use the OpenAI API or ChatGPT?

Usually yes, and dramatically so on the cheaper models. GPT-5.6 Luna is priced at $0.20 per million input tokens and $1.20 per million output, so the same twenty long chats would run about 27 cents against $20 for Plus. On GPT-5.6 Terra ($2/$12) the break-even is near 151 chats a month; on Sol ($5/$30) it arrives much sooner.

Is ChatGPT Plus worth it in 2026?

It depends on what you use beyond text. Plus at $20 adds advanced reasoning with GPT-5.6, expanded uploads and images, projects and scheduled tasks, and no ads. If you only ask questions and read answers, OpenAI’s free tier now lists unlimited text chats with GPT-5.6 Luna, and you may be paying for headroom you never reach.

What is the difference between ChatGPT Go and Plus?

Go costs $8 and buys capacity — more messages, uploads, images, voice and longer memory. Plus costs $20 and adds the advanced reasoning models, projects and scheduled tasks. One more difference matters: OpenAI’s pricing page flags Go, not Free, with “This plan may include ads.” Plus is the cheapest tier carrying a written no-ads promise.

How much does Claude cost per month?

Claude Pro is $17 a month on the annual plan ($200 billed up front) or $20 month to month, as listed on claude.com/pricing on August 19, 2026. Claude Max starts at $100. There is also a genuine free tier. Paying per token instead, Claude Sonnet 5 is $2 per million input tokens and $10 per million output.

Can I use Claude AI for free?

Yes. Claude’s pricing page lists a free plan at $0 with usage limits. It is a real working tier rather than a timed trial, though heavy or very long sessions will reach its ceilings sooner than a paid plan would.

Are AI prices going up on September 1, 2026?

Not for Claude Sonnet 5. Its $2/$10 rate was announced as introductory pricing through August 31, 2026, but Anthropic’s pricing documentation now states that this is the standard price and that the planned increase to $3/$15 on September 1 will not occur. Neither company has announced a consumer plan price change as of August 19, 2026.

How many tokens does one conversation use?

Far more than people expect, because every turn resends the whole thread. Using 100-word questions and 400-word answers, a three-turn chat runs about 3,300 input and 1,600 output tokens; a ten-turn chat about 34,333 input and 5,333 output; a twenty-turn chat about 135,333 input and 10,667 output. As a rule of thumb, one token is about four characters, or three-quarters of an English word.

Sources

Keep reading

The Briefing runs every weekday morning: one thing in AI and technology that touches what you pay, with the arithmetic shown and the sources dated. No hype, no panic — just the math.

Similar Posts