Your AI feature costs you money every time someone uses it. The number is falling fast. You should be delighted. Most founders I talk to are, right up to the moment I ask them what happens to their pricing when their competitor's costs fall too.
Cheap inference is not a gift to you. It is a gift to everyone, including the person who wants your customers.

The numbers are absurd, and they are real
Start with the price of thinking. a16z's LLMflation analysis found inference cost for a model of equivalent performance drops by roughly 10x every year. GPT-3 ran at $60 per million tokens in November 2021. Llama 3.2 3B hit $0.06 per million. A thousandfold drop in three years.
More recent work sharpens the picture. Data Today, citing Epoch AI, reports the price to run a model at fixed performance has fallen about 40 times per year, with the full range spanning "between 9 and 900 times per year depending on which capability milestone you measure."
Look at what you rent today. Together.ai's published rates list GPT-OSS 20B at $0.05 in and $0.20 out per million tokens. Llama 4 Maverick sits at $0.15 and $0.60. Qwen 3 235B runs $0.20 and $0.60. These are open-weight models. Anyone with a card rents them. Anyone with a GPU runs them.
Now ask the uncomfortable question. If the model behind your product costs a fifth of what it did last year, and your price has not moved, are you a business or a temporary arbitrage?
The margin gap nobody wants to talk about
Here is where it gets sharp. Cheap tokens do not automatically mean fat margins.
Analysis from Preuve.ai on AI wrapper economics puts wrapper gross margins at 50-60%, against 70-90% for classic SaaS. The same piece reports wrappers need roughly 3.2x more funding to reach profitability, and operators who survive keep API cost at 30-50% of revenue through model routing.
Sit with those numbers. Falling token prices have been running for years, and margins are still 20 to 30 points below normal software. The savings did not land in your pocket. They landed in your price, because your competitor cut theirs first.
Retention tells the same story from a different angle. Growth Unhinged's churn data, drawn from roughly 200 AI-native companies alongside 2,700 B2B SaaS firms, found median gross revenue retention for AI-native companies at 40% and net revenue retention at 48%. B2B SaaS median NRR was 82%. Split by price point, products under $50 a month held 23% GRR. Products over $250 a month held 70%.
Cheap products churn. Cheap products are what commodity inference produces. You see the loop.

Where prices are not falling
This is the part most people skip, and it is the most useful thing on this page.
The decline is not uniform. Per the Epoch AI figures, reaching GPT-3.5 quality on general knowledge got up to 900 times cheaper per year. Holding PhD-level science accuracy steady got cheaper at roughly 9 times a year. Mid-range work sits near 40x.
Read those numbers as a business signal, not a benchmark. Easy work is collapsing toward free. Hard work is still expensive, because staying at the top of the capability curve still needs a large, current model.
So the strategic question is not "how do I get cheaper tokens." It is: is the work my product does easy work or hard work?
If your product summarises documents, drafts emails, classifies tickets, or rewrites copy, you are on the 900x slope. Your pricing power is evaporating on a schedule. If your product does multi-step reasoning over messy proprietary context with real consequences when it is wrong, you are on the 9x slope, and you have years, not months.
Most founders think they are on the 9x slope. Most are on the 900x one.
The build versus buy call for 2027
I have made this decision badly before, and the wrong version always sounds smart in the room. Let me give you the version I use now.
Question one: would you survive being copied by your model provider?
The acid test from Preuve.ai is blunt: "If a foundation model shipped your pitch as a default next release, would customers cancel? If yes, you are building a feature, not a company." They cite an AI writing tool whose $50K MRR fell 70% in 60 days after a single model update.
Not a hypothetical risk. A Tuesday.
Question two: does your cost structure survive a competitor at half your price?
Run the sum with your actual token spend. If a rival routes to an open-weight model at $0.15 per million input tokens and undercuts you by half, does anything remain? If the answer is "we would cut price," you never had pricing power. You had a head start.
Question three: what compounds?
Falling token prices are a rising tide under everyone. Only one category resists commoditising: things which accumulate. Proprietary data collected by being inside the workflow. Corrections and outcomes a competitor cannot scrape. Workflow depth, where you become the system of record instead of a single API call. Preuve.ai cites a document-automation company whose churn dropped 80% once customers had 1,000-plus processed documents inside the product.
Nobody switches away from their own history. They switch away from a text box.

What to do about it on Monday
Four things. None of them require a strategy offsite.
Instrument your token cost per customer. Not aggregate spend. Per customer, per feature. You cannot reason about a cliff you have not measured. Most teams I meet know their monthly API bill and nothing else.
Build the routing layer now. Every request should be able to move between models without a code change. Not because you plan to switch tomorrow, but because the day you need to, you will need to do it in an afternoon. Model routing is what keeps API cost inside 30-50% of revenue.
Pick your slope deliberately. Decide whether you are competing on hard reasoning or on convenience. Both are viable businesses. Confusing the two is not. Convenience businesses need distribution and low cost. Hard-reasoning businesses need depth and proof.
Start hoarding what compounds. If your product does not currently capture a signal a competitor cannot buy, there is your roadmap. Everything else is a feature.
The uncomfortable summary
Falling inference cost is not a tailwind. It is a tide. It lifts your competitor's boat by exactly the same amount it lifts yours, and it does it every year without asking your permission.
Businesses which survive the pricing cliff are not the ones with the cheapest tokens. They are the ones where tokens stopped being the product.
So look at your product honestly. If your provider halved their prices tomorrow, would you win or would your customers simply expect a discount?
Answer it before your board does.