The cost of running a top-tier AI query dropped 280-fold in roughly 18 months — from $20 per million tokens in late 2022 to seven cents by late 2024, a price collapse with no real parallel in the history of computing

Date:

The Quiet Revolution in AI Costs: What Startups Need to Know

For startups navigating the fast-evolving AI landscape, the most consequential chart might not be the one flaunting larger models, record-breaking benchmark scores, or billion-dollar training budgets. Instead, it could be the one showing a quieter but more impactful shift: the dramatic collapse in the cost of querying capable AI models.

According to Stanford’s 2025 AI Index Report, which analyzed data from Epoch AI and Artificial Analysis, the cost of running inference on a model with GPT-3.5-level performance on the MMLU benchmark has plummeted from about $20 per million tokens in November 2022 to just $0.07 per million tokens by October 2024. This represents a staggering reduction of more than 280 times in less than 18 months.

While the term “per million tokens” may sound technical, its business implications are straightforward. What used to be a significant line item in product budgets has become almost negligible, enabling AI-powered features to blend seamlessly into everyday software applications without breaking the bank.

It’s important to note that this comparison isn’t simply between a single model and its upgraded version. The AI Index deliberately measures the cost of achieving a consistent capability threshold—equivalent to GPT-3.5’s performance on MMLU—rather than the sticker price of any one vendor’s offering. The $0.07 figure comes from Gemini-1.5-Flash-8B, a smaller and more cost-efficient model capable of matching GPT-3.5’s earlier benchmark.

This distinction is crucial because it highlights the broader industry trend: smaller, cheaper models are increasingly capable of delivering powerful AI functionalities at fraction of the previous costs.

The Weird Economics of Falling Intelligence Costs

Typically, software pricing improves gradually over years or decades—cloud storage becomes cheaper, bandwidth speeds increase, and chip density follows Moore’s Law. However, AI inference costs have bucked this trend in an extraordinary way.

The AI Index reveals that depending on the task, prices for large language model inference have fallen anywhere from nine to 900 times annually. An additional analysis by Epoch AI confirmed that these steep reductions span multiple domains, including knowledge retrieval, reasoning, mathematics, and software engineering.

Several factors drive this rapid cost decline. Hardware advancements play a significant role: GPUs, specialized accelerators, and optimized data-center operations have all improved. Competitive dynamics also contribute, as multiple labs offering similar capabilities intensify price pressure. Yet, arguably the greatest influence comes from innovations in model design—techniques like distillation, routing, quantization, caching, and smarter serving systems allow models to do more with less compute.

Put simply, what was once cutting-edge GPT-3.5-level intelligence in late 2022 can now be delivered by smaller, more efficient models at a tiny fraction of the cost. This shift redefines what AI-powered products can achieve economically.

Implications for Startups

For startups, the plunge in inference costs fundamentally reshapes product possibilities. In 2023, AI features often faced strict unit economics constraints: every user query carried a calculable cost, limiting the scale and frequency of AI-powered interactions.

With inference costs falling by a factor of 280 at a fixed performance level, many previous restrictions ease. Features that formerly required rationing can now run continuously. Support bots can summarize all tickets instead of just escalated ones. Writing assistants can generate multiple drafts unobtrusively in the background. Compliance tools can scan more documents. Educational apps can provide richer, more frequent feedback. Internal knowledge systems can answer small queries that once weren’t cost-effective.

One immediate benefit is experimentation. Lower costs reduce the risk of testing novel AI features, freeing product teams from the burden of complex cost spreadsheets per interaction. User experience also improves: instead of waiting for one expensive AI response, products can invisibly chain multiple smaller AI calls—classify, retrieve, rewrite, check, personalize, summarize—creating a smoother, more intelligent feel.

This evolution transforms AI from a costly feature into an ambient, integrated layer within software. Users may not see the AI in action multiple times during their workflow, but they will sense a product that’s more responsive, organized, and contextually aware.

The Catch: Cheap Tokens Don’t Mean Cheap AI

Despite the promising cost collapse, there’s an important caveat. Cheaper inference costs don’t automatically translate into lower overall AI bills. History shows that when marginal costs decrease, usage tends to expand, often dramatically. Just as cheaper storage and cloud compute led to increased demand, AI consumption is likely to grow as prices fall.

Moreover, there remains a “frontier premium.” The AI Index notes that cutting-edge models, such as OpenAI’s o1 and Anthropic’s Claude 3.5 Sonnet, still carry significantly higher token costs than smaller, cheaper alternatives. The cheapest model capable of a task isn’t always the one suited for the most complex or sensitive jobs.

This reality forces startups to ask a more nuanced question: “Which parts of our product require frontier-level AI intelligence, and which can rely on good-enough, low-cost models?”

Companies that master this balance can build stronger margins. They’ll reserve expensive, high-accuracy models for complex reasoning, critical decisions, challenging code generation, or high-value enterprise workflows, while deploying cheaper models for routine classification, extraction, rewriting, routing, formatting, and everyday assistance.

Commoditization Arrives Fast

The steep price drop also signals a shift in competitive moats. If GPT-3.5-level AI becomes nearly free, then “having AI” ceases to be a defensible advantage. Startups can no longer rely on exclusive access to models as their main differentiator.

This shifts strategic focus to other areas: distribution channels, proprietary data, deep workflow integration, trust, compliance, and user experience. While the AI model remains important, it is less likely to be the sole source of sustainable competitive advantage.

This rapid commoditization can be challenging for founders. What once seemed like a technical edge can vanish quickly as the capability curve accelerates faster than product development cycles. Features that were a selling point in a seed round pitch deck can become table stakes before the next fundraising.

Yet, this compression also opens new opportunities. Startups can now attack workflows previously deemed uneconomical, offering AI-rich products at prices comparable to traditional software. They can automate long-tail tasks that were either too small to justify human labor or too costly for earlier AI models.

The Real Collapse Is in Permission

Perhaps the most profound change is psychological. High inference costs historically forced teams to “ask permission” before deploying AI broadly—every feature had to justify its economic impact, and background AI processes were scrutinized as potential budget items.

At just seven cents per million tokens for GPT-3.5-level performance, that mindset shifts. Instead of wondering whether AI usage is affordable, teams can start asking, “Which tasks should be intelligent by default?” Not every task requires cutting-edge AI, but many now merit some degree of model assistance.

This transition is unprecedented in computing history. It’s not just that one service became cheaper—it’s that natural-language intelligence, a fundamentally new form of software input, moved from a premium, scarce resource to a widely available commodity in under two years.

The winners in this new era won’t be those who spend the least on tokens. They’ll be the companies that deeply understand what cheap intelligence enables—and what it renders obsolete.

In late 2022, AI at GPT-3.5-level was a precious resource developers used sparingly. By late 2024, that same capability level is priced like “infrastructure dust.” This is not the end of the AI business model—it’s the start of a more demanding one: where intelligence is cheap, products must be genuinely useful to succeed.

Read the full analysis here.

LEAVE A REPLY

Please enter your comment!
Please enter your name here

Share post:

Popular

More like this
Related