Twenty dollars. Ask anyone how much AI costs per month and that's the number you'll hear, because it's the list price of ChatGPT Plus, Claude Pro, Google's AI Pro plan and Cursor Pro. It correctly answers a narrow question: what one person pays to chat with a frontier model. For a company the range is much wider, from $20 per seat to several thousand dollars a month. Three things decide where your bill lands: how many of your people use AI heavily, whether AI runs inside your product for customers, and who keeps it working after it ships.
This is the answer I give clients, using list prices at the time of writing. A standard seat for a knowledge worker costs $20–30 a month, and business tiers with admin controls and data-retention terms cost $25–30. Power users hit the usage caps and move up to the $100–200 tiers. These are engineers running coding agents all day, or analysts working in long-context research. An AI feature inside your product usually costs $50 to $1,000 a month in inference for a small or mid-sized SaaS. Nearly all of that spread comes from which model tier you use and how much context you send with each request. Then there is the engineering to keep the feature correct: evaluations, monitoring, and rewriting prompts when a model is retired. In our experience that takes 2–5 engineer-days a month, which is $800–4,000 at nearshore rates of $50–99 an hour. That last line is usually the biggest one, and it's the one most budgets leave out.
The usual answer to what AI costs rests on five claims. Each one is true within limits. Here is where each stops being true.
Claim One: AI Costs $20 a Month
The $20 subscription exists, but it's rationed. Every consumer tier caps how many messages or how much compute you get in a rolling window. Someone doing sustained work, like feeding long documents to a model or running an agent across a codebase, can use up that allowance before lunch. That's why every major vendor now sells a $100 or $200 tier: Claude Max, ChatGPT Pro, Cursor's top plan. Those prices reflect what sustained use actually costs, and the vendors priced them on real usage data.
At company level the seats add up. Take a 20-person company with eight engineers. Business-tier chat seats at $30 come to $600 a month. GitHub Copilot Enterprise at $39 for each engineer adds $312. Microsoft 365 Copilot is $30 per user on top of the base Microsoft 365 license, so anyone who also wants AI inside Office adds another line. That's about $900 a month, or roughly $45 per head, before a single engineer moves to a $200 tier. Most of the leak in seat spend comes from overlap: three tools doing the same job for the same person, each renewed separately.
Claim Two: The API Costs Fractions of a Cent
Per call, that's accurate. Mid-tier frontier models cost around $3 per million input tokens and $15 per million output tokens, and small models cost $0.10–0.60. A single question and answer costs a fraction of a cent. The trouble is volume.
Here's a realistic support assistant built on retrieval-augmented generation. It handles 10,000 conversations a month at six turns each. Every turn sends about 4,000 tokens of retrieved documentation and conversation history and gets back about 300 tokens. That's 240 million input tokens and 18 million output tokens a month. On a mid-tier frontier model, the bill is $720 for input plus $270 for output, about $990 a month. On a small model the same traffic costs about $47. It's the same feature at a twentyfold price difference, and whether you need the expensive model is a question you can answer with testing.
Look at where the tokens come from in that example. Most are input, and most of that input is context the user never sees: retrieved chunks, system prompts, and conversation history that gets sent again on every turn. A pipeline that retrieves ten chunks when three would answer the question pays more than three times as much for input. The biggest cost levers are architecture choices. Prompt caching, which providers discount heavily for repeated prefixes, is one. Trimming history is another. So is routing easy requests to a small model and sending only the hard ones up to a frontier model. Getting these right can move a monthly bill by an order of magnitude without changing what the user sees.
Claim Three: Inference Is the Main Cost
In most production AI features I've seen, inference is the smallest recurring cost. The bigger one is keeping the output correct over time. Providers deprecate model versions on their own schedule. The replacement answers differently, and your prompts were tuned for the old model. If you don't have an evaluation set, meaning a few hundred real inputs with known good outputs that you run against every change, you'll learn about the regression from customers.
The upkeep has several parts. The retrieval index has to be rebuilt when the source documents change; the embeddings cost pennies per million tokens, but the pipeline work doesn't. Someone has to review tracing in a tool like Langfuse or LangSmith. Evals have to run on every model or prompt change. Guardrails need adjusting when users find inputs nobody expected. Our Python and AI/ML practice has run production AI pipelines for about two years, and the pattern hasn't changed: quality drifts in the months when a pipeline has no clear owner.
That changes the number to compare against value. A feature with $100 of monthly inference and 2–5 engineer-days of upkeep costs $900–4,100 a month. It's still often worth it. But the business case has to be made against that full figure, not the API invoice.
Claim Four: Prices Keep Falling, So the Bill Will Shrink
Per-token prices really have fallen. GPT-4 launched in March 2023 at $30 per million input tokens and $60 per million output tokens. The small models released in 2024 charge around $0.15 per million input tokens, about 200 times less, for the many tasks they handle well.
Total spend keeps rising anyway, because usage grows faster than prices fall. Reasoning models bill their internal thinking as output tokens, so a response with 300 visible tokens can be billed for thousands. Agents loop: a coding agent that reads a repository, runs tests, reads the failures and tries again can use more tokens in an afternoon than a chat user does in a month. The cost of each unit keeps dropping while the number of units climbs. The planning assumption I'd defend is that the cost per task falls year on year while spend per power user rises. The $200 tiers exist because that pattern already shows up in the data.
Claim Five: AI Pays for Itself
The math on a single seat is simple. A $200-a-month power tier, set against an engineer billed at $50–99 an hour, pays for itself if it saves two to four hours a month. For engineers who use Claude Code or Cursor every day, that bar is low.
That payback depends on one condition: someone reviews the output. At EltexSoft, every pull request is reviewed by at least one other senior engineer before merge, whether a person or an agent wrote the first draft. That review is what makes AI tools an accelerator we can defend. When generated code isn't reviewed, the cost doesn't go away. It moves from the tool budget to rework, where it's bigger and shows up months later. Product features work the same way: an AI answer nobody checks is a support ticket you haven't received yet.
When AI Is Worth the Monthly Bill
AI pays off when three things are true: the task repeats at volume, a wrong answer is cheap to spot, and you've measured the baseline so you can prove the improvement. Support deflection, document extraction, internal knowledge search, and code drafting with review in place all fit. In those cases the monthly cost grows with the value delivered, and the upkeep is spread across thousands of transactions.
It doesn't pay off when the volume is low, the stakes are high, and nobody owns the evals. It also doesn't pay off when the feature is on the roadmap mainly because the roadmap was expected to include AI. In those cases the $900–4,100 a month buys risk. The fix is scope discipline, which applies to AI features like anything else. When we rescued Meal4U's stalled build, we re-estimated the scope about ten times, cutting 30–50% on each pass until we found the smallest product worth shipping. AI features respond to the same approach. The smallest useful version is usually a small model on a narrow task with an eval set from day one, and that version tells you whether the frontier model on everything was ever needed.
The Number That Belongs in the Budget
Put together, the monthly cost of AI has four lines. Seats come to $20–45 per head blended across chat and coding tools. Power users cost $100–200 each. Inference for a production feature runs $50–1,000 at small-to-mid SaaS volume, driven by model tier and context size. Upkeep for that feature takes 2–5 engineer-days a month. A quote for a production AI feature that covers only the inference line is covering the smallest of the four.
That's why we cost an AI feature before we build it. During EltexSoft's free discovery week, we model inference against your real traffic, covering tokens per request, model tier, caching and routing, and we estimate the upkeep alongside it. If the numbers hold up, a paid pilot with no lock-in puts the smallest useful version into production, and we continue from there in 2-week sprints. The next step is to send us the feature you're considering and a month of usage data. We'll give you back the monthly figure with every driver itemized.
Last updated October 2, 2026