The essentials

Tokens are the smallest running item. The example agent costs 3,302.54 euros over twelve months: 378.85 for tokens, 1,217.81 for rework, 1,127.60 to build. It breaks even at 965 cases a year.

Three items are missing from almost every quote

A quote for an AI agent lists the build and the subscription. It rarely lists token usage in steady operation, hardly ever the hours for the next model migration, and almost never the time spent on the cases the agent gets wrong.

In the calculation below, those three items are 58.6 percent of the twelve-month cost. The example agent runs to 3,302.54 euros in year one, or 22.9 cents per case. That the uncertainty persists shows in the Bitkom survey of 14 September 2026: 54 percent of AI users name cost uncertainty as a hurdle.

Every line was computed with python3 from list prices dated 28 September 2026. The volumes sit in the next section and are yours to swap out.

The example agent, so the numbers stay checkable

The agent drafts answers to service requests. Every volume here is a figure set for the example so you can replace it; nothing was measured for this article.

Twelve months, five lines, 3,302.54 euros

The twelve-month total comes to 3,302.54 euros. Each line reads: item, arithmetic, annual total, share.

  • Tokens: 3.00 US cents per case, 36.00 dollars a month, 31.57 euros; twelve months 378.85 euros, a share of 11.5 percent.
  • Automation platform: 20.00 euros a month; twelve months 240.00 euros, a share of 7.3 percent.
  • Model migration: once a year, 12 hours to retest and adjust the instructions, 338.28 euros, a share of 10.2 percent.
  • Rework: 3 percent of cases at 6 minutes each, so 3.6 hours and 101.48 euros a month; twelve months 1,217.81 euros, a share of 36.9 percent.
  • Build, one-off: 40 hours, 1,127.60 euros, a share of 34.1 percent.
  • Strip out the build and 2,174.94 euros of running cost remain, which is 15.1 cents per case.

Tokens are 11.5 percent, not 90

The token price is the item a vendor is happy to itemise. Companies read it correctly: in the Bitkom survey of 14 September 2026, 8 percent named token usage as a major cost block, against 51 percent for infrastructure and 50 percent for data preparation.

Switching models moves little. The same volumes cost 727.39 euros a year in tokens on Claude Opus 5.5 instead of 378.85 on Sonnet 5, and only 189.42 euros on Claude Haiku 4.5. The spread between cheapest and dearest model is 537.97 euros a year, less than six months of rework.

Caching matters more than the model choice. Without a cache all 18,000 input tokens per case would be fresh, and the token line climbs to 644.04 euros a year, a factor of 1.70. The per-case breakdown is in working out AI cost per case.

The migration date belongs to the vendor, not to you

Your model will be retired, on somebody else's schedule. Anthropic promises at least 60 days of notice before a model retires and retired three in 2026: Claude Sonnet 3.7 on 19 February, Claude Sonnet 4 and Opus 4 on 15 June, Opus 4.1 on 5 August. Calls to a retired model fail.

A migration is not one line of configuration. Setting temperature, top_p or top_k returns a 400 error on Claude Opus 4.7 and later. Models from Claude 4.7 onward also use a newer tokenizer that, according to Anthropic, produces roughly 30 percent more tokens for the same text, which lifts the token bill with no price change at all.

Hence the 12 hours a year in the table: update the instructions, replay twenty old cases against the new model, review the differences, sign off. Skip that budget and you pay it during an outage, at outage rates.

Rework is the biggest running item

Rework is 36.9 percent of the annual cost. At a 3 percent error rate and 6 minutes per case that is 36 cases and 3.6 hours a month, or 1,217.81 euros a year at 28.19 euros an hour. The rate is set, not measured, and it is the most sensitive figure in the whole calculation.

Work the range yourself. At 5 percent errors and 10 minutes of rework the item rises to 3,382.80 euros a year, the annual total to 5,467.53 euros, and the price per case from 22.9 to 38.0 cents. An agent gets expensive through its hit rate, not through tokens.

That yields one condition for the contract: the agent has to flag the cases it is unsure about. An agent without that list costs you a review of all 1,200 cases rather than 36.

The formula for your own quote

One line, five numbers, no spreadsheet required. Annual cost equals 12 times (cases per month times token price per case plus subscription) plus migration hours times hourly rate plus 12 times cases times error rate times rework minutes divided by 60 times hourly rate plus build hours times hourly rate.

Four questions get the missing numbers into the quote. Send them word for word.

  • How many input and output tokens does one case consume on average, measured on ten real cases from our own traffic?
  • Which model, with which retirement date, and who pays the hours on migration day?
  • What error rate do you commit to, and how do we see the uncertain cases without reading all of them?
  • What does leaving cost: do the instructions, rules and logs stay with us, and in what format?

From 965 cases a year the agent pays for itself

The comparison is manual handling. The same 14,400 cases at four minutes each are 960 hours and, at the gross hourly wage of 28.19 euros, 27,062.40 euros, a factor of 8.2 against the agent. Per case that is 1.88 euros by hand against 11.1 cents of variable agent cost.

So the tipping point is fixed: 965 cases a year including the build, and 327 cases from year two onward. Yes to the agent when the workflow runs above 1,000 similar cases a year and somebody reviews the flagged ones.

No to the agent in three situations: below roughly 1,000 cases a year, when every case looks different, and whenever a fixed rule would do. That sorting exercise is in AI automation in mid-sized companies.

Common questions

Three questions about pricing models that come up often.

Next step: four lines into the next quote

Take the quote on your desk and add four lines: tokens per case, hours per model migration, the committed error rate, and rework hours per month. Without those four numbers, the price at the bottom covers the build alone.

Then run the formula above and compare it with your manual figure. If the saving lands below 20 percent, negotiate the error rate rather than the price.

If you would rather walk the numbers through with someone, we do that in a free first call, and we build the workflow as an internal AI service.

Sources and status

Sources last checked: 28 September 2026. Vendor statements and our own reading of them are kept apart in the text.

  1. Anthropic: Preise der Claude-Modelle
  2. Anthropic: Abschaltung von Modellen
  3. OpenAI: Preise der API-Modelle
  4. n8n: Preise
  5. Microsoft: Preise für Power Automate
  6. Statistisches Bundesamt: Durchschnittliche Bruttoverdienste, Stichmonat April 2025
  7. Europäische Zentralbank: Euro-Referenzkurs US-Dollar
  8. Bitkom: Erstmals nutzt die Mehrheit der Unternehmen KI

Corrections: [email protected].

What does this mean for your business?

We go through one concrete workflow with you and check whether AI can help.

Book a free first call →