All projects

Case study · AI operations · Aug to Sep 2026

LLM Cost Control

At the end of August the shared LLM key for our team hit its monthly budget and every flow using it started failing. The platform team pointed at one of my agent workflows. I found out why it was expensive, fixed it, and then got prompt caching exposed in the company's n8n model node so the rest of our agents could use it.

63× cheaper per runAI operations
n8nn8nClaudeClaudeAWS BedrockAWS BedrockGrafanaGrafanaPrometheusPrometheus
My roleDiagnosis, fix, verification and the caching request. I own the affected workflows.
When28 Aug to 22 Sep 2026
Built withn8n, the company LiteLLM proxy, Claude Sonnet on Bedrock, Grafana and Prometheus
StatusFixed. Spend is reviewed weekly per workflow.
63×
cheaper per run on the same input document: $4.37 to $0.069
−74%
daily spend on the team key, 16 to 28 Aug against 1 to 22 Sep
−40%
cost per request after prompt caching, while requests grew about 40%
96%
of the spend traced to a single unfiltered tool

What was happening

The workflow reads corporate trip requests and signed work orders with two Claude agents, and each agent has a few tools to read and update rows in the operations sheet (described in the B2B pipeline). I pulled 59 real executions from the previous four days and costed each step.

One lookup tool, meant to find the row for a given trip, had no filter set. Every call returned the entire sheet: 381 rows, about 207,000 tokens. The agent called it several times per run, and because an agent loop resends the whole conversation on every iteration, that payload was paid for again and again. That single tool was 96% of the spend.

I also ruled out the obvious suspects before touching anything: files processed twice (six distinct versions, checked by hash), retries on the tool nodes (none configured) and column naming (not the cause).

The fix

  • The lookup column is now set by expression and its value comes from the model's own argument, with "return only the first matching row". The tool now returns one row instead of 381.
  • The agents' iteration cap went from 40 to 12, which is still well above what a normal run needs.
  • I re-ran the same document before and after: 381 rows returned became 1, five tool calls became three, and input tokens went from 1,443,324 to 9,623. Cost per run: $4.37 before, $0.069 after.
key blocked,fix shipped05010015020010 Aug · index 5510 Aug11 Aug · index 10012 Aug · index 8613 Aug · index 4414 Aug · index 10615 Aug · index 2816 Aug · index 1317 Aug · index 9517 Aug18 Aug · index 10419 Aug · index 10620 Aug · index 7821 Aug · index 19022 Aug · index 1623 Aug · index 1124 Aug · index 11324 Aug25 Aug · index 8826 Aug · index 9627 Aug · index 9428 Aug · index 19731 Aug01 Sep · index 2502 Sep · index 3103 Sep · index 3504 Sep · index 4505 Sep · index 1706 Sep · index 1607 Sep · index 3307 Sep08 Sep · index 3109 Sep · index 3010 Sep · index 2711 Sep · index 1312 Sep · index 913 Sep · index 714 Sep · index 1514 Sep15 Sep · index 3916 Sep · index 5117 Sep · index 4018 Sep · index 3819 Sep · index 620 Sep · index 621 Sep · index 2221 Sepaverage 16–28 Aug = 100average 1–22 Sep = 26prompt caching on (8 Sep)budget exhausted 28 Aug
Billed spend per day on the team key (UTC days), indexed so that the 16 to 28 August average is 100. The key was blocked for three days after hitting its budget, which is when the fix shipped. Source: the LiteLLM spend metrics behind the team's Grafana board.
Show the daily values as a table
DayIndex
Mon 10 Aug55
Tue 11 Aug100
Wed 12 Aug86
Thu 13 Aug44
Fri 14 Aug106
Sat 15 Aug28
Sun 16 Aug13
Mon 17 Aug95
Tue 18 Aug104
Wed 19 Aug106
Thu 20 Aug78
Fri 21 Aug190
Sat 22 Aug16
Sun 23 Aug11
Mon 24 Aug113
Tue 25 Aug88
Wed 26 Aug96
Thu 27 Aug94
Fri 28 Aug197
Sat 29 Aug0
Sun 30 Aug0
Mon 31 Aug0
Tue 01 Sep25
Wed 02 Sep31
Thu 03 Sep35
Fri 04 Sep45
Sat 05 Sep17
Sun 06 Sep16
Mon 07 Sep33
Tue 08 Sep31
Wed 09 Sep30
Thu 10 Sep27
Fri 11 Sep13
Sat 12 Sep9
Sun 13 Sep7
Mon 14 Sep15
Tue 15 Sep39
Wed 16 Sep51
Thu 17 Sep40
Fri 18 Sep38
Sat 19 Sep6
Sun 20 Sep6
Mon 21 Sep22

Prompt caching

After the fix, CabiBot, the Help Center agent, became the largest spender: a long, stable system prompt resent on every message. That is exactly what prompt caching is for, but the n8n node we use to call the company's LiteLLM proxy only exposed the model, temperature and max tokens.

I confirmed with the proxy owner that it supports cache injection points, opened an issue with the team that maintains our n8n, and they shipped the change. It went live on our agents on 8 September.

36%
of input tokens now served from cache (44% excluding one anomalous hour)
$0.034 → $0.020
average cost per request, before caching (1 to 8 Sep) and after (8 to 22 Sep)
34.9k → 8.5k
input tokens per request on the key, before and after the August fix

What I took from it

  • Cost per run is a metric that has to be looked at per workflow, weekly. A monthly total hides a single tool that is 96% of the bill.
  • Tool output has to be bounded by design. The model has no way to tell that a lookup returned too much, so the limit has to live in the tool.
  • The trip workflow's spend has crept back up in September as volume grew, to roughly a third of its August level. That is the next thing I am putting an alert on.

Figures are spend on my team's key only. They describe my own automations, not company revenue or costs.