Case study · AI operations · Aug to Sep 2026
LLM Cost Control
At the end of August the shared LLM key for our team hit its monthly budget and every flow using it started failing. The platform team pointed at one of my agent workflows. I found out why it was expensive, fixed it, and then got prompt caching exposed in the company's n8n model node so the rest of our agents could use it.
What was happening
The workflow reads corporate trip requests and signed work orders with two Claude agents, and each agent has a few tools to read and update rows in the operations sheet (described in the B2B pipeline). I pulled 59 real executions from the previous four days and costed each step.
One lookup tool, meant to find the row for a given trip, had no filter set. Every call returned the entire sheet: 381 rows, about 207,000 tokens. The agent called it several times per run, and because an agent loop resends the whole conversation on every iteration, that payload was paid for again and again. That single tool was 96% of the spend.
I also ruled out the obvious suspects before touching anything: files processed twice (six distinct versions, checked by hash), retries on the tool nodes (none configured) and column naming (not the cause).
The fix
- The lookup column is now set by expression and its value comes from the model's own argument, with "return only the first matching row". The tool now returns one row instead of 381.
- The agents' iteration cap went from 40 to 12, which is still well above what a normal run needs.
- I re-ran the same document before and after: 381 rows returned became 1, five tool calls became three, and input tokens went from 1,443,324 to 9,623. Cost per run: $4.37 before, $0.069 after.
Show the daily values as a table
| Day | Index |
|---|---|
| Mon 10 Aug | 55 |
| Tue 11 Aug | 100 |
| Wed 12 Aug | 86 |
| Thu 13 Aug | 44 |
| Fri 14 Aug | 106 |
| Sat 15 Aug | 28 |
| Sun 16 Aug | 13 |
| Mon 17 Aug | 95 |
| Tue 18 Aug | 104 |
| Wed 19 Aug | 106 |
| Thu 20 Aug | 78 |
| Fri 21 Aug | 190 |
| Sat 22 Aug | 16 |
| Sun 23 Aug | 11 |
| Mon 24 Aug | 113 |
| Tue 25 Aug | 88 |
| Wed 26 Aug | 96 |
| Thu 27 Aug | 94 |
| Fri 28 Aug | 197 |
| Sat 29 Aug | 0 |
| Sun 30 Aug | 0 |
| Mon 31 Aug | 0 |
| Tue 01 Sep | 25 |
| Wed 02 Sep | 31 |
| Thu 03 Sep | 35 |
| Fri 04 Sep | 45 |
| Sat 05 Sep | 17 |
| Sun 06 Sep | 16 |
| Mon 07 Sep | 33 |
| Tue 08 Sep | 31 |
| Wed 09 Sep | 30 |
| Thu 10 Sep | 27 |
| Fri 11 Sep | 13 |
| Sat 12 Sep | 9 |
| Sun 13 Sep | 7 |
| Mon 14 Sep | 15 |
| Tue 15 Sep | 39 |
| Wed 16 Sep | 51 |
| Thu 17 Sep | 40 |
| Fri 18 Sep | 38 |
| Sat 19 Sep | 6 |
| Sun 20 Sep | 6 |
| Mon 21 Sep | 22 |
Prompt caching
After the fix, CabiBot, the Help Center agent, became the largest spender: a long, stable system prompt resent on every message. That is exactly what prompt caching is for, but the n8n node we use to call the company's LiteLLM proxy only exposed the model, temperature and max tokens.
I confirmed with the proxy owner that it supports cache injection points, opened an issue with the team that maintains our n8n, and they shipped the change. It went live on our agents on 8 September.
What I took from it
- Cost per run is a metric that has to be looked at per workflow, weekly. A monthly total hides a single tool that is 96% of the bill.
- Tool output has to be bounded by design. The model has no way to tell that a lookup returned too much, so the limit has to live in the tool.
- The trip workflow's spend has crept back up in September as volume grew, to roughly a third of its August level. That is the next thing I am putting an alert on.
Figures are spend on my team's key only. They describe my own automations, not company revenue or costs.