LLM Cost Reduction: How We Cut Our Inference Cost per Session by 6x
Dennis Zagiansky8 min read
In the first week of September, traffic from our largest tenant jumped. Our LLM spend on that tenant was cut in half.
That combination is the whole story, and it is also why "our LLM bill went down" is the wrong way to tell it. Between the late-August peak and mid-September 2026, our fleet-wide LLM cost per processed session fell roughly 6x, while traffic went up.
We build cost attribution for teams running AI agents in production. This is what happened when we pointed it at our own analytics pipeline: what moved the number, what didn't, and the part of the drop that was not efficiency at all.
Why Absolute Spend Lies
If you only track total spend, you are measuring traffic as much as engineering.
Our traffic has a weekday-to-weekend cycle of about 5.3x. A cheaper weekend tells you nothing about your code, and a busy Tuesday can make a real optimization look like a regression. For a long time that cycle buried the signal. We could ship a genuine improvement and never see it in a spend graph.
So before optimizing anything, we rebuilt our cost dashboard around one normalized metric: cost per processed session, trailing 7 days. Dividing spend by a unit of work removes traffic from the picture, and the seven-day window absorbs the weekly cycle. When that line goes down, the pipeline is doing the same work for less, or doing less work. We will come back to that distinction.
We also attributed every LLM call to a component, a model and a tenant. You cannot optimize what your dashboard cannot show you, and a single blended number cannot tell you which feature is spending the money. It is the same reason traditional SaaS metrics fail for AI agents: an LLM call is not a generic request, and it needs its own accounting.
The Results
A note on the numbers first. Our dollar figures are in-app estimates priced from a model price table, and that table has known errors: some models carry a different provider's rates than the ones we actually pay. Both sides of every comparison carry the same bias, so ratios and percentages hold; absolute dollars do not. That is why everything below is a ratio.
- Fleet-wide: on a trailing-7-day basis, cost per session fell about 60% while session volume rose about 62% and absolute spend fell about 35%. From the late-August peak to September 10, cost per session fell roughly 6x.
- Our largest tenant: cost per session fell about 82% from its late-August peak to September 10, roughly 5.5x.
- The best single week: for that tenant, from the week ending August 30 to the week ending September 6, sessions rose 45%, spend fell 51%, and cost per session fell 66%.
No single change did this. It took three engineering mechanisms plus a deliberate scoping decision.
What Actually Drove It
In order of impact.
1. Vertex AI Flex Tier
The largest lever by far was moving analytics inference onto Vertex AI's Flex tier.
Flex is Google's discounted serving tier: 50% off list price, in exchange for higher and less predictable latency. That is a bad trade for a chat interface and a reasonable one for much of a background analytics pipeline, but "much of" is doing real work in that sentence, and it varies by component.
We rolled it out over about five weeks, starting behind a per-service flag. The share of our largest tenant's LLM calls served on Flex:
- August 2: ~3%
- August 5: ~55%
- August 23: ~77%
- September 1: ~79%
- September 4: ~99%, the day Flex became the default tier for analytics LLM calls
- September 10: ~100%
The price cut was the easy part. The engineering work was making the pipeline tolerate a slower, burstier tier:
- Sizing concurrency to what Flex actually costs. Concurrency limits tuned for the standard tier's latency are wrong for Flex. We re-sized the Flex components' concurrency to Flex's own cost and latency profile rather than carrying the old numbers over.
- Keeping context caches alive through the queue. We serve static prompt prefixes from Gemini context caches. On Flex, a call can wait in queue long enough for its cache to expire before the request is served, which silently turns a cached call into a full-price one. We fixed caches expiring mid-flight on the Flex tier.
- Making the metric honest. We applied the Flex discount inside our estimated-cost metric, so the dashboards reflected what we were actually paying for Flex calls rather than list price.
Flex is not free. It costs latency, and we are not claiming otherwise. Our dashboard tracks the latency penalty per component at p50, p95 and p99, because the only sensible way to use a discounted tier is to decide component by component whether the discount is worth the wait.
2. Context Caching for Static Prompt Prefixes
Our analyzer prompts carry a large static preamble. Served from a Gemini context cache, those tokens bill at cached rates instead of full input rates.
We built a finding that flags cacheable prompt prefixes, moved several analyzers' static prompts into context caches, and made the dashboard say whether each cache is working and what it is worth.
To keep this in proportion: our largest tenant's cached-token share went from about 13% to about 18.6%. That is a real gain of roughly five percentage points, and it is a supporting contributor, not what produced a 5.5x. If you are trying to reduce LLM inference costs, caching is worth doing, but check how much of your input is actually static before expecting it to carry the result.
3. Sending the Model Less
The third lever was simply sending less, and calling the model less often:
- Capping input text. We capped how much user text goes into session-summary generation.
- Skipping empty work. Trace and span batches with no analytics work to do no longer reach a model at all.
- Removing a pass. We removed a prompt-injection guard from our session classifier.
- Measuring what goes unused. We started measuring how much of each tool payload the model never uses, which tells us where the next round of trimming should come from.
The Part That Isn't Efficiency
Not all of the 6x came from doing the same work more cheaply. Some of it came from doing less work, on purpose.
The clearest example is session summaries. They used to be generated for every processed session. We moved them to on-demand generation behind a per-service flag. It is the same feature, with less of it running when nobody has asked for it.
That is a legitimate product decision and a real cost lever. It is not an optimization, and it would be misleading to fold it into the efficiency story.
The general lesson: in your own reporting, keep "cheaper per unit of work" separate from "less work." Both lower cost per session. Only one of them means your pipeline got better. If you blend them, you will eventually credit your engineering for a feature you quietly switched off.
An LLM Cost Reduction Checklist
What we would tell a team starting on this today:
- Normalize by a unit of work. Cost per session, per task, or per successful outcome. Absolute spend mostly measures traffic.
- Attribute cost by component, model and tenant. A blended number cannot tell you where to look.
- Evaluate discounted tiers per component. Measure p50, p95 and p99 latency before and after, and move only what can absorb the wait.
- Know your price table. In-app estimates are good for ratios and trends. They are not a bill. If you need a real dollar figure, pull it from your cloud billing.
- Separate optimization from scope change. Label which is which in the same dashboard, so nobody has to reconstruct it later.
The Bottom Line
The biggest share of our 6x came from a discounted serving tier and the engineering needed to use it well, with caching and leaner prompts behind it, and a deliberate decision to run expensive analysis only when someone asks for it. None of it was visible until we measured cost per session and attributed it by component.
That instrumentation is the product we build. Brizz shows teams what each agent component and skill costs per session and per success, and what a classification run or backfill will cost before it runs. If your inference bill is climbing and you cannot yet say which part of your agent is driving it, that is where Brizz starts.