Skip to content

We stopped paying an LLM to answer yes/no questions: 21x cheaper with Jev

Aviv TabakAviv Tabak5 min read
We stopped paying an LLM to answer yes/no questions: 21x cheaper with Jev

On a live backfill this week, our session classifier ran at about $4 an hour. The same workload used to cost about $89 an hour. That's 21x cheaper, and the average request went from 49 seconds to 3.

We didn't get there with a cheaper LLM or a tighter prompt. We stopped using a generative model for that job and moved it to Jev, TypeSafe's "System One" model.

Where the money was going

Brizz classifies every session an AI agent has. Did the user get what they came for? Did the agent hand off? Did a capability break? Every one of those questions has a fixed set of answers.

For a long time, a generative LLM was the obvious choice for this. Agent conversations are long and messy, and answering these questions means reading the whole exchange in context. An LLM read each session and gave us the label.

What changed is that there's now a real alternative for this kind of question. A generative model is built to write, so even when the answer is one option from a fixed list, you're paying for a model that produces text and a step that turns that text back into a label. At thousands of sessions a day, that came to about $89 an hour on a busy run.

What Jev does differently

Jev doesn't generate text. You give it a question with a closed set of answers and it returns the answer directly: a calibrated probability for a yes/no question, or a pick among competing options. There's no prose to pay for and nothing to parse.

For a classifier, that's the whole job. Our label sets didn't change, the questions didn't change, and the output our customers see is the same. Only the engine behind the answer changed.

Here's one run of 17,840 classification requests, compared with the same work on the generative model we used before (Gemini, for the curious):

  • Cost: $1.40 on Jev versus $29.67 on the generative model, 21x cheaper. At full backfill speed that's about $4 an hour instead of $89.
  • Mean latency: 3.0 seconds per request, down from 49.0 seconds, 16x faster
  • Errors: 0 across all 17,840 requests

The next bill: screening before we pay

Session classification was the first job we moved. The second is how we catch broken tools and integrations in agent conversations. That change is in review now, and the savings there are bigger in dollar terms.

That check looks at each window of conversation where the agent called a tool, and describes anything that broke. Today it makes a generative call on every one of those windows to find out whether there's anything to describe. For one large customer that's 1.19 million calls a month, and only about one window in 470 turns up a finding. We were paying full price to hear "nothing here" over and over.

The new version asks Jev first: does this window contain anything worth describing? Only the windows that pass go to the generative description call, which is unchanged.

  • Jev screen: $0.000102 per window
  • Generative description: $0.009742 per window, 95x more

On that customer's real traffic, the 1.19 million described windows drop to about 270,000 a month, and spend on this step drops from $11,594 to $2,759, or 76%.

Cheaper only matters if the findings survive, so we ran it against our golden set with the real LLM and the real engine. Both versions found 15 of 19 known issues, and both had zero false positives. No issue was lost because the screen blocked it.

We also sized the savings on real traffic rather than the golden set. The golden set is packed with interesting cases, so it passes far more windows through the screen than production does, and using it would have made the numbers look better than they are.

What we kept from the old setup

Two decisions kept this safe to ship.

The LLM stays as the fallback. If a Jev call fails, the session gets classified by the generative model as before. During the rollout we hit a burst of connection failures under heavy load (we traced them to a shared egress IP, and the vendor is looking into it). Nothing was lost. Those windows just cost us twice until we tightened our retries.

The screen fails open. If the engine is down or returns something unexpected, the window gets described as usual. The screen protects a bill, not a finding, so when it can't answer, the right move is to spend the money and keep the finding.

We also didn't flip everyone at once. The engine can be turned on per service, per customer or per plan.

How to find this in your own stack

If you run agents or LLM pipelines in production, there's a good chance some of your spend looks like ours did. A few ways to spot it:

  • Look for calls whose output you immediately parse into a fixed label. Routing, intent detection, yes/no checks, "is this worth looking at" gates. Those are closed questions, and a model that answers them directly can be much cheaper.
  • Look for expensive calls that mostly come back empty. A cheap screen in front of them can cut most of the volume without touching the calls that matter.
  • Keep the LLM as the fallback. It makes the switch low-risk and lets you move one workload at a time.
  • Measure on real traffic. Your test set is probably richer than production, and savings measured on it will look better than they really are.

We're working through the rest of our pipeline the same way, one closed question at a time.