Decision Models vs Agentic RAG: Where Jev Cuts Token Costs

In agentic RAG, an LLM repeatedly reads retrieved passages to judge relevance, decide whether to search again, and check whether sources support an answer. Those are narrow judgments. A decision model such as TypeSafe's Jev can make them for far less, so the LLM reads only what passes. Savings depend on how much of the pipeline is judgment.

Where agentic RAG spends tokens

Basic RAG retrieves a few passages and asks the model to answer. Agentic RAG adds loops: the model inspects results, decides whether they are relevant, reformulates queries, searches again, checks whether sources support its claims, and only then writes. Each inspection means reading passages into context. For document-heavy domains such as finance, legal, or compliance, the reading can dwarf the writing.

Most of those inspections are narrow questions with constrained answers. Is this passage relevant to the question? Which of these documents is the most recent filing? Does this source support this sentence? An LLM can answer them, but it reads every passage in full to do so, at LLM prices.

What a decision model does instead

A decision model reads state and returns typed answers to predefined questions rather than generating text. TypeSafe's Jev, launched in September 2026, is one example: per its documentation it supports yes or no questions with a probability, choices among options with a probability for each, and rubric scores. TypeSafe's published figures, reported by Flavio Copes and Firecrawl, put latency at roughly 70 to 500 milliseconds and launch list pricing at 0.042 dollars per million input tokens with no charge for output.

In an agentic RAG pipeline, that model can score retrieved passages for relevance, filter out off-topic results, and check citation support, so the LLM reads only the passages that survive.

Where the savings are real

  • Reranking large result sets. Scoring dozens or hundreds of retrieved chunks and passing the top few to the LLM.
  • Filtering before reading. Discarding irrelevant, outdated, or duplicate documents early.
  • Citation checks. Asking whether each source supports each claim instead of having the LLM reread sources.
  • Routing. Deciding whether a question needs retrieval at all, or which index to search.

The common feature is volume: many small decisions that would otherwise each cost an LLM read.

Where they are not

Savings shrink when retrieval already returns a few precise passages, when the work is mostly synthesis and writing, or when decisions require arithmetic, date logic, or reading exact figures, which TypeSafe lists among Jev's weak spots. In finance especially, comparing numbers across filings belongs in code, not in any probabilistic model. Claims of very large token reductions come from specific vendor workloads; Firecrawl notes TypeSafe's own caveat that its comparison workflows were built by its staff. Benchmark on your own documents and questions.

How to measure the effect

Pick a fixed set of real questions with known good answers. Run the pipeline as it is and record LLM input tokens per question, answer accuracy, latency, and total cost. Add the decision step, for example scoring retrieved passages and passing only those above a threshold, and run the same questions again. Compare all four numbers, and inspect the questions where accuracy dropped: they show where the threshold is too aggressive or the judgment question is poorly framed.

Framing good judgment questions

Decision models respond best to specific, constrained questions about the state you provide. Ask whether a passage answers a stated question, not whether it is useful. Offer explicit options when classifying, such as current filing, prior filing, or unrelated document. Keep the state focused on the passage and the question, rather than the whole conversation. Well-framed questions are also easier to test and calibrate.

The layer underneath still matters most

A judge can only choose among what retrieval returns. Poor chunking, a weak embedding model, missing keyword search, or a stale index produce candidates no judge can fix. The most reliable cost reductions usually come in order: retrieve better, so fewer and more relevant passages arrive; reuse stored answers to repeated questions instead of regenerating them; then add a decision model to make the remaining judgments cheaply.

Frequently asked questions

What is agentic RAG?
Agentic RAG is retrieval augmented generation where the model works in a loop: it searches, inspects results, decides whether they are relevant or sufficient, refines its queries, searches again, and checks sources before answering. It handles complex questions better than single-pass RAG, at the cost of many more model calls and tokens.
Can Jev replace an LLM in a RAG pipeline?
Only for judgment steps. Jev returns typed answers, such as relevance scores or yes or no checks, and does not write text. It can handle reranking, filtering, and verification so the LLM reads less, but the LLM still synthesises and writes the answer. Arithmetic and exact-figure comparisons belong in code.
How much can a decision step actually save?
It depends entirely on the workload. The large multiples in circulation, such as the 100x in press coverage, describe cost or speed per decision against frontier models, not token reduction in a retrieval pipeline. Published pipeline figures are far smaller, in the range of retrieving a third to three quarters less material. Pipelines dominated by judging large volumes of retrieved text can save a lot; pipelines dominated by writing save little. Measure your own: track LLM input tokens per answered question before and after adding a decision step.
Which steps in a RAG pipeline should use a decision model?
Steps that ask the same narrow question many times: is this passage relevant, is this document current, does this source support this claim, does this question need retrieval at all. Leave synthesis, explanation, and anything involving calculation to the LLM or to code, and keep thresholds tuned against measured accuracy.