Deep Dive · Architecture

Decision models vs rerankers

They look like different categories — one is a retrieval component, the other is the model that decides things for your agent. They are not. Both read an input, score it, and write nothing. The gap between them is narrower than the labels suggest, and the one place they genuinely diverge is a number that most pipelines throw away.

By Hussain Nazary Published 6 October 2026 Models named 23 Reading time ~30 min Level Expert
0
output tokens, either model
+12 vs +7
Recall@20 pts: decision model vs best cross-encoder, measured
logit diff
what a reranker returns
Calibrated
what a decision model returns
50 → 10
the reranker’s standard job
255
max options in one decision
32K
per pair on a modern reranker

01 — Two models that refuse to write

The similarity is not superficial, and neither is the divergence.

Ask a general-purpose model a question and it writes you a paragraph. Ask a reranker how relevant a passage is and it returns a number. Ask a decision model whether a customer is likely to cancel and it returns a probability. Three behaviours, two of them apparently boring — and it is the two boring ones that turn out to be siblings.

The shared trait is structural, not cosmetic. Neither a reranker nor a decision model ever generates a token. Both run a forward pass, read the logits sitting at one position of the prompt, and hand them back. There is no decode loop, no sampling, no output string to parse, and no cost that grows with how long the answer would have been. That is a real architectural property, and it is why the two families feel interchangeable the moment you first look at them.

The one-sentence versionA reranker answers “how relevant is this document to my query?” once per document, and the number it gives you is a logit you may only sort by. A decision model answers “which of these options is it, and how sure are you?” for many questions at once, and the number it gives you is a probability you can threshold. Same mechanism, different contract.

That contract is what the rest of this article is about. Chapters 2 and 3 put each model under the microscope, chapter 4 lines them up on six axes, chapters 5 and 6 take the two that actually decide the answer — calibration and compute shape — and chapters 7 and 8 run the substitution test in both directions, because the honest answer is not symmetric: one direction works with conditions attached, and the other does not work at all.

One note before starting: both families have been moving fast. The reranker figures here come from our own top-20 reranker ranking and from the model cards themselves; the decision-model figures come from our System One deep dive and the independent benchmark behind it. Every number is dated so you can re-check the ones that will have aged. One thing has changed since the substitution tests were first written: third parties have now ranked both families over the same candidate window and published the numbers, so chapters 7 and 8 quote those measurements directly, name who ran them, and say what they do not cover.

02 — What a reranker actually does

A cross-encoder, one forward pass per pair, and a number that is not a probability.

A reranker — more precisely a cross-encoder — takes your query and one candidate document, concatenates them into a single input, and runs them through a transformer where every token can attend to every other token. It then reads a score from that pass. Because the query and the document see each other from the first layer, this is strictly more expressive than the alternative, where each is embedded separately and compared afterwards.

Two things follow from that design. First, the cost is per pair: to score k candidates you run k forward passes (batched together, but still k inputs). This is why the accepted practice is to embed your way to 50–100 candidates and rerank only those down to the top 5–10 — never to rerank a corpus. Second, pair order matters: most cross-encoders are asymmetric, so the query goes first.

What the score is made of

This is the detail that decides the whole comparison, so it is worth being exact about it. On Qwen3-Reranker, the published implementation scores a pair with the raw logit difference logit("yes") − logit("no") read at the final non-padding position. The model card is explicit about the consequence: “By default, scores are raw logit differences. To get 0-1 probability scores, pass a Sigmoid activation function.” That sigmoid is a squashing function, not a calibration step — it puts the number in a range, it does not give it a meaning.

Older pair-classification rerankers work the same way: a single relevance logit, optionally pushed through a sigmoid. And the failure mode of treating that as a probability is documented well enough to quote directly — raw reranker outputs are logits, and they are not comparable across models or across queries. A score of 0.82 from one query says nothing about a score of 0.82 from another.

The trap everyone falls into onceBoth families return a floating-point number between 0 and 1, so the two look interchangeable in a log file. They are not. One is a monotone transform of a logit that you may sort by; the other is a probability that was trained to mean something. Put a threshold on the first without fitting a mapping first and you have written a bug that only shows up on production traffic.

How the field is actually shaped in 2026

Two architectural generations sit side by side. The classic generation scores one pair at a time. The newer listwise generation scores a whole list of documents in one pass, which is a direct attack on the k-forward-passes cost — Jina Reranker v3.5 is the current poster child at 0.6B parameters with an 8K–32K listwise window, and it reportedly rivals models many times its size on BEIR. A third approach, late interaction (ColBERT-style), pre-computes per-token document embeddings once and only does the expensive comparison at query time, which is what makes million-document corpora tractable.

RerankerSizeContextFlavour
Jina Reranker v3.50.6B8K–32KListwise — scores a list in one pass
Qwen3-Reranker-0.6B / 4B / 8B0.6B–8B32KCausal yes/no logit per pair, instruction-conditioned
BGE-Reranker-v2-M3568M512+XLM-RoBERTa, 100+ languages, the default everywhere
BGE-Reranker-v2.5-Gemma2~2B2K–8KReasoning backbone, scores contradictions
mxbai-rerank-large-v2~1.5–2B2K+RL-tuned to cut false positives
Zerank-2~1.7–4BExtendedInstruction-conditioned, tops Agentset ELO
cross-encoder/ms-marco-MiniLM-L-6-v2—512The classic pair classifier
cross-encoder/ms-marco-TinyBERT-L-24M512Ultra-fast CPU scoring

Sizes, context windows and flavours as published by each project and tabulated in this site’s reranker ranking (updated 4 August 2026). Note what the table has in common: every entry is bounded by a context window per pair, and none of them promises that a score means anything beyond the list it just scored.

03 — What a decision model does

One state, many typed questions, one distribution per question.

A decision model takes a block of state — a support ticket, a pull request, a log line, a game frame — plus a set of typed questions about that state, and returns a probability for each option you defined. It never writes text, so there is nothing to parse: your schema is the output space.

The three primitives are the whole interface. Choice picks one of up to 255 described options and returns the full distribution over them. Noul asks a yes/no question and returns a single probability. Score places the state on an ordinal rubric of 2–10 described levels and returns the probability-weighted position. Labels arrive at request time, so adding a category means editing a dictionary, not retraining.

ModelSizeBackboneNote
JevundisclosedCausal, closedThe reference; 32k state plus longest question
Von395MModernBERT encoderBest open zero-shot entrant; CPU, under 15 ms
Laya421MModernBERT + head100+ languages; fine-tune it, don’t zero-shot it
Kev0.8B–9BQwen3.5/3.8 + LoRAServes the same wire contract as the hosted API
SemIfnone addedFrozen Qwen3.5-4BTrains nothing; reads option logits directly
Nimble9BQwen3.5-9B + LoRAPublished contrastive data recipe; 8,192 context
Rizzo Flow4B / 1.7BSpark-X2.5 + LoRAllama.cpp, default 8,192 context, built-in abstention

The detail that separates these from rerankers is why they output probabilities rather than scores. They are trained against a proper scoring rule — a reward that peaks at the truth rather than at confidence — so a model that reports 0.8 is being pushed, during training, toward actually being right 80% of the time. A reranker is trained to order relevant above irrelevant, and ordering does not care what the numbers are, only how they sort.

Why this is not a free lunchCalibration is a property of a population of predictions, not of a single answer, and it moves with your data. The independent evaluation of Jev found its choice probabilities well calibrated and its binary probabilities ranking well but sitting badly against a fixed 0.5 threshold. So even here the rule stands: fit the temperature on your own labelled data before any number drives an action. The difference is that you are tuning a real probability, not rescuing a logit.

04 — The differences that matter

Six axes. Two of them are technicalities; four decide which one you want.

AxisRerankerDecision model
What is scoredOne (query, document) pair at a timeOne (state, question) pair — many questions against one shared state
What comes backA single scalar per documentA distribution over the options you wrote, plus expected values
Competitive or independentIndependent — every document can score highMutually exclusive — probabilities sum to 1
Training targetOrdering over relevance labelsA proper scoring rule, so the number is calibrated
Typical k50–100 candidates in, 5–10 out — designed to scale upA handful of questions — designed to batch horizontally
Output contractA sort orderTyped JSON: a label, a probability, an abstention

The row everyone skips: competitive or independent

This is the most quietly consequential difference in the table, so it deserves its own paragraph. A reranker scores each document on its own merits: if all 100 candidates are relevant, all 100 can score 0.9, because nothing forces the scores to compete.

A decision model’s Choice is a softmax over the options. If you hand it 10 candidate documents as 10 options, it must distribute its probability mass across them — it will confidently tell you which one is best even when every one of them is irrelevant, and the moment you add an 11th bad option the other ten shift. For picking a single winner that is exactly what you want. For judging a set of candidates where the correct answer may be “none of these”, it is a trap — and the workaround (adding an explicit “none of them is relevant” option) is a workaround, not a design.

The softmax trap, in one lineAsk a Choice which document is most relevant and it must pick one. Ask a reranker which documents are relevant and it will cheerfully say none of them. If your pipeline can legitimately return an empty result set, that difference is the whole decision.

05 — Calibration, the real divide

Where a ranking instrument and a measuring instrument come apart.

Everything up to here has been about shape. This chapter is about meaning, and it is the reason the substitution tests come out asymmetric.

A reranker’s job is to produce an ordering. Given one query and one list, the only thing you do with the scores is sort them and take the top k. For that job, calibration is irrelevant — multiplying every score by 7 changes nothing about the sort. So rerankers are not trained for it, and their scores are not comparable across queries: a 0.9 on one query and a 0.4 on another can describe equally relevant documents.

A decision model’s job is to produce a threshold. The moment an application says “act when the probability exceeds 0.8”, the number stops being an ordering and starts being a measurement, and it has to survive being compared to outcomes. That requires the reward to peak at the truth: a proper scoring rule such as the log score or the Brier score, where reporting 0.7 when the true rate is 0.7 scores better than reporting 0.9 or 0.5. A 0/1 reward on a sampled answer has no such property — it peaks at 1.0 regardless of the truth — which is precisely why you cannot get calibration by accident.

You want to…With a rerankerWith a decision model
Put the best candidate firstNative. That is what it is forWorks for a short list, wastes the probability
Return nothing when nothing fitsNot possible without fitting a cutoffNative. Abstention is part of the contract
Act only above 0.8 confidenceNot possible without fitting a calibratorNative, after a temperature fit on your labels
Compare confidence across requestsNot meaningfulMeaningful by construction, drifts with domain
Return a label a program can switch onSort order onlyNative. Typed JSON, no parsing

The middle column is not a criticism of rerankers — nobody chose a badly designed contract. It is that “rank these for me” and “tell me how likely this is” are different questions, and only one of them has an ordering for an answer.

If you genuinely need a threshold and only have a rerankerYou can have one, but you pay for it: collect labelled pairs from your own traffic, fit a calibrator — Platt scaling, isotonic regression, or a temperature fit — and refit whenever the model, the prompt or the data drifts. This is entirely legitimate and it is what careful teams do. It is also exactly the work that a model trained on a proper scoring rule does for you during training instead.

06 — Compute shape

Both skip generation. Only one of them gets cheaper as you ask more questions.

It would be easy to assume that a decision model’s advantage is “no decode loop”. It is not, because rerankers have no decode loop either — they also read a single position and stop. So the interesting question is what happens to cost as the amount of work grows.

For a reranker the cost is k forward passes over k (query, document) pairs, and the query is repeated in every one of them. Batch them and you amortise the launches, not the arithmetic. Listwise models fix part of this by scoring a list at once, and late-interaction models pre-compute the document side entirely — both are real answers, and both narrow the gap this section used to own.

For a decision model the state is prefilled once and the questions branch off it, sharing its key/value cells instead of copying them, read at one position each. The measured shape of that is flat in the question count:

  • 2 options 76 ms, 200 options 75.5 ms — the option list barely registers.
  • 5,000 questions over a 22.8k-token state in 1.8 s — batching, not serialisation.
  • One prefill plus two micro-batches for 8 yes/no questions over a 218-token contract, 136 ms total, instead of eight separate passes.
  • Reusing one long state across many criteria lifted throughput from 2.33 to 20.03 decisions per second in SemIf’s measurement.
The honest limitation on “flat”Flatness applies to questions about shared state, not to state you have not paid for yet. If your 100 candidates are 500 tokens each, someone pays for 50,000 tokens of prefill either way — the decision model does not get that for free, and at that size it is also over the 8,192 context of most open models (Jev documents 32k). The real asymmetry is the one that matters for this article: many short questions over one long document is where a decision model wins outright. Many long documents each needing their own judgement is a reranker’s shape, and listwise models have already closed most of it.

07 — Test A: can a decision model replace your reranker?

Short list: yes, with conditions. Real retrieval: no.

Let us be concrete. You have a query and ten short candidate passages, and you want the best one — or you want to know whether any of them is good enough to use. There are two ways to put that to a decision model.

# Option 1 — one Choice over the candidates: cheap, competitive
{"state": {"query": "does the refund policy cover partial refunds?"},
 "questions": {"best_doc": {"type": "choice",
   "criteria": {"doc_1": "<passage>", "doc_2": "<passage>", ...}}}}

# Option 2 — one Noul per candidate in a single request: independent,
#            each returns a probability you can threshold
{"state": {"query": "...", "doc_1": "<passage>", "doc_2": "<passage>"},
 "questions": {"relevant_1": {"type": "noul", "instructions": "Is doc_1 on topic?"},
               "relevant_2": {"type": "noul", "instructions": "Is doc_2 on topic?"}}}

Option 2 is the one worth noticing. The state is prefilled once, the questions are isolated from each other by design — a question cannot see a sibling’s answer — so each candidate is judged independently and returns its own calibrated probability. That is the property the reranker also has, plus the property the reranker lacks. You can also ask, in the same single request, things a reranker has no way to express: is it relevant, is it authoritative, is it stale, should we ship it.

Where it stops working

  • Candidate count. A Choice caps at 255 options, and that is before context. Retrieval workloads are sized in the hundreds. The measured 200-candidate run below got there with one Noul per document in six shared batches rather than one Choice — which works, but charges you a question and its share of context per candidate.
  • Token budget per question. Laya’s options for a single question must fit a 192-token budget in total — which is exactly why a 77-label problem leaves three or four tokens per label. Ten full passages will not fit that budget in any of these models.
  • Context per pair. Qwen3-Reranker takes 32K per pair. The decision models default to 8,192 (Nimble, Rizzo Flow) or document 32k including the longest question (Jev), and Laya’s English checkpoint truncates the state at 512 tokens in the ONNX client.
  • Independent scoring. Option 1 hits the softmax trap from chapter 4. Option 2 avoids it, but you are now running one question per candidate and paying for the question suffixes.
  • Ecosystem. Rerankers are wired into FlagEmbedding, sentence-transformers, every RAG framework, and support late interaction and pre-computed document embeddings. Decision models have no recall stage at all — you still need something else to produce the candidates.
Bar chart of published context limits comparing rerankers per query-document pair against decision models for the whole state: Qwen3-Reranker 4B and 8B and Jev at 32,768 tokens, Nimble 9B and Rizzo Flow 4B at 8,192, and BGE-Reranker-v2-M3 and the Laya English ONNX client at 512
The budgets are not measured the same way, and that is the point: a reranker gets its window per pair, re-reading each document from scratch, while a decision model spends one window across a shared state plus every question asked against it. Ten passages do not fit a 192-token option budget, and one long document costs a decision model most of its window — which is where the substitution test stops.

What has actually been measured

Up to here the substitution test has been argued from contracts. Third parties have now run part of it, and the numbers are worth having in front of you. In September 2026 the search team at Opine took 64 sales queries, built one candidate window per query (the top 60 and the top 200 from their own hybrid retrieval), graded every candidate with an independent model (GPT-6 Astra), and ranked the same window several ways: with Jev used as a reranker — one Noul question per document, batches of 20–30 — against dedicated cross-encoders (Cohere Rerank 3.5, 4 Fast and 4 Pro, Voyage Rerank 3) and against frontier LLMs ranking the whole list at once (GPT-6 Luna, GPT-6 Sol, Claude Opus 5.5).

Same candidate set, same graderDecision model as reranker (Jev)Best dedicated cross-encoderFrontier LLMs
Recall@20 gain over hybrid search, 60 candidates+12 pts+7 (Voyage Rerank 3)GPT-6 Luna +11; Sol and Opus 2–3 pts higher
Recall@20 gain, 200 candidates+17 ptsnot reported at that window—
nDCG@20, 200 candidates0.8760.815 (Voyage Rerank 3)0.864 (Luna); Sol and Opus ~0.92
Median latency per search0.4 s parallel (0.9 s at 60, 1.9 s at 200 sequential)0.3–0.7 s—
Cost per 1,000 searches$0.80 at 60, $2.19 at 200list pricessignificantly higher
Top-20 overlap with the best frontier model57% (Opus), 58% (Sol)—Opus and Sol agree 72%

Opine, September 2026 — 64 sales-related queries, 60- and 200-result windows from one hybrid retriever, documents graded by GPT-6 Astra rather than by people. A second, smaller measurement from MindStudio in the same month put a decision-model rerank step on top of a plain BM25 baseline and moved top-1 accuracy from 21% to 54% without touching the retriever, at about 17 decisions per second sequentially (roughly 40 s per batch), falling to 7.6 s with 64 requests in flight. Both are small-n, LLM-graded evaluations run by teams using the model rather than by its makers — Opine calls its own a “vibe check” and lists human label validation among its next steps — so read the magnitudes as indicative and the ordering as the interesting part.

Notice as well what the table does not cover: no test touched the recall stage, none went past 200 candidates, none ran multilingual corpora or pre-computable embeddings, and all of them sat inside the limits listed above (roughly 1,000-character documents, a window that fits a shared state). The measured numbers firm up the shortlist half of the verdict; they do not probe the retrieval half, which remains an argument from structure.

The verdict on Test AFor a shortlist of a handful of short passages where you want a threshold or an abstention, a decision model replaces a reranker well, and gives you back something the reranker cannot. For the retrieval job itself — hundreds of candidates, long chunks, multilingual recall, pre-computable embeddings — it does not, and you should not want it to. The honest recommendation is that they sit in sequence: rerank to a shortlist, then decide. Measured, the first half now has numbers behind it: on Opine’s 64-query window the decision model beat every dedicated cross-encoder at 60 candidates (+12 vs +7 Recall@20 points) inside their latency band and under $1 per 1,000 searches — while the second half stays an argument from structure, because nobody has published that test.

08 — Test B: can a reranker replace a decision model?

The shorter test, because the answer is shorter.

A reranker is not as fixed as it first appears. Its instruction field lets you define what “relevant” means for this query, and instruction-conditioned models like Zerank-2 take that seriously. So you can point one at a binary proposition: set the instruction to “an urgent customer support ticket” and read its yes/no logit. That is, near enough, what Qwen3-Reranker does internally on every pair.

So it fakes one binary decision. What it cannot fake:

  • Calibration. You get a logit. You cannot threshold it without fitting a mapping, so the number cannot drive an action on its own.
  • Typed output. No label to switch on, no expected value on a rubric, no probability per option — one scalar per pair, and the scalar only orders.
  • Multiple questions over shared state. A decision model returns a state, five independent answers and five abstention flags in one request. The reranker returns five pairs’ worth of scores from five inputs.
  • Abstention. There is no “I cannot tell from this state” in a relevance score, and no way for a threshold to express “this is out of domain” without a fit behind it.
  • Rubrics. “How severe is this incident, 1–5?” has no document to rank and no ordering to produce. It has an answer.

What each family is willing to publish

There is a measured version of this verdict, and it is a measurement of what gets measured. Decision models publish calibration error next to accuracy, because their numbers are meant to be thresholds. On the JevBench v1.4 board Von 1.2 posts a sealed-half expected calibration error of 0.107 against Laya’s 0.172, under one protocol. Laya’s own model card reports 0.766 accuracy, Brier 0.062 and ECE 0.213 over 2,000 decisions, breaks it down by primitive (noul 0.857, choice 0.733, score 0.723) and then warns in its own limits section that the confidence is still uncalibrated and should be refit on held-out data. Jev’s figures, as quoted in that same card, are accuracy 0.727 with ECE 0.144. Latency is reported per decision: Von runs a p50 of 0.096 s on four vCPUs under OpenVINO and 0.023 s on an A10G, and 32.8 ms p50 across the 38-benchmark, ~150k-request Decision Index.

Reranker boards, by contrast, publish ordering and only ordering. BEIR reports nDCG@10 — the Jina Reranker v3 paper quotes 61.94 nDCG@10 as the best among the rerankers it evaluated — and no expected calibration error, no Brier score and no cutoff quality appears anywhere in the standard tables. It is not an omission: a raw logit is not a probability until somebody fits a mapping, so the boards have no number they could report. That asymmetry is the Test B result, measured. You can audit whether a decision model’s number may drive an action; for a reranker the audit does not exist until you have fitted the calibrator that the decision model was trained to make unnecessary.

The verdict on Test BNo. A reranker can imitate a single binary decision badly, and nothing else. If what you actually need is a ranking, a ranking model is the correct and cheaper tool; if what you need is a judgement with a number attached, the reranker will hand you a logit wearing a probability’s clothing. That specific confusion is the most common way these two categories get mixed up in production. The measured form of “no” is the row that does not exist: nobody publishes a calibration error or a cutoff quality for a reranker, because until a mapping is fitted there is no probability to score — and the moment you fit one, you have rebuilt by hand the head the decision model ships with.

09 — Where each one belongs

They are not competitors. They are adjacent stages.

Once the substitution tests are done, the placement is easy: the two models sit on either side of a boundary neither of them crosses.

StageModelThe job
RecallEmbedding model / hybrid searchTurn a corpus into 50–100 plausible candidates
PrecisionRerankerOrder those candidates, keep the top 5–10
GenerationFrontier LLMWrite the answer from the retrieved context
ArbitrationDecision modelRoute, guardrail, verify and threshold — around the model, not over the corpus

The distinction to carry away is one of direction. A reranker looks outward at a candidate set that the world supplied and asks which of those matters. A decision model looks at a state you assembled and asks a question you formulated. Retrieval is a problem with an answer set already in hand; judgement is a problem where you have to define the options first. Same mechanism, opposite direction of travel.

Where they genuinely overlapExactly one place: a small candidate set plus a decision about it. “Which of these five approaches should we take?” “Is any of these three documents safe to show a customer?” Both families answer that, and a decision model answers it with an extra property you will eventually want. Everywhere else they are complements rather than substitutes — and in an agent harness the reranker’s only appearance is normally inside memory retrieval, where it was going to live anyway.
Four-stage pipeline showing where each model sits: recall by an embedding model producing 50 to 100 candidates, precision by the reranker keeping the top 5 to 10, generation by a frontier model that writes the answer, and arbitration by the decision model that routes, guards, verifies and thresholds around the model
The two families never meet, because they face opposite directions. A reranker looks outward at candidates the world supplied and asks which of those matters; a decision model looks at a state you assembled and answers a question you formulated. Same mechanism, opposite direction of travel.

10 — Pick by job

The shortest useful table in the article.

If you want…UseWhy
50–100 retrieved chunks trimmed to the top 5–10RerankerDesigned for it; listwise models make it cheap
Long documents, 32K per pairRerankerQwen3-Reranker takes 32K per pair today
Millions of documents, pre-compute onceLate-interaction rerankerColBERT-style embeddings are pre-computable
Multilingual recall with one small modelRerankerBGE-Reranker-v2-M3: 568M, 100+ languages, default everywhere
Route a ticket into typed categoriesDecision modelLabels at request time, probabilities per option
Score on an ordinal rubricDecision modelScore has no reranker equivalent
Act only when confidence clears 0.8Decision modelCalibrated by construction, after a local fit
Return nothing when nothing qualifiesDecision modelAbstention is native; a reranker must be fitted first
Many checks over one long document in one callDecision modelOne prefill, questions nearly free
“Is this document relevant?” for a shortlist, then decideBoth, in that orderRerank to shortlist, then judge with thresholds

11 — FAQ

The five questions that come up every time.

Is a reranker just a decision model with one question?

Architecturally, closer than the naming suggests: both read logits at one position and generate nothing. Contractually, no. One returns a scalar that only orders, trained to rank; the other returns a distribution that means something, trained on a proper scoring rule. The difference is not the architecture, it is the target.

Should I threshold a reranker score?

Not as it stands. Reranker outputs are raw logits, not comparable across models or queries, and a sigmoid only squashes them. If you need a cutoff, fit a calibrator on your own labelled pairs and refit on drift. If you need a threshold often, that is a signal you wanted a decision model.

Could a decision model do my RAG reranking?

For a handful of short candidates, yes — and at that scale the answer is now measured rather than argued: in Opine’s 64-query test a decision model reranking 60–200 candidates beat every dedicated cross-encoder it faced on Recall@20 and nDCG@20, at 0.4 s median latency and under $1 per 1,000 searches. For the retrieval job, no: candidate counts, the 192-token option budget, context limits, independent scoring and the absence of recall-stage tooling all say no. Rerank to a shortlist, then decide.

Do both really generate zero tokens?

Yes, and it is the reason they get confused. Neither pays a decode loop, so neither has output that grows with answer length, and neither produces text you have to parse. It is the single strongest thing the two families have in common — and the reason “it gives me a number in 0–1” feels like a shared contract when it is not.

Which one is cheaper?

It depends on the shape of the work, not the model. A reranker costs you one forward per candidate and nothing else. A decision model costs a prefill over the state and then almost nothing per question. So: many questions about one shared state favours the decision model; many independent long documents favour the reranker, especially listwise ones.

12 — Sources

Where each claim came from.

#SourceUsed for
1Qwen/Qwen3-Reranker model cards (0.6B / 4B / 8B) and the NVIDIA NeMo reranker component docsThe raw logit difference logit("yes") − logit("no"), and the card’s “pass a Sigmoid to get 0-1” caveat
2AIMultiple reranker benchmark; published guidance on reranking deploymentsCausal yes/no scoring, and the warning that raw outputs are not comparable across models or queries
3This site, Top 20 Reranker Models 2026 (4 Aug 2026)The reranker table, listwise and late-interaction flavours, the 50→10 practice, pair ordering
4This site, Embedding, Reranker & OCR Models 2026Reranker sizes, licences and hardware requirements
5This site, System One decision models deep diveThe decision-model family, primitives, KV-sharing mechanism and hosting paths
6jabr classifier-benchmark (CC0), v2 suite: 49 tasks / 866 casesThe independent accuracy numbers behind the decision models discussed
7Deußer, Sparrenberg & Sifa, Evaluating and Benchmarking the System One Model Jev (arXiv:2609.37647)32k state limit, 255-option Choice, price, and the calibration/selective-prediction findings
8Laya ONNX client README (receptron/laya) and Laya model cardsThe 192-token option budget and the 512-token state truncation
9Rizzo Flow, SemIf, Nimble and Kev repositoriesThe batching, throughput and context figures quoted in chapter 6
10Opine, Using Jev to improve search quality (September 2026)The same-candidate-set head-to-head in chapter 7: Recall@20 and nDCG@20 at 60 and 200 candidates, latency, cost per 1,000 searches, top-20 overlap
11MindStudio, JEV as a Steerable Reranker (21 September 2026)BM25 top-1 accuracy 21% → 54% with a decision-model rerank step; ~17 decisions/s sequential, 40 s → 7.6 s at 64-way concurrency
12Von model card (wfzyx/von, Hugging Face), JevBench v1.4 boardSealed ECE 0.107 vs Laya 0.172; 0.096 s p50 on 4 vCPU, 0.023 s on A10G; Decision Index p50 32.8 ms over ~150k requests
13Laya typed-decisions model card (convaiinnovations/laya-typed-decisions)Accuracy 0.766, Brier 0.062, ECE 0.213, the per-primitive breakdown, the card’s own uncalibrated-confidence warning, and Jev’s quoted 0.727 / ECE 0.144
14Ma et al., jina-reranker-v3: Last but Not Late Interaction (arXiv:2509.25085)61.94 nDCG@10 on BEIR — the ordering-only metric reranker boards report, and the reason chapter 8’s calibration row is empty
15This site, scripts/benchmark-substitution.pyThe runnable protocol behind chapters 7–8: one shared candidate window, seeded fit/eval split, threshold, abstention, agreement and latency reporting, with an offline --self-test for its metric core

Assembled 6 October 2026. Reranker sizes and context windows are as published by each project; the reranker ranking was last updated 4 August 2026 and will have moved. Decision-model calibration figures are the projects’ own unless the source is named as independent. The verdicts in chapters 7 and 8 are ours — but they no longer rest on argument alone: third parties have now ranked both families over the same candidate window (Opine, MindStudio) and published the numbers quoted in chapter 7, and the calibration row in chapter 8 is filled in from the model cards themselves. What still does not exist is a formal two-family study with human-labelled relevance at retrieval scale; the protocol for one — build the window once, split queries into fit and eval halves, report ranking, cutoff quality, abstention, agreement and latency — ships with this site as scripts/benchmark-substitution.py, and its metric core can be verified offline with --self-test before any model or dataset is involved.

Need help with this?

Tell me what you’re working on and I’ll help you work through it — where you got stuck, what you’re trying to build, which model to pick. Your message arrives with this article attached, so I’ll know exactly what you’re reading.