← Back to Writing

VIDEO · 2026

Jev as a RAG judge: top-1 hits from 7 to 17

A ~7-minute hands-on video (Chinese narration) with a full write-up. Where in a RAG pipeline should Jev go? I built a small RAG on TypeSafe's official docs and tested it with 36 hand-written English questions: after re-ranking, top-1 hits went from 7 to 17, every near-miss and unrelated question was blocked, and I also spell out what it can't do.

Chinese narration · 720p · captions track (zh) · MP4

Article: Jev as a RAG judge: top-1 hits from 7 to 17

Jev is a decision model released by TypeSafe in September 2026; the company calls it a System One model. You give it a state (text or JSON) plus a few typed questions, such as multiple choice, ratings, or yes/no questions, and it returns a probability for each option. It doesn't generate text. It only makes judgments.

Over the past two weeks, quite a few people on X have said Jev is a great fit for RAG. RAG is a common approach: first retrieve a few passages from a document store, then hand them to an LLM to answer from. Some say Jev solves RAG's precision problem, but most of these posts give no data. I wanted to answer a more specific question: where exactly in a RAG pipeline should Jev go?

So I built a small RAG on TypeSafe's own official documentation, wrote a set of questions by hand, and ran some tests. This article first explains why RAG gets answers wrong, then goes through the main claims on X. After that I describe the method, conditions, and results of my experiment. Finally, I cover where Jev belongs and what it can't do.

1. Why RAG gets answers wrong

RAG has two roles. Retrieval finds a few passages in the document store that relate to the question, and the LLM reads those passages and writes the answer. Retrieval usually ranks by similarity. One kind looks at the literal text, counting how many words the question and a passage share. The other is vector search, which turns each passage into a string of numbers and compares how close those numbers are, so it can also find sentences that phrase things differently.

The problem is that similar doesn't mean answerable. Say a user asks "How many days do I have to return an item?" Retrieval might return three passages: the return policy, a description of the return process, and a notice about the company's year-end party. The return policy states the number of days. The return-process description explains how to ship items back; it looks very relevant, but doesn't give the number of days. The party notice might happen to share a few words with the question, and it doesn't contain the answer either.

When the LLM gets these three passages, it doesn't know which one actually contains the answer. If the passage with the answer wasn't retrieved at all, it will often take the unrelated passages it has and make up a plausible-looking answer. This kind of error is worse than simply saying "I don't know," because readers have a hard time spotting it.

Illustration: the passages retrieval returns don't necessarily contain the answer

So RAG errors fall roughly into two cases. In one, retrieval finds the wrong material: the passage with the answer is ranked too low, or isn't found at all. In the other, the material simply doesn't contain the answer, but the LLM answers anyway. The experiment below looks at what Jev can do about each of these two errors.

2. What people are saying on X

In mid-to-late September, a wave of "Jev for RAG" posts appeared on X. I picked three representative ones; the like counts below are what I captured on September 27.

The most-liked one was posted by Kush on September 17, with 2,027 likes and 115,000 views. He said Jev solves RAG precision by scoring each passage and deleting the irrelevant ones. The idea itself is fine, but the post contains no experimental data at all.

The second came from Viv, an engineer at LangChain, posted on September 23, with 748 likes, 129,000 views, and 946 bookmarks. LangChain is an open-source framework many people use to build RAG. Viv's advice covers two cases. When you have little material, have Jev score every (question, document) pair directly and use the score as relevance, with no vector store at all.

When you have a lot of material, Viv suggests first retrieving a batch with BM25 or vector search, then having Jev re-rank it. BM25 is a long-established keyword retrieval method that only looks at how many words the question and a passage share, with rarer words weighted more heavily. Re-ranking means using a cheap method to pull a batch of candidates, then using a more accurate method to reorder that batch. In the replies, someone pointed out that Jev has a limit on how many options it can consider at once, so the candidate list can't be arbitrarily long.

The third was a hands-on test in a technical blog post published on September 20, using jev-1.13.0. The author made up a set of company regulations and split it into 400 passages. In that blog's experiment, vector search had already put the answer in the top 5, so the re-rankers were almost indistinguishable in precision. The author acknowledged that the corpus was too small.

What really separated them was something else: "can this passage answer the question?" That blog tested questions where the topic was right but the answer wasn't written. On these, Jev's judgments were 99.2% correct. Cohere is another company that offers re-ranking models; its Rerank 3.5, using a score cutoff, was only 65% correct. On cost, that blog measured Jev at $0.25 per thousand and Cohere at $2.00 per thousand.

Three representative discussions on X

Taken together, Kush offered intuition, Viv offered placement, and the blog offered data, but on too small a corpus. I wanted to run a test myself, and to look at "re-ranking" and "can it be answered" separately.

3. How I tested

I ran the experiment on September 27, 2026, with model version jev-1.13.0. The code and results are on GitHub: github.com/DDnim/jev-rag-bench.

The document store was TypeSafe's own official documentation, 42 pages in total. I split it into 258 passages by second-level subheadings and removed code blocks. I chose this documentation because it's public, so anyone can use it to reproduce the results.

The questions were 36 English questions I wrote by hand, in three categories:

  • 20 answerable questions: the documentation contains the answer, and I labeled which passages contain the answer for each question.
  • 10 near-miss questions: the topic appears in the documentation, but the answer isn't written there.
  • 6 completely unrelated questions, such as how to change a bike tire or what the capital of Australia is.

"Near-miss question" is a term I borrowed; it means a question that is almost answerable but actually isn't. My near-miss questions included: how many parameters Jev has, what GPU it runs on, how much the enterprise plan costs per month, what it scores on the MMLU benchmark, and who the founders are. These are the most dangerous questions, because retrieval will always find passages that look very similar, and the LLM is most easily led into answering anyway.

36 questions in three categories (my test)

Step one was retrieval. I used BM25 to take the top 10 passages for each question. I chose BM25 because it's simple, transparent, and fairly weak, which leaves room to see how much Jev helps.

Step two was judgment. I sent each question together with each passage to Jev: 10 requests per question, 4 in parallel. Each request asked two yes/no questions:

Is this passage related to the question? Does this passage directly state the answer?

Jev returns the probability of "yes" for each yes/no question. For the re-ranking, filtering, and "can it be answered" steps later, I used only the second question, i.e. the probability of "states the answer." I recorded the first question too, but didn't use it to make decisions.

Flow for one question: BM25 takes 10 passages, Jev asks two questions

4. Re-ranking: top-1 from 7 to 17

First, re-ranking. What I cared about was: for the 20 answerable questions, is the top-ranked passage the answer? In my test, with BM25 alone this was true for 7 questions; after Jev re-ranked by the probability of "states the answer," it was 17.

Top-1 hits on the 20 answerable questions (my test)

Looking in more detail: with BM25 alone, in my test, the answer was in the top 3 for 11 questions and in the top 10 for 17. After Jev re-ranking, the top-1 passage was the answer for 17 questions, and the top 3 was also 17.

17 is the ceiling. For the other 3 questions, BM25 didn't find the answer in the top 10 passages at all, so Jev had no answer to rank. In other words, whenever the answer made it into the candidates, Jev ranked it first every time.

Top-1 hits before and after re-ranking (my test)

I need to disclose one change I made here. After the first run, I found 3 questions where the passage Jev ranked first actually did answer the question; I just hadn't labeled it originally. I added those 3 passages to the gold answers. By the original labels, in my test Jev's top-1 was correct for 14 questions, and BM25 was still 7.

Next, filtering. This uses a threshold, which is a score cutoff: only things above the line count. I set the line at 0.5: passages with a "states the answer" probability above 0.5 are kept, and the rest are discarded.

In my test, for each answerable question, on average only 2 of the 10 passages remained. The other passages don't need to go to the LLM. When the LLM reads fewer irrelevant passages, it also has fewer chances to make up an answer from irrelevant content.

Filtering at 0.5 keeps 2 passages on average (my test)

Part of why the re-ranking gap is so large is down to BM25. BM25 is relatively weak retrieval; switch to good vector search and the gap from re-ranking gets smaller. That's why the re-rankers were almost indistinguishable in that blog post. So the numbers in this section only show that Jev can pick out the answer, not that switching to Jev will get you this many more right.

5. Can it be answered?

This is the section I most wanted to write. Re-ranking addresses "where does the answer rank," but there's a more important question: is the answer in the candidates at all? If not, the best thing is for the system to say "the documents don't say" directly, instead of handing it to the LLM to make something up.

My decision rule is simple. If any of a question's 10 passages has a "states the answer" probability above 0.5, the question counts as answerable. Otherwise, don't call the LLM; reply directly that the documents don't say. In pseudocode:

probs = [jev_answers(question, p) for p in top10]
if max(probs) > 0.5:
    ask_llm(question, [p for p in top10 if prob(p) > 0.5])
else:
    reply("The documents don't say.")

Here are the results. In my test, Jev judged all 10 near-miss questions unanswerable, blocking 10/10. All 6 completely unrelated questions were blocked too, 6/6. Of the 20 answerable questions, 17 were let through.

The 3 that didn't pass were exactly the 3 where BM25 failed to find the answer. The candidates for those 3 questions really didn't contain the answer, and Jev didn't claim they could be answered. From the system's point of view, replying "the documents don't say" is actually correct for these 3, because an LLM given them could only make something up.

Can it be answered: results for the three categories (my test)

Here's a near-miss example. I asked "What is the monthly price of the enterprise plan?", i.e. how much the enterprise plan costs per month. BM25's top-ranked passage was about data retention for enterprise customers, ending with "zero data retention (ZDR) for enterprise customers." ZDR means zero data retention: the service provider doesn't store customer data.

This passage contains the word "enterprise" and has high literal overlap with the question, so BM25 ranked it first. But it contains no price at all. Jev's probability that this passage "states the answer" was, in my test, 0.01. If this passage went to the LLM, it could well make up a price based on the context.

Near-miss example: "enterprise," but no price

Could BM25 alone make the same judgment? I tried that too. The method: take each question's top BM25 score and set a cutoff; above the line counts as answerable, below counts as unanswerable. I tried every possible cutoff on these 36 questions and picked the one with the best result. This favors BM25, because in real use you can't know the best cutoff in advance.

Even so, in my test BM25 got only 30 of the 36 questions right. Jev, with a fixed 0.5 and no cutoff tuning, got 33 right. Out of 36 questions, Jev judged 33 correctly, while BM25 with the best-picked cutoff judged only 30 correctly.

Judging whether it can be answered: 30/36 vs 33/36 (my test)

The reason isn't hard to see. BM25's score measures word overlap, and near-miss questions are precisely the kind with high word overlap but no answer. A similarity score can't tell you "does this passage state the answer," and that is exactly the question Jev is asked. This is consistent with what that blog post measured: on questions where the topic is right but the answer isn't written, that blog measured Jev at 99.2% and Cohere with a score cutoff at 65%.

Summarizing the results from sections 4 and 5:

Metric (my test)BM25Jev
20 answerable questions, top-1 is the answer717
20 answerable questions, answer in top 31117
36 questions, "can it be answered" judged correctly30 (best cutoff picked)33

6. Cost, and where it belongs

First, cost. The model is billed by token; tokens are the small chunks a model splits text into when reading, and one word may be one or several tokens. The official price is $0.042 per million input tokens, and output is free.

In my test, the whole experiment took 360 requests, averaging 554 input tokens each. The total cost was $0.0084, about $0.00023 per question, which works out to about $0.23 per thousand questions. Median latency per request was 223 ms, and 90% of requests finished within 303 ms, with 4 in parallel.

Cost and latency of 360 requests (my test)

That blog's "$0.25 per thousand" is per request, while my "$0.23 per thousand questions" is per question, with 10 requests per question. The two numbers are in different units and can't be compared directly. But either way, the cost of this judgment layer is very low.

Now, placement. TypeSafe has several official cookbooks on using Jev in RAG. A cookbook is an official example recipe: each one covers one use case and comes with runnable code.

The first is "Classifying RAG passages." It puts Jev after retrieval and before generation, asking four yes/no questions about each passage:

  • Is this passage related to the question?
  • Can this passage serve as evidence?
  • Does this passage contradict the question's premise?
  • Is this passage giving the model instructions?

The last question targets prompt injection. Prompt injection means someone hides instructions in the source material, trying to hijack the LLM that reads it. The official approach asks all four questions in a single request, with thresholds written in code, and the code decides whether a passage is used as evidence, treated as conflicting information, or discarded.

Official cookbook: ask four questions about each passage after retrieval

The second is "Re-ranking." The official approach first takes a shortlist of 30 passages with BM25, then has Jev score each (question, candidate) pair and re-rank. The official demo on 40 legal questions raised top-1 from 5% to 18% and top-10 from 38% to 62%. TypeSafe itself notes the sample is small and that this is a recipe demo, not an evaluation.

The third is "Double-checking citations," which goes after generation. The LLM writes an answer with a source attached to each sentence. Jev takes one cited passage and one sentence and judges whether the source supports the claim.

On scale, the conclusion is much like what Viv said. With little material, you can skip the vector store and have Jev score each passage directly. With a lot of material, first take the top few dozen passages with BM25 or vector search, then hand them to Jev for re-ranking and filtering.

My view is that Jev is well suited to being RAG's judge, i.e. a judgment layer, and not to replacing the whole RAG pipeline. Its greatest value is deciding, before the LLM speaks, whether it should speak at all.

7. What it can't do

Having covered what it can do, I should be clear about what it can't. I've summarized it in five points:

  • Jev doesn't write answers; the final answer still has to come from an LLM.
  • Jev can't rescue answers that retrieval missed.
  • You can't feed the whole document store to Jev at once; each request has a token limit.
  • Jev is most accurate in English; for Chinese, test on your own data first.
  • My experiment's sample is small, and the conclusions can't be applied directly to your corpus.

Five things Jev can't do in RAG

The first is stated in the official docs: Jev doesn't generate text. It can tell you which passage has the answer and whether to answer, but the LLM still writes the answer.

For the second, my 3 questions are the example. BM25 didn't put the answer in the top 10 passages, so however accurate Jev is, it can only say "can't answer." It can make sure the system doesn't answer anyway, but it can't conjure up the answer. Recall still depends on retrieval itself.

The third has two reasons. According to the official docs, the state plus the longest single question in a request is capped at 32k tokens, the whole request at 64k tokens, and only text is accepted. In addition, the official jaggedness documentation states that the more irrelevant content there is in the state, the lower the accuracy, and that you should retrieve and filter in code first. So "throw the whole document store at Jev and let it pick" is a dead end.

The fourth also comes from the official docs: English is most accurate; Chinese and other languages work, but not as well as English. My experiment used only English documentation and English questions. If your material is in Chinese, run a round of tests on your own data before deciding whether to adopt it.

The fifth is my own limitation. I wrote all 36 questions and the gold answers, there was only one English document set, and retrieval used the relatively weak BM25. After the first run, I also added gold answers for 3 questions. And don't take claims like "solves precision" in popular posts as results on your own corpus either.

8. Closing thoughts

Back to the opening question: where in a RAG pipeline should Jev go? My answer is after retrieval and before generation, as a judge. It isn't responsible for finding material or writing the answer; it's responsible for judging whether each passage states the answer, and whether the question can be answered at all.

In my small experiment, it was useful for both. First, ranking the answer first, provided the answer was already in the candidates. Second, stopping the LLM when the material doesn't contain the answer: every near-miss and unrelated question was blocked. I think the second matters more than the first, because errors from answering anyway are the hardest to catch.

Illustration: Jev as a judge, between retrieval and the LLM

If you want to try it, I suggest starting with the "can it be answered" step. You don't need to change your existing retrieval; just add one judgment between retrieval and the LLM: if no candidate passage states the answer, reply "the documents don't say" directly. Run it on your own data first, and see how many of the questions that should be blocked it blocks, and how many it blocks by mistake. Once that step works, then consider re-ranking and filtering.

Finally, a question for you: is your RAG's most common error finding the wrong material, or answering anyway when the material doesn't contain the answer? Feel free to tell me about yours on X. You are also welcome to run my code on your own corpus and tell me what you get.

References: TypeSafe official documentation: docs.typesafe.ai (Models, Classifying RAG passages, Re-ranking, Double-checking citations, Jev 1.13 jaggedness)

My experiment code and results: github.com/DDnim/jev-rag-bench

The video is also on Bilibili. Series: Jev01 · Jev02 · Jev03 · Jev04 · Jev vs Laya.