selfjev/GitHub
THE OPEN NOTEBOOK / RESEARCH

Show your work.

What can it do? How well does it work?
The results, and the experiments behind them.

56 recorded experiments3 ways to test decisions1 current model
WHAT THE SCORES MEAN

Test the work
you need it to do.

01 / TEXT DECISIONS

Read. Understand. Choose.

Can it answer questions about a piece of text, including exceptions and several possible answers?

1,991 questions · yes/no, pick one, select all
02 / AI RESPONSE REVIEW

Check the AI’s work.

Can it judge answer quality, check supporting evidence, apply a policy, or spot an attempt to bypass instructions?

946 questions · quality, accuracy & safety
03 / BROADER TEXT TASKS

Go beyond one use case.

Can it recognize intent, topics, sentiment, and other text patterns? These tasks helped guide development.

3,471 questions · development progress only

Run the evaluation yourself. Decision Bench publishes Text Decisions, AI Response Review, and a separate 720-question Record Reasoning suite. The broader development benchmark above is not included.

Dataset, scoring tools & methodology
AVAILABLE ON HUGGING FACE

Get the weights. Explore the evidence.

Jwuthrich

How often did the model choose the expected answer? The same 1,991 questions test yes/no decisions, choosing one answer, and selecting all that apply.

MODELANSWERS MATCHED
01
JevTypeSafe’s hosted decision model
97.2%
02
SelfJevOur current model · self-hosted
95.7%

Scores are the share of answers matching the expected result; for “select all,” every choice must match. Only models with a recorded result are shown. Click a row for its report.

A useful signal, with a clear scope.

Each score is the percentage of answers that match the expected result on that test. Jev still leads on text decisions. SelfJev brings comparable results on these questions to a model you can run yourself.

Our two focused tests were written and checked by AI, whose answers can still be wrong. They help us assess progress; they do not guarantee the same accuracy on your data. We have not tested every model on Hugging Face.

Test sources, methodology & limitations

“Text decisions” is the project’s eval2 test; “AI response review” is eval_llm. Both use fixed questions kept out of training, with expected answers checked by two AI judges. They have informed the research direction, so a fresh final test is still needed. “Broader text tasks” combines the public and authored datasets in our development benchmark, reused for many decisions.

For questions with multiple correct choices, every choice must match. Results are single runs; repeating a training recipe moved scores by roughly a percentage point. Earlier experiments differ in training data, objectives, and serving software, so the table does not isolate the effect of architecture alone. Missing results are omitted, not treated as zero.

The current SelfJev uses TreeServer. Its original evaluation used different serving software with the same trained model; both are preserved in the full experiment archive.

Full methodology and paired comparisons
Explore all 56 experiments Model names, run IDs & source reports

How often did the model choose the expected answer? The same 1,991 questions test yes/no decisions, choosing one answer, and selecting all that apply.

MODEL / EXPERIMENTANSWERS MATCHED
01
Jev~typesafe/jev-latest · ~typesafe/jev-latest
97.2%
02
SelfJev-4B · original evaluationqwen35_4b_tree_scratch_jevall_ · Qwen3.5-4B
95.8%
03
qwen35_4b_tree_sft_jevall__lastqwen35_4b_tree_sft_jevall__last · Qwen3.5-4B
95.7%
04
qwen35_4b_tree_cost_jevall__lastqwen35_4b_tree_cost_jevall__last · Qwen3.5-4B
95.7%
05
SelfJevselfjev_4b_treeserver · Qwen3.5-4B
95.7%
06
qwen35_4b_tree_rlcd_fresh_qwen35_4b_tree_rlcd_fresh_ · Qwen3.5-4B
95.6%
07
qwen35_4b_tree_scratch_jevall__lastqwen35_4b_tree_scratch_jevall__last · Qwen3.5-4B
95.6%
08
Earlier SelfJevqwen35_4b_tree · Qwen3.5-4B
95.6%
09
qwen35_4b_tree_sft_fresh_qwen35_4b_tree_sft_fresh_ · Qwen3.5-4B
95.4%
10
qwen35_4b_tree_rlcd_jevall__lastqwen35_4b_tree_rlcd_jevall__last · Qwen3.5-4B
95.4%
11
selfjev_4b_v2selfjev_4b_v2 · Qwen3.5-4B
95.4%
12
selfjev_4b_reproselfjev_4b_repro · Qwen3.5-4B
95.1%
13
qwen35_4b_comboqwen35_4b_combo · Qwen3.5-4B
94.5%
14
tree_4b_combotree_4b_combo · Qwen3-4B-Instruct-2507
94.5%
15
qwen35_4b_r2x64qwen35_4b_r2x64 · Qwen3.5-4B
93.7%
16
tree_4b_instruct_r3tree_4b_instruct_r3 · Qwen3-4B-Instruct-2507
93.3%
17
tree_4b_combo_r2tree_4b_combo_r2 · Qwen3-4B-Instruct-2507
92.9%
18
tree_4b_instruct_r2x64tree_4b_instruct_r2x64 · Qwen3-4B-Instruct-2507
92.7%
19
tree_4b_combo_ptrtree_4b_combo_ptr · Qwen3-4B-Instruct-2507
92.0%
20
tree_4b_ovatree_4b_ova · Qwen3-Reranker-4B
91.6%
21
curve/tree_4b_r2b_r64_mlpcurve/tree_4b_r2b_r64_mlp · Qwen3-Reranker-4B
91.3%
22
tree_4b_ova_kdtree_4b_ova_kd · Qwen3-Reranker-4B
91.0%
23
tree_4b_instruct_r3_step300tree_4b_instruct_r3_step300 · Qwen3-4B-Instruct-2507
90.8%
24
tree_4b_r2btree_4b_r2b · Qwen3-Reranker-4B
90.6%
25
tree_4b_r2tree_4b_r2 · Qwen3-Reranker-4B
90.4%
26
curve/tree_4b_r2b_r64curve/tree_4b_r2b_r64 · Qwen3-Reranker-4B
90.3%
27
tree_4b_instructtree_4b_instruct · Qwen3-4B-Instruct-2507
88.1%
28
curve/tree_4b_r64curve/tree_4b_r64 · Qwen3-Reranker-4B
87.2%
29
Fine-tuned text rerankerlora_4b · Qwen3-Reranker-4B
86.5%
30
curve/tree_4b_mlpcurve/tree_4b_mlp · Qwen3-Reranker-4B
86.3%
31
tree_4btree_4b · Qwen3-Reranker-4B
85.1%
32
t5_round2b_2026-09-24/decoder_only_audit/t5_r2bt5_round2b_2026-09-24/decoder_only_audit/t5_r2b · google/t5gemma-2-1b-1b
76.8%
33
jina_r2bjina_r2b · jinaai/jina-reranker-v3.5
73.3%
34
t5_round2b_2026-09-24/t5_r1t5_round2b_2026-09-24/t5_r1 · google/t5gemma-2-1b-1b
73.0%
35
lora_pilotlora_pilot · Qwen3-Reranker-0.6B
68.8%

Scores are the share of answers matching the expected result; for “select all,” every choice must match. Only models with a recorded result are shown. Click a row for its report.

WHAT WE LEARNED

The route was
anything but straight.

01
Start with search models

Finding relevant text isn’t the same as deciding.

We began with models built to rank search results. Training them on decisions helped, but moving to a larger version brought little improvement. The task needed more than a bigger model.

Read the experiment
02
Try a shortcut

The question and the text need to meet.

We tried reading the text and question separately, then combining what the model learned. It reduced repeated work, but missed too much detail. How a question relates to the text matters.

Read the experiment
03
Reuse the reading

Read once. Look at it from every angle.

We kept the rich interaction between question and text, while sharing the work of reading the document. This became the foundation of SelfJev’s architecture.

Read the experiment
04
Try smaller alternatives

Smaller wasn’t always better.

We tested compact models, including Jina and an adapted T5Gemma. In those experiments, efficiency came with lower decision accuracy. These results describe our versions and training, not every use of those models.

Read the experiment
05
Make the choices clearer

Teach it to weigh the alternatives.

A stronger starting model, all possible answers shown together, and training on longer texts improved the recipe. The model could learn what makes one choice fit better than another.

Read the experiment
06
Train for the actual work

Better examples. Better decisions.

The current model learns from nearly 80,000 questions, including difficult cases and AI response review. It combines checked training answers with Jev’s probability estimates to learn how to weigh the choices.

Read the experiment

The dead ends are part of the result.

Larger models and more training data did not always help. We keep the failures alongside the wins, so the reasoning behind the current design stays visible.

What didn’t work