Read. Understand. Choose.
Can it answer questions about a piece of text, including exceptions and several possible answers?
1,991 questions · yes/no, pick one, select allWhat can it do? How well does it work?
The results, and the experiments behind them.
Can it answer questions about a piece of text, including exceptions and several possible answers?
1,991 questions · yes/no, pick one, select allCan it judge answer quality, check supporting evidence, apply a policy, or spot an attempt to bypass instructions?
946 questions · quality, accuracy & safetyCan it recognize intent, topics, sentiment, and other text patterns? These tasks helped guide development.
3,471 questions · development progress onlyRun the evaluation yourself. Decision Bench publishes Text Decisions, AI Response Review, and a separate 720-question Record Reasoning suite. The broader development benchmark above is not included.
Dataset, scoring tools & methodologyThe complete merged weights, with tokenizer and model configuration.
Get the model ↗LORA ADAPTERThe trained adapter used by the native tree engine, with its model card.
Get the adapter ↗EVALUATION DATASET3,657 questions across text decisions, AI response review, and record reasoning.
Explore the dataset ↗How often did the model choose the expected answer? The same 1,991 questions test yes/no decisions, choosing one answer, and selecting all that apply.
Scores are the share of answers matching the expected result; for “select all,” every choice must match. Only models with a recorded result are shown. Click a row for its report.
Each score is the percentage of answers that match the expected result on that test. Jev still leads on text decisions. SelfJev brings comparable results on these questions to a model you can run yourself.
Our two focused tests were written and checked by AI, whose answers can still be wrong. They help us assess progress; they do not guarantee the same accuracy on your data. We have not tested every model on Hugging Face.
“Text decisions” is the project’s eval2 test; “AI response review” is eval_llm. Both use fixed questions kept out of training, with expected answers checked by two AI judges. They have informed the research direction, so a fresh final test is still needed. “Broader text tasks” combines the public and authored datasets in our development benchmark, reused for many decisions.
For questions with multiple correct choices, every choice must match. Results are single runs; repeating a training recipe moved scores by roughly a percentage point. Earlier experiments differ in training data, objectives, and serving software, so the table does not isolate the effect of architecture alone. Missing results are omitted, not treated as zero.
The current SelfJev uses TreeServer. Its original evaluation used different serving software with the same trained model; both are preserved in the full experiment archive.
Full methodology and paired comparisonsHow often did the model choose the expected answer? The same 1,991 questions test yes/no decisions, choosing one answer, and selecting all that apply.
Scores are the share of answers matching the expected result; for “select all,” every choice must match. Only models with a recorded result are shown. Click a row for its report.
We began with models built to rank search results. Training them on decisions helped, but moving to a larger version brought little improvement. The task needed more than a bigger model.
Read the experimentWe tried reading the text and question separately, then combining what the model learned. It reduced repeated work, but missed too much detail. How a question relates to the text matters.
Read the experimentWe kept the rich interaction between question and text, while sharing the work of reading the document. This became the foundation of SelfJev’s architecture.
Read the experimentWe tested compact models, including Jina and an adapted T5Gemma. In those experiments, efficiency came with lower decision accuracy. These results describe our versions and training, not every use of those models.
Read the experimentA stronger starting model, all possible answers shown together, and training on longer texts improved the recipe. The model could learn what makes one choice fit better than another.
Read the experimentThe current model learns from nearly 80,000 questions, including difficult cases and AI response review. It combines checked training answers with Jev’s probability estimates to learn how to weigh the choices.
Read the experimentLarger models and more training data did not always help. We keep the failures alongside the wins, so the reasoning behind the current design stays visible.