selfjev/GitHub
THE SELF-HOSTED DECISION MODEL v0.2

Intelligence,
decided.

Turn context into decisions.
A small AI model that reads once, answers many,
and runs on infrastructure you control.

Weights available on Hugging Face
Runs on one GPUNo text generationJev-compatible API
ONE TEXT. MANY ANSWERS.01 — 03
YOUR TEXTREAD ONCE

“I was charged twice for my subscription. Please refund the duplicate payment.”

ask your questions
Q01 yes / no

Needs a refund?

yes
Q02 pick one

Which team?

billing
Q03 pick many

Which topics?

payments, refund
Answers your code can useNo generated text
Example decisions · not a live model response
TEXT DECISIONS95.7%Matches expected answers in our tests AI RESPONSE REVIEW93.1%Checks quality, accuracy & safety READ YOUR TEXT ONCE1 readAcross every question in a request GENERATED TEXT0Answers your code can use

Accuracy on our project’s test questions, not a guarantee for every use case. What did we test? ↗

01 / BUILT TO DECIDE

Some things need an answer.
Not another conversation.

Route a request. Check an agent’s work. Apply a policy. SelfJev turns your text and questions into structured decisions with probabilities—without waiting for a generated response.

Four answer types.

Yes/no, pick one, rate a result, or select all that apply. Ready to use in your application.

Your data stays yours.

Run the model on your own GPU. Process sensitive text without sending it to an external model API.

A familiar interface.

Already using Jev? Keep the same request format and point your application to your own server.

02 / HOW IT WORKS

One context.
Every angle.

A support message can raise several questions: what happened, who should handle it, and what to do next. SelfJev reads the message once and reuses that work for every answer.

SELFJEV / SHARED-PREFIX TREE

Read once. Decide many.

SelfJev shared-prefix treeOne customer message: I was charged twice. Please refund the duplicate payment. The shared document feeds three different question types. Binary: Refund requested? One Yes candidate is scored with a yes/no readout, returning true. Multiclass: Which team? Billing and Support compete; Billing is selected. Multilabel: Which topics? Payments, Refund and Login are scored independently; Payments and Refund are selected. These are illustrative answers, not live model predictions. Moving light illustrates shared computation and isolated candidate paths.One customer message“I was charged twice. Please refundthe duplicate payment.”Read once. Reused by all three questions.BINARY · YES / NORefund requested?One yes / no decisionYesyes / no readoutYesA boolean for your codeMULTICLASS · PICK ONEWhich team?Billing or SupportBilling✓ selectedSupportnot selectedBillingOne team selectedMULTILABEL · PICK MANYWhich topics?Payments, Refund, LoginPayments✓ selectedRefund✓ selectedLoginnot selectedPayments + RefundEvery matching topic selectedThree answers. Ready for your code.refund: trueteam: "billing"topics: ["payments", "refund"]One shared message.Different answer types.Probabilities alsoreturned by the API.
Illustrative answers · not a live model runSwipe to explore →Inside the engine

01. Give it contextA message, document, or AI response, together with the questions you need answered.

02. Ask several questionsEach question uses the same text, so the model avoids reading it from scratch each time.

03. Act on the answersGet clear choices and probabilities to route, filter, or review in your own application.

YOUR MODEL / YOUR REQUIREMENTS

Make room.
Make it yours.

Your documents, your vocabulary, your edge cases. Control the context budget and adapt the model to the decisions that matter to you.

ROOM FOR LONGER CONTEXT

A bigger foundation.
A limit you control.

Jev allows 32K tokens for your text plus the longest question. SelfJev has a configurable limit, built on a model with a 262,144-token native context window.

Jev · text + longest question32K
SelfJev · base-model capacity262K

SelfJev defaults to 32,768 tokens; larger windows need sufficient memory and validation. The full base-model window has not been validated in our engine. Jev also allows 64K across a whole request.

Context & sizing
BUILT TO ADAPT

Your use case.
Your fine-tune.

Not getting the decisions you need? Fine-tune SelfJev on examples from your own workflow: your labels, your policies, your definition of a good answer.

  1. 01
    Show it what good looks like.Pair real inputs with verified answers.
  2. 02
    Train a lightweight adapter.Use the included supervised fine-tuning tools.
  3. 03
    Evaluate. Then deploy.Check held-out examples and serve your model.
Fine-tune for your use case
03 / THE EVIDENCE

Strong results.
Receipts included.

Can it choose the right answer from a piece of text? We tested yes/no decisions, choosing between options, and selecting every answer that applies.

How often did the model choose the expected answer? The same 1,991 questions test yes/no decisions, choosing one answer, and selecting all that apply.

MODELANSWERS MATCHED
01
JevTypeSafe’s hosted decision model
97.2%
02
SelfJevOur current model · self-hosted
95.7%

Scores are the share of answers matching the expected result; for “select all,” every choice must match. Only models with a recorded result are shown. Click a row for its report.

SelfJev approaches Jev’s accuracy on these tests and runs on your own infrastructure. The questions and expected answers were created and checked by AI; results on your own data may differ.

How we tested it
↳

The breakthrough was learning the right task.

Handling exceptions, weighing alternatives, and checking AI responses all needed targeted training. Improving those examples mattered more than simply making the model bigger.

Read the findings
04 / THE SPEED STORY

Less repeated work.
More decisions.

From a local Mac to a dedicated GPU. Explore measured response times, adjust your workload, and see what fits your infrastructure.

22ms

H100 · measured processing time
Very short input · one question

How we measured speed
SELF-HOSTING / HARDWARE LATENCY

Choose the machine.

Typical processing time for one request. Lower is faster.

MEASURED LATENCY
Input length · 512 text tokens
Questions about the same text
A10G24 GB GPU memory
136ms
L40S48 GB GPU memory
55ms
H10080 GB GPU memory
30ms
M5 Pro48 GB unified · MPS

Local Apple Silicon

2,018ms
AVAILABLE ON HUGGING FACE

Get the weights. Explore the evidence.

Jwuthrich
05 / YOUR INFRASTRUCTURE

From a question
to your first call.

A lightweight Python SDK. One GPU to serve. Docker and an AWS deployment command, with practical guides for Runpod and Google Cloud.

Python / your first decision
from selfjev import SelfJev, Noul, Choice

# Use the secret you set as SELFJEV_API_KEYS on your server.

client = SelfJev(
    base_url="http://localhost:8000",
    api_key="your-server-key",
)

result = client.system_one(
    state="I was charged twice. Please refund me.",
    questions={
        "refund": Noul("Does the customer want a refund?"),
        "team": Choice("Which team should handle this?", {
            "billing": "payments and refunds",
            "support": "technical support",
        }),
    },
)
THE MODEL IS SMALL. THE NOTEBOOK IS OPEN.

Make the next decision
on your terms.

Self-host SelfJev