selfjev/GitHub
GET STARTED / SELFJEV

Hardware & sizing

GPU, CPU, memory, and what is actually tested.

Use an NVIDIA GPU for the supported server. selfjev serve selects CUDA by default and does not offer a CPU or MPS device flag. A full-model Apple Silicon experiment is measured below through the lower-level tree engine; it is not a validated Mac serving setup.

A practical starting point

WorkloadGPU / VRAMSystem RAMCPUStatus
Light inferenceNVIDIA, 24 GB16–32 GB4 vCPUsPlanning recommendation; A10G engine correctness checked
Smaller-memory experimentNVIDIA, 16 GB16–32 GB4+ vCPUsDocumented floor; workload-dependent, not a measured minimum
Long inputs or more concurrencyNVIDIA, 48 GB32 GB+8 vCPUsConservative starting point; measure your workload
Fine-tuningL40S, 48 GB64 GB used on g6e.2xlarge8 vCPUsTraining recipe run on this class of machine
CPU-onlyNone32 GB budget for an experimentModern multicore CPUUnsupported CLI path; no measured latency or minimum
Marketing and docs siteNoneNo dedicated runtime requiredNone at runtimeStatic Next.js export

System RAM and GPU VRAM are separate resources. Extra system RAM does not automatically compensate for too little VRAM in this engine.

Compare measured GPU response times

The interactive hardware comparison shows server-side median request times on A10G, L40S, and H100 for short through long inputs and either 1 or 16 questions about the same text. Each result is the median of 10 warmed requests, with three answer options per question and network time excluded. The GPU runs use vLLM; the Mac run uses the native tree engine on MPS. These are measured configurations, not a hardware-only comparison. The chart links each GPU to its raw request log; full model and runtime configurations are recorded in the benchmark methodology.

Apple Silicon experiment

We also ran the current SelfJev-4B on a local M5 Pro with a 20-core GPU and 48 GB of unified memory. The adapter was merged into Qwen3.5-4B in bf16 and scored through TreeServer on PyTorch MPS. The test used synthetic repeated text and three answer options per question. These are local call times, including tokenization but excluding HTTP and network time; each cell is the median of 10 calls after two warm-ups.

Text length1 question16 questions
8 tokens688 ms6,786 ms
512 tokens2,018 ms8,599 ms

The MPS path uses PyTorch's slower reference implementation for the model's gated recurrent operation. The raw samples and benchmark script make this small experiment reproducible. The supported CLI still requires CUDA.

Why not an 8 GB machine?

A nominal 4-billion-parameter model needs approximately 8 GB for bf16 weights alone, or 16 GB for float32. That excludes the adapter, temporary loading copies, activations, attention masks, recurrent state, and the Python runtime. The downloaded base checkpoint is about 9 GB.

A 16 GB RAM CPU machine is not a supported minimum. A 32 GB budget gives more room for experimentation, but long inputs can still exceed it and the CPU speed is unknown. No quantized CPU artifact or llama.cpp/GGUF serving path is provided. Supporting one is engineering work, not a configuration switch.

Context length and concurrency matter

SelfJev's context limit is configurable with --max-length; the default is 32,768 tokens for the formatted document plus the longest question/candidate path. Prompt formatting and options also consume this budget.

The underlying Qwen3.5-4B configuration has a native 262,144-token context window. That is base-model capacity, not a validated SelfJev serving limit: training used texts up to 16K, and we have not validated full-window inference or accuracy in this engine. Raising the flag alone does not establish usable capacity.

For comparison, Jev's published limits are 32K tokens for state plus the longest question, and 64K tokens for state plus all questions combined (checked September 28, 2026). Self-hosting lets you experiment with a larger budget, subject to memory and validation on your documents.

Memory use grows with document length, the number of question/candidate branches, and batching. The current tree builds a dense attention mask; long packed sequences can be expensive even when the weights fit.

Start with short inputs and conservative batching. The controls are:

bash
uv run --no-sync selfjev serve \
  --max-length 4096 \
  --max-batch-tokens 4096

--max-length limits the state plus the longest question/candidate path. --max-batch-tokens controls packing; it is not a hard memory cap for a single large request tree. Reduce question and candidate counts as well if you run out of memory. Longer-than-allowed input is rejected, not silently truncated.

Disk and first startup

Budget 50 GB of free disk for the checkout, Python/CUDA dependencies, adapter, model cache, and temporary files. This is a planning allowance, not a measured minimum. Building a Docker image with baked-in weights may require additional space.

What has actually been measured?

The hardware explorer includes latency runs on A10G, L40S, H100, and M5 Pro. The native tree engine also has an A10G correctness check, and the training recipe ran on an L40S. Mac serving and full-model CPU inference remain unvalidated. See the linked reports for the configuration and scope of each measurement.

Sources: deployment notes, engine loader, tree packing, GPU engine check, and latency methodology.