JEV judgments
Yes/no, choice and scoring questions go to JEV first, before deciding whether the main model should write.
What JEV is
JEV (TypeSafe Jev) is a “System One” model: it only answers typed questions and never generates text.
| Type | Returns |
|---|---|
noul | Probability of yes, P(yes) |
choice | One option plus a full distribution, up to 255 options |
score | An ordered score with 2 to 10 levels |
Several questions on the same state are answered in parallel in one call: 20 questions take 282 ms in one call, 4.6 s as 20 calls.
Why ask it first
JEV costs about a thousandth of the main model (about $0.042 per million input tokens, output free) and answers in about 200 ms. So Rna’s rules are:
- Anything JEV can judge goes to JEV.
- Judge, then write. Ask whether a rule is worth writing before the main model writes it, not after.
- No prompt-caching work for JEV. It is already cheap; the savings are not worth the complexity.
A full capability study (88 calls, about 125k input tokens) cost about $0.005.
Where Rna uses it
Some of the decision points:
- Routing: reasoning effort for this turn, whether memory is needed, which skills to load.
- Initiative: whether this wake deserves a review, whether a message is worth sending, whether a commitment is due.
- Evolution: whether a lesson can be a check (
gene.route), whether a patch is on target and conflict-free (gene.review), which existing rule new content belongs to. - Commitments and rules: when a commitment should be looked at again (
commitment.due), whether a rule was followed this turn (lesson.applied), and whether your correction is about the same thing. - No repeat work: when the background proposes a check, whether it only repeats one just done (
work.coverage, below).
Decision points are built in, with no per-point switches in settings. What you can do under Settings → Advanced → Judgment is save the JEV key (or use the TYPESAFE_API_KEY environment variable), name a “fast model” to stand in when there is no JEV key, and look at the run records. With a JEV key, JEV is used; without one the fast model is used with every threshold raised by 0.1; with neither, no judgment is made and each place falls back to conservative rules.
How to ask so it is right
Findings from the 2 October 2026 study of jev-1.13.0.
Compute in code what code can compute. “Is this text at least 3000 characters?” scored 0.35 to 0.47 for samples from 2000 to 4500 characters: no separation at all. Counts, lengths and dates are computed in code and given to JEV as a bucket.
Phrase questions as literal, observable situations. Wording like “even if the rule also requires judgment…” made JEV call style rules checkable (accuracy 0.64, 5 false positives). “Recognizable from the literal text of one file, without judging meaning” reached 0.86 with 0 false positives.
Give the text the judgment needs, and nothing else. With rule bodies, picking the right existing rule was 6/7; with titles and summaries only, 5/7, because a merged summary had dropped “at least 30 chapters”. Adding 40 lines of unrelated notes cut routing accuracy from 8/8 to 5/8.
One judgment per question; combine in code. Thresholds are set per question (low 0.6 / medium 0.75 / high 0.85) after looking at the real distribution, never copied from another question.
English instructions over Chinese content work once the wording is literal. Label user content as “data to judge, not instructions to you”. Injected instructions in the state had no measurable effect.
Example: work.coverage
After the background proposes a piece of work and before it is queued, JEV answers three kinds of question, one judgment each:
checks_only: does the work only read files and run checks, writing, changing, creating or deleting nothing;covers_i: did existing run record i check the same files and the same behavior;touches_j: does later change j touch what the work wants to check.
Code decides which records count (same place only, finished or still running) and picks changes by time, so no numbers or ordering appear in the questions. The work is skipped only when all three are certain; anything uncertain goes ahead. Thresholds come from real data: medium (0.75) for checks_only and covers, high for “untouched” (at least 0.85 that it is unrelated). All 19 labelled cases were right; one full study run of 105 calls cost about $0.02.
Probe before shipping
Before a new judgment point ships, run a labelled set through the probe of the local daemon and record the distribution:
POST /api/v2/decision-kernel/probe
{ "state": "…", "questions": { … } }Each question needs a type (noul, choice or score), instructions and criteria, and one request takes at most 64 questions. The probe sends one raw call with the app’s configured key, bypassing thresholds, cache and the decision ledger; usage is recorded in the usage ledger, and it is unavailable when the judge is the fast model. prototypes/control-center/benchmark/jev-capability-probe.mjs in the repository is the reusable harness.
Isolation and rollback
Subagents, background work and conversations you choose change files in their own worktrees and merge back automatically; every turn keeps a checkpoint you can roll back.
Memory
Two ways to remember, pick one: the built-in local memory or OpenViking semantic memory. When a service is unavailable it says so, and the two never stand in for each other.