Skip to content

Architecture

Context

How Jev works: the System One thesis (judgment ≠ generation), the RLCD training method behind calibrated confidence, the three question primitives and their exact response shapes, the parallel-in-isolation sampling design, and what independent evaluations do and don't confirm.

The System One Thesis

TypeSafe's founder (Diogo Almeida, previously at OpenAI on the instruction-following research behind ChatGPT) frames generation and judgment as different workloads. LLMs are trained on North Stars, and each North Star bends the model differently:

Training paradigm North Star Pathology
RLHF Human preference on generated text Mode collapse/mode drop; miscalibration "poisons" the probability space (confident style wins rewards)
RLVR Programmatically verifiable outputs Superhuman on benchmarks, "fractal" jaggedness elsewhere
RLCD (Reinforcement Learning for Calibrated Decisions) Calibrated decisions for programmatic use — answers whose confidence tracks observed accuracy New; the claim is epistemically honest probabilities on System One tasks

Jev gives up string generation entirely. The output space is fixed at request time, which is what makes the type-safety guarantee structural rather than empirical: a response is a distribution over options you declared, so "a malformed value" is not a representable outcome — TypeSafe calls the 0% type-error rate "mathematically impossible" to falsify. The sharp edge of that claim: the model can still choose a wrong valid option with high confidence. Schema safety ≠ judgment quality.

Request/Response Model

Named after Kahneman's Thinking, Fast and Slow (System 1 = fast intuition; the model class is the intuition layer under System 2 reasoning models); "Jev" is William Stanley Jevons — the Jevons paradox: cheaper intelligence unlocks proportionally more demand.

sequenceDiagram
    autonumber
    participant App as Your code
    participant API as api.typesafe.ai/v1/systemone
    participant M as Jev (jev-1.13.0)

    App->>API: POST { state, questions }
    Note over API: state: string / JSON / JSON array of text<br/>questions: { id: { type, instructions, criteria } }
    API->>M: evaluate ALL questions in parallel,<br/>each in isolation against the same state
    M-->>API: per-question typed answers<br/>+ probabilities + confidence
    API-->>App: one round trip, 70-500ms
    Note over App: code branches on values + thresholds:<br/>act / escalate / abstain

Key semantics (from docs + independent verification):

  • One round trip, N questions. All questions ride in a single call; TypeSafe reports response time barely moves as questions are added.
  • Parallel AND in isolation. Every question is evaluated independently against the same state — one answer never becomes context for another, so there is no context-rot and no cross-question contamination. Independent tests found no batching effect beyond sampling noise.
  • Question IDs never reach the model. The id keys your code branches on are local; the model sees only instructions (and criteria for choice/score). Write the full question in instructions.
  • Code owns the branch. Jev returns estimates; your program compares against thresholds and decides to act, escalate, or abstain.

The Three Primitives

Primitive Question shape Returns Constraints
choice Which of these N options? winning key, per-option probabilities, confidence up to 255 options; high-cardinality choices use a 2-stage internal system (score options independently, then an explicit choice) — occasionally slower
score Where on this ordered rubric? fractional score, level legend, full probability distribution, confidence 2-10 levels
noul Is this true? probability 0.0-1.0 TypeSafe's Boolean decision type; criteria optional

A single call freely mixes all three. The response example from the DDDS walkthrough shows the contract:

{
  "choice": "billing",
  "probabilities": { "billing": 0.52, "technical": 0.46, "sales": 0.02 },
  "confidence": 0.18
}

The winner was 6 points ahead — the distribution, not the label, is the automation signal. Systems need an abstention policy alongside the label: act automatically when confidence is high and consequence is small; confirm or escalate to a stronger model when middling; route to a human when low. Thresholds belong in code and should scale with consequence.

Type Safety and Calibration

Two orthogonal guarantees, often conflated:

  1. Schema safety (absolute): the output cannot violate the declared types — a consequence of the constrained output space, not of intelligence.
  2. Calibration (empirical, per-deployment): predicted confidence should track observed accuracy because of RLCD training — but that relationship must be measured on your traffic, per decision type, and re-checked after any change to model version, questions, or input distribution.

Internal Architecture

Not published in detail. What TypeSafe discloses: a new model architecture (not a downsized LLM — the FAQ explicitly rebuts "is Jev just a smaller LLM"), a parallel sampler that emits all outputs in one hardware-aware query instead of sequential token generation, and RLCD training. The efficiency story is structural: no autoregressive decoding means output tokens cost nothing to meter and latency does not scale with the number of questions.

Component Breakdown

Component Role Notes
state input The context under judgment string, JSON object, or JSON array of text; no multimodal input
questions map Typed decisions to evaluate { type, instructions, criteria }; atomic questions work best — decompose multi-factor judgments and combine in code
Parallel sampler One-pass evaluation of all questions replaces autoregressive decoding; 70-500ms end-to-end
RLCD training stack Calibrated confidence the differentiator vs RLHF/RLVR models
Versioned model IDs jev-1.13.0, jev-latest, jev-preview responses report the answering version — log it, pin thresholds to it
Token budgets ~64K shared across state + questions; ~32K for state + longest question (~150K chars English) irrelevant context measurably hurts accuracy — send minimal state
Workflow evals (evals.typesafe.ai) TypeSafe's benchmark: fixed compute graphs, reference = average of GPT-6 Astra + Fable 5.1 source of the 193.6x faster / 444.6x cheaper headline claims

Benchmarks and Evidence

TypeSafe's workflow evals (self-reported)

New evaluation type: a correct compute graph is assumed (the "workflow" is in code), and models are scored against the average predictions of the largest external frontier models (GPT-6 Astra, Fable 5.1) using the same workflow. Jev claims the Pareto frontier for almost 2 orders of magnitude on cost/accuracy. Disclosed caveats: workflows built by TypeSafe's own model-capabilities team (bias possible, though not in the training distribution); reference answers bias toward OpenAI/Anthropic (undercounting DeepSeek and Jev alike); headline multipliers are "on the higher end of real world gains"; speed measured from West Coast laptops; pricing sustainability unproven ("can't prove it isn't subsidized"). LLM baselines run through TypeSafe's own structured-decision wrapper (system-one-adapter-python), which the company argues is the most accurate way to get decisions from LLMs but is slower and costlier.

Independent evaluation: SREGym-Lite (September 2026)

SREGym (MIT-licensed SRE incident benchmark) integrated Jev as decision support for a Codex-harness agent on gpt-5.6-luna: jev_plan ranks 3-5 competing diagnostic hypotheses via choice+score questions; jev_submit reviews evidence before diagnosis/mitigation submission, with every required question needing probability >= 0.70 and rejected submissions forcing new evidence gathering.

  • Result: 24/50 vs 20/50 attempts (40% -> 48%) across 10 problems x 5 attempts; internal-traffic-policy went 0/5 -> 3/5; two problems regressed.
  • Where it helped: distinguishing causal mechanism from believable noise (OpenTelemetry errors, unrelated workloads) — Jev added "useful friction before premature diagnosis".
  • Where it failed: accepted evidence of current functionality without testing the invariant that makes a repair durable (e.g., fixing a pod but leaving maxUnavailable: 100%); could not rescue a missing hypothesis (if the right test never enters the candidate set, ranking cannot recover it).
  • Caveats: small sample, single source, self-published by the benchmark's authors; pass-rate only (time-to-diagnosis untested).

Ecosystem signal

Beacon (Asymptote Labs, MIT, 1.1k+ stars) uses Jev to evaluate agent traces across 19+ coding-agent harnesses and extract reusable skills — evidence the latency/cost envelope enables "judge every run" workloads that were previously impractical.

What Jev Cannot Do (per TypeSafe's own jaggedness disclosures)

  • No generation of any kind; no reasoning chains, no explanations.
  • No arithmetic, counting, or date comparison; no precise string manipulation (cannot compare #FF4B0A to another hex color).
  • Literal reading: it interprets your words, not your intent — ambiguous questions get literal answers.
  • No unknown-value extraction: it chooses among supplied candidates only.
  • Text-only state; multimodal is future work.
  • Calibration numbers are self-reported; independent replication is thin as of September 2026.

Response Shape Reference

Exact fields per primitive (from docs and the DDDS walkthrough) — the contract your code branches on:

Primitive Fields Notes
noul noul (0.0-1.0) probability the statement is true; nothing else
choice choice, probabilities (per option), confidence confidence summarizes how close the race was — a 0.52 winner with 0.46 runner-up yields low confidence
score score (fractional), legend (levels), probabilities (per level), confidence fractional scores come from the level distribution

Worked noul example from the DDDS ticket-routing walkthrough — urgent returns {"noul": 0.92} while owner returns the choice shape above; the calling code pages on-call only when urgent > 0.9 AND the winning team matches, and sends to human review when confidence < 0.6. The response also carries the model version ID that answered it — treat it as part of the audit record.

Latency and Cost Anatomy

Why a fixed output space is fast: autoregressive generation pays per output token sequentially; Jev's parallel sampler emits all answers in one hardware-aware pass, so latency is dominated by one forward pass over the state — not by question count or answer length. Consequences:

  • Latency range 70-500ms is roughly flat in the number of questions (docs: "adding questions barely changes the response time").
  • Output is unmetered ("too cheap to meter") — the bill scales with state tokens only.
  • Cost per decision is dominated by engineering costs: rubric design, shadow evaluation, threshold maintenance. TypeSafe's own guidance is to measure end-to-end cost per decision including escalations and false outcomes.
  • The exception to flat latency: very high-cardinality choices (up to 255 options) use a 2-stage internal process (independent option scoring, then explicit choice), which is visibly slower.

The Open Counterfactual: AnyJev (Nokia Applied Research, September 2026)

AnyJev (Apache-2.0, Nokia + Tencent Hunyuan) reproduces the Jev interface over any open LLM by reading decisions straight off next-token log-probabilities — no generation, no fine-tuning — and fixing the two biases that make raw logit readouts unusable:

  1. Cyclic-shift marginalization — a K-option list is shown in K rotations so every option sits at every position once, combined in log space (removes position bias).
  2. Prior correction — the model's label prior is estimated label-free (same prompt, content replaced by N/A) and divided out. Example from the README: a spam noul reads P(Yes) = 0.62 on content but 0.70 on N/A — the model leans Yes regardless — and dividing out the prior flips the judgment to 0.41.

Levels: raw (restricted softmax over label tokens — what the simple clones do), L0 (label-free debiasing; not calibrated), L1 (temperature scaling on top of L0, needing 100-500 labels per question; calibrated within its distribution). Every decision carries its level so downstream code can refuse the wrong one.

Measured, Qwen3-8B on BANKING77 20-way (300 items): order-flip rate 0.230 -> 0.073, accuracy 0.747 -> 0.803, ECE 0.240 -> 0.095 (L1), and the operationally decisive row — auto-decidable share at <=5% error: 7.7% raw -> 52.0% L1.

Independent Jev numbers. AnyJev's bench includes Jev 1.13.0 as measured by a third party (Laya's typed-decisions set, 2,000 decisions): accuracy 0.727, ECE 0.144, Brier 0.148. On the same set, Qwen3-32B + AnyJev L1 reaches 0.699 accuracy (2.8 points behind) with ECE 0.036 (4x better calibrated); the fine-tuned Laya checkpoint wins argmax accuracy (0.768) but its ECE is 6x AnyJev L1's. First-party caveat: the benchmark is AnyJev's own; the fine-tuned-Laya row reproduces its published number, and the Jev row is quoted from its authors, not rerun.

Honest limits from AnyJev's own README: L0 is not a free win everywhere (it lowers 5%-risk coverage on one prompt-injection split); the batch prior needs >= 8 items and hurts when the true majority label exceeds ~65%; calibration makes uncertainty legible, not smaller (on Minesweeper no readout beats random); 26-option cap in the letter readout; only Qwen rows measured so far; L1 does not survive distribution shift.

Source Discrepancies

  • Pricing: the official figure is $0.042/MTok input (TypeSafe blog, Vercel gateway listing, flaviocopes). The unofficial reseller site jevtypesafeai.com shows $0.25-$0.42/M — reseller markup, not official pricing. Official = typesafe.ai.
  • Headline multipliers: "up to 200x lower latency / 400x lower cost" (product page) vs 193.6x/444.6x (workflow evals derivation) vs "40x-200x faster" (announcement table) — all self-reported, different comparison points; treat as upper bounds.
  • "Cannot hallucinate": TypeSafe means schema impossibility; the DDDS walkthrough correctly narrows it to "cannot break the declared output schema — it can still choose the wrong valid option with high confidence."

Sources