Field guide · 6 chunks · ~20 min
A model that answers in types, not paragraphs.
TypeSafe's Jev takes a piece of state and a set of typed questions, and hands back values your code can branch on — with probabilities attached. This guide breaks that idea into pieces, shows it in real industry JSON, and gives you a way to judge it against everything else you could use instead.
you send
state + questions
One payload. The material to judge, and every judgment you want about it.
jev evaluates
in parallel, in isolation
Each question is answered against the same state without seeing the others.
you receive
typed answers + probabilities
No prose, no parsing. Your code branches, sorts, routes.
Chunk it
Six ideas, in the order they build on each other
Open one at a time. Each ends with the thing to actually remember and a self-check you can apply to your own use case.
A normal language model is trained to produce text a person will read. When you need a judgment your software will act on — which queue, how severe, is this a refund request — you end up asking for text, then parsing it back into something your code can trust.
That round trip is where the fragility lives: prompt drift, formatting slips, retries, and no honest sense of how sure the model was.
Jev skips generation entirely. You send the material to judge and the questions to ask; you get back typed values and probability distributions.
Remember
Jev is not a chatbot. It is a decision endpoint.
Self-check
If the output ever needs to be read as prose by a person, that is a job for a generative model, not Jev.
The whole shape in one look
A request is state, questions, answers
This is the smallest complete example: one support message, three questions of three different types, one round trip. Everything else in this guide is a variation on it.
- 01Question IDs (
department) are yours — they key the response and are not shown to the model. - 02
criteriadefines the answer space. If a value is not in there, it cannot come back. - 03Score returns a fractional position (a
1.4is meaningful) plus the legend it was placed against.
{
"model": "jev-latest",
"state": "Hi, I've been trying to connect my Stripe
account for 3 days and the integration keeps
failing. I'm losing sales. Please help ASAP.",
"questions": {
"department": {
"type": "choice",
"instructions": "Which team should handle this",
"criteria": {
"billing": "Payment or subscription issues",
"technical": "Bugs or integration problems",
"sales": "Pricing or account questions"
}
},
"frustration": {
"type": "score",
"instructions": "How frustrated the customer appears",
"criteria": [
"Calm, just stating facts",
"Frustrated but civil",
"Very angry, strong language"
]
},
"is_urgent": {
"type": "noul",
"instructions": "The message conveys urgency"
}
}
}{
"answers": {
"department": { "choice": "technical", "confidence": 0.78,
"probabilities": { "technical": 0.85,
"billing": 0.15,
"sales": 0.0 } },
"frustration": { "score": 1.0, "confidence": 1.0 },
"is_urgent": { "noul": 1.0 }
},
"usage": { "input_tokens": 392, "output_tokens": 65 }
}Vocabulary
Ten words you need
- System One
- A class of model built for fast, structured decisions rather than generated text, named after Kahneman's fast, intuitive System 1 thinking. Jev is TypeSafe's first.
- State
- The content being evaluated in a request: a string, a JSON object, or an array. All questions in the request see the same state.
- Question
- One typed judgment with an ID, a type, instructions, and (for Choice and Score) criteria.
- Criteria
- The answer space: a map of options for Choice, an ordered list of levels for Score, an optional clarification of yes/no for Noul.
- Probabilities
- The distribution across the answer space. Often more useful than the winning value alone.
- Confidence
- How concentrated a Choice or Score distribution is. Not a correctness guarantee.
- Noul
- A question type returning a single 0–1 probability that a statement is true.
- Calibration
- The property that predicted probabilities match observed frequencies across groups of predictions.
- Context rot
- Degradation when many unrelated tasks share one prompt. Isolated per-question evaluation avoids it.
- Fan-out
- Sending many questions — including speculative ones — in a single request and letting code pick what is relevant.