1systemone/field-guide

Comparison board

Four ways to make the same decision

Nobody picks a decision API on accuracy alone. These are the categories that tend to decide a real project — and the honest answer is that Jev loses several of them. Tap any row to read the reasoning cell by cell.

Rules & regex

Hand-written conditions, keyword lists, lookup tables.

Your own code, a rules engine, a spreadsheet of keywords

LLM + structured output

A general text model asked to return JSON matching a schema.

GPT-class chat/Responses APIs with JSON schema, function calling

Trained classifier

A small supervised model fitted on your own labelled data.

Fine-tuned encoder, gradient boosting over embeddings, in-house ML

System One (Jev)

Typed questions evaluated against a state, returned with probabilities.

TypeSafe Jev via POST /v1/systemone

Legend

strongdependsweak
CategoryRules & regexLLM + structured outputTrained classifierSystem One (Jev)

Output guarantee

Can my code consume the result without defensive parsing?

strongdependsstrongstrong

Calibrated uncertainty

Does it tell me honestly how sure it is?

weakweakdependsstrong

Semantic understanding

Can it handle phrasing nobody anticipated?

weakstrongdependsstrong

Cold start

What do I need before I get a first useful answer?

dependsstrongweakstrong

Many judgments at once

What happens when I need twelve decisions about the same input?

strongdependsdependsstrong

Latency profile

Can it sit in a synchronous request path?

strongweakstrongstrong

Changing the policy

What does it cost to shift priorities next quarter?

dependsweakweakstrong

Explanation & audit

Can I tell a regulator or a customer why this happened?

strongdependsdependsdepends

Writing text for people

Can it draft the reply, the summary, the redline?

weakstrongweakweak

Cost shape

How does spend scale with volume?

strongweakdependsdepends

Input modalities

What can I feed it?

weakstrongdependsweak

Output guarantee

Can my code consume the result without defensive parsing?

Rules & regex

strong

You wrote the output type. It is whatever you returned.

LLM + structured output

depends

Schema modes make the shape reliable, but the values inside are still generated text and can be plausible-but-wrong.

Trained classifier

strong

Fixed label set by construction.

System One (Jev)

strong

Typed values from a closed answer space you defined. No parsing step exists.

Reading the board

The realistic answer is a stack, not a winner

Layer 1

Rules

Every hard constraint that must never be probabilistic: limits, dates, entitlements, regulatory non-negotiables. Cheap, instant, auditable.

Layer 2

Typed judgments

The routine reading-comprehension calls, batched into one request, with confidence deciding what gets handled automatically.

Layer 3

Reasoning model or human

The low-confidence tail, plus anything that needs prose written or a genuinely novel call made. Measure what percentage lands here — that number is your business case.

Honest caveats about this board

  • — Verdicts are structural, not benchmarked. They describe how each approach behaves by design, not how it scores on your data.
  • — "Weak" is not disqualifying. Jev is weak at writing text because it was built not to; that only matters if your product needs text.
  • — Cost and latency move constantly. Treat those two rows as shape, and measure the actual numbers yourself.
  • — The next page turns this into something you can run: a step-by-step evaluation you can execute in a week.