1systemone/field-guide

Do it yourself

How to actually run this comparison

A vendor board tells you how things differ in principle. This turns it into a week of work that produces a defensible answer for your own project — including the case where the answer is 'none of them, keep the rules'.

  1. 01

    Write down the decision, not the technology

    Start from what the software will do differently: which queue, which handler, approve or hold, show or hide. If nothing in the product changes based on the output, there is no decision to automate and no comparison to run.

    output → One sentence per decision, plus the action each outcome triggers.

  2. 02

    Split it into code-work and judgment-work

    Amount limits, date windows, entitlement checks, lookups — those stay in code forever, whichever vendor you pick. Circle only the parts that need reading comprehension. That circle is what you are actually shopping for.

    output → Two columns: deterministic rules, and semantic judgments.

  3. 03

    Build a gold set before you call anything

    Collect 100–300 real inputs and label the correct outcome by hand, including the awkward ones. Reserve a slice you never look at while tuning. Without this, every comparison is vibes.

    output → A labelled CSV with a held-out split.

  4. 04

    Define the categories that matter to you

    Accuracy is never the only axis. Pick the four or five that would actually decide the purchase: p95 latency, cost per thousand decisions, effort to change policy, auditability, modality support, data residency.

    output → A scoring sheet with weights agreed before results come in.

  5. 05

    Run every candidate on identical inputs

    Same gold set, same state, same day. For each approach record the outcome, the reported uncertainty, the wall-clock latency, and the token or compute spend. Log the raw responses so you can re-score later without re-running.

    output → One results table per approach, joined on input ID.

  6. 06

    Score for the cost of being wrong, not for accuracy

    Separate the error types. A missed emergency, a wrongly declined merchant, and a mis-tagged newsletter are not equivalent. Weight the confusion matrix by real consequence, then find the confidence threshold where automation still beats the human baseline.

    output → Cost-weighted error per approach and a chosen operating threshold.

  7. 07

    Test the change you know is coming

    Pick a plausible policy shift — a new prohibited category, a re-weighted priority — and implement it in each candidate. Time it. This is usually where the approaches separate more than accuracy does.

    output → Hours-to-change, measured rather than estimated.

  8. 08

    Decide the hybrid, not the winner

    The realistic answer is usually layered: rules for the hard constraints, a fast typed model for the routine judgments, a reasoning model or a person for the low-confidence tail. Write down which layer owns what and what the escalation rates are.

    output → An architecture diagram with a measured escalation rate per branch.

Step 05, concretely

One results row per input, per approach

Keep the raw response alongside the parsed outcome. Six weeks later, when someone asks whether a different threshold would have helped, you want to answer from the log rather than re-running and re-paying for everything.

Join everything on a stable input ID so the same case can be compared across all four approaches, including the human baseline.

results schema
{
  "input_id": "case_0417",
  "approach": "jev" | "llm_json" | "classifier" | "rules",
  "gold_label": "unauthorised",
  "predicted": "unauthorised",
  "uncertainty": 0.91,          // confidence, noul, or softmax max
  "latency_ms": 310,
  "unit_cost_usd": 0.0009,
  "raw_response": { /* verbatim */ },
  "run_at": "2026-09-30T14:02:11Z"
}

// scoring, later and offline:
//   cost_weighted_error = Σ error_cost[gold][predicted]
//   automation_rate     = share above your threshold
//   escalation_quality  = accuracy of the escalated tail

Before you commit

Questions worth asking any vendor

Are the probabilities calibrated, and measured how?

Ask for the reliability curve, not a single accuracy figure. Calibration is a property of groups of predictions.

What is p95 latency at my payload size?

Averages hide the cases that break a synchronous request path. Test with your real state, not a one-line string.

What happens to cost as I add judgments?

Batched-in-one-request and one-call-per-judgment have very different curves at a million decisions a month.

Where does my data go, and for how long?

Residency and retention usually decide regulated projects before accuracy ever enters the conversation.

How do I version a question?

Changing criteria changes meaning. You need a way to pin, diff and roll back question definitions.

What is the migration cost if this is wrong?

Keeping the composition logic in your own code — the whole point of the pattern — is what makes swapping the model survivable.