1systemone/field-guide

Applied · six industries

The same shape, in problems you recognise

Each scenario walks the full round trip: what goes in as state, which questions get asked, what comes back, and — the part that matters most — the code that turns answers into a decision. The answer values are illustrative, written to show realistic distributions rather than claimed model output.

Fintech

Card dispute triage

A cardholder files a dispute. Someone has to decide within minutes whether it is a likely-fraud case, a merchant service failure, or a duplicate charge — and whether to provisionally credit the account before investigating.

Why a typed judgment fits

The hard parts — the $500 limit, the dispute-history check, the two-business-day clock — are already rules you can write. What code cannot do is read 'the merchant name means nothing to me' and tell you it is an unauthorised-use claim with checkable detail. That single semantic gap is the whole job.

Watch out

A fraud_signal split 45/45 between levels 2 and 3 is exactly the case a threshold would hide. Route the near-tie to an analyst rather than rounding it.

state — the dispute case
{
  "claim": "I never made this charge. I was at work and my card was in my wallet. The merchant name means nothing to me.",
  "transaction": {
    "merchant": "GLBL*DIGI SRV LTD",
    "amount_usd": 249.00,
    "mcc": "5817",
    "country": "MT",
    "entry_mode": "ecommerce_no_3ds",
    "timestamp": "2026-09-14T03:12:00Z"
  },
  "cardholder_history": {
    "account_age_months": 61,
    "prior_disputes_12mo": 0,
    "typical_country": "US",
    "avg_ticket_usd": 43.20
  },
  "policy": "Provisional credit within 2 business days for unauthorised-use claims under $500 with no dispute history."
}

Named JSON fields let a question point at a specific path. Everything a question needs must be here — nothing carries over between requests.

Pattern across all six

Three things repeat every time

01

The rules never move

Amount caps, date windows, final-sale flags, entitlement checks. These stay in code in every single scenario. The model is only ever asked to read.

02

One Choice, one Score, two Nouls

A rough default shape: what kind of thing is this, how bad is it, and two independent yes/no properties that the routing depends on.

03

Low confidence buys more compute

Below your threshold, escalate — to a reasoning model, a specialist, or a person. That branch is the product decision, not the model's.