Triage at the
speed of the desk
A broker sends a messy email. Five agents pull it apart, check it against your appetite ruleset, and hand back a scored risk — with every step of the reasoning on the table.
New habitational risk for your landlord program. Insured is Maple Row Apartments LLC, a 24-unit garden-style apartment complex at 4410 Kenway Blvd, Columbus, OH 43220 (Franklin County). Joisted masonry construction, built 1998, fully sprinklered. TIV is $4,200,000.
- Intake
- Extraction0/6 fields
- Source grounding
- Appetite check
- Decision
Five agents, one
auditable path
Each node does one job and shows its work. No single prompt is asked to be clever about everything at once.
Intake
The broker email and its attachment land together. Nothing is normalised yet — the raw text is kept, because every later claim has to point back at it.
Extraction
Structured fields stream out token by token: insured, address, state, county, units, TIV, construction, year built.
Source grounding
Four guardrails run: schema, grounding, required fields, and an aggregate confidence gate at 0.70. Any one of them can stop the risk here.
Appetite check
Six rules from the Member's own ruleset execute as tool calls — allowed states, unit cap, TIV band, excluded construction, coastal counties, premium band.
Decision
A fit score and a bind-ready summary, or a written referral note naming exactly which check failed and why.
Watch it think,
not just answer
Tokens stream as they are generated. Tool calls appear with their arguments. Guardrails flash the moment they pass or trip. An underwriter can see which sentence in the broker's email produced which field — and disagree with it.
The interesting
case is the one
it refuses
A model that always answers is a liability. When the email says twelve units and the attachment totals eight, the grounding check trips, confidence drops under the gate, and the risk goes to the referral queue with the contradiction written out.
Claims you can check
A 26-case golden set with committed results. The extraction prompt is version-frozen so the regression stays honest across changes.
insuredName, state, county, units, TIV, construction, year built — scored per field, not in aggregate.
Cleared vs referred, against the golden label for all 26 cases.
Full 26-case sweep. Avg 3,306 in / 965 out tokens per submission.
Run one yourself
Ten seed submissions are loaded, including the two that are supposed to fail. No signup, no key needed to look around.
Open the demo