EXCHANGE TRIAGE/ Member workspace

Member: HABITATIONAL SPECIALTY MGA · Session
Exchange onlineBind desk: 4 open
Eval bench · golden set of 26 labeled submissions · prompt v1 vs v2
last full run Jul 27, 10:29 PM
Field accuracy
100.0%
Decision accuracy
100.0%
Cost / submission
$0.0244
Latency p50 / p95
17.6s / 22.6s

Vertafore self-reports 87% accuracy on submission processing. This harness is how we'd continuously know ours — same pipeline, labeled golden set, per-field scoring. (Mock data; methodology, not a comparison.)

Per-field extraction accuracyv1v2
Fieldv1v2Δn
insuredName92.0%100.0%+8.0pp25
state92.3%100.0%+7.7pp26
county91.7%100.0%+8.3pp24
units92.0%100.0%+8.0pp25
tiv96.2%100.0%+3.8pp26
constructionType70.6%100.0%+29.4pp17
yearBuilt92.0%100.0%+8.0pp25
targetPremium94.4%100.0%+5.6pp18
subclass100.0%100.0%0.0pp8
Live spot-check

Re-runs 5 golden cases through the live pipeline with prompt v2 — real API calls, scored on arrival. Full 26-case runs happen at build time (`pnpm evals`) and are committed.