Exchange TriageMember workspace

HABITATIONAL SPECIALTY MGA│Session —
Exchange onlineBind desk4 open

Eval bench

golden set of 27 labeled submissions│prompt v1 vs v2 vs v3
last full run Aug 15, 03:26 PM
Field accuracy
100.0%
Decision accuracy
100.0%
Cost / submission
$0.0277
Latency p50 / p95
19.3s / 23.7s

Vertafore self-reports 87% accuracy on submission processing. This harness is how we'd continuously know ours — same pipeline, labeled golden set, per-field scoring. (Mock data; methodology, not a comparison.)

Per-field extraction accuracyv1 → v3
Fieldv1v3Δn
insuredName92.3%100.0%+7.7pp26
state92.6%100.0%+7.4pp27
county92.0%100.0%+8.0pp25
units92.3%100.0%+7.7pp26
tiv96.3%100.0%+3.7pp27
constructionType72.2%100.0%+27.8pp18
yearBuilt92.3%100.0%+7.7pp26
targetPremium94.7%100.0%+5.3pp19
subclass88.9%100.0%+11.1pp9
Live spot-check

Re-runs 5 golden cases through the live pipeline with prompt v3 — real API calls, scored on arrival. Full 27-case runs happen at build time (`pnpm evals`) and are committed.

  1. 01queued
  2. 02queued
  3. 03queued
  4. 04queued
  5. 05queued