Eval bench · golden set of 26 labeled submissions · prompt v1 vs v2
last full run Jul 27, 10:29 PMField accuracy
100.0%
Decision accuracy
100.0%
Cost / submission
$0.0244
Latency p50 / p95
17.6s / 22.6s
Vertafore self-reports 87% accuracy on submission processing. This harness is how we'd continuously know ours — same pipeline, labeled golden set, per-field scoring. (Mock data; methodology, not a comparison.)
Per-field extraction accuracyv1 → v2
| Field | v1 | v2 | Δ | n |
|---|---|---|---|---|
| insuredName | 92.0% | 100.0% | +8.0pp | 25 |
| state | 92.3% | 100.0% | +7.7pp | 26 |
| county | 91.7% | 100.0% | +8.3pp | 24 |
| units | 92.0% | 100.0% | +8.0pp | 25 |
| tiv | 96.2% | 100.0% | +3.8pp | 26 |
| constructionType | 70.6% | 100.0% | +29.4pp | 17 |
| yearBuilt | 92.0% | 100.0% | +8.0pp | 25 |
| targetPremium | 94.4% | 100.0% | +5.6pp | 18 |
| subclass | 100.0% | 100.0% | 0.0pp | 8 |
Live spot-check
Re-runs 5 golden cases through the live pipeline with prompt v2 — real API calls, scored on arrival. Full 26-case runs happen at build time (`pnpm evals`) and are committed.