Eval bench
golden set of 27 labeled submissions│prompt v1 vs v2 vs v3
Field accuracy
100.0%
Decision accuracy
100.0%
Cost / submission
$0.0277
Latency p50 / p95
19.3s / 23.7s
Vertafore self-reports 87% accuracy on submission processing. This harness is how we'd continuously know ours — same pipeline, labeled golden set, per-field scoring. (Mock data; methodology, not a comparison.)
Per-field extraction accuracyv1 → v3
| Field | v1 | v3 | Δ | n |
|---|---|---|---|---|
| insuredName | 92.3% | 100.0% | +7.7pp | 26 |
| state | 92.6% | 100.0% | +7.4pp | 27 |
| county | 92.0% | 100.0% | +8.0pp | 25 |
| units | 92.3% | 100.0% | +7.7pp | 26 |
| tiv | 96.3% | 100.0% | +3.7pp | 27 |
| constructionType | 72.2% | 100.0% | +27.8pp | 18 |
| yearBuilt | 92.3% | 100.0% | +7.7pp | 26 |
| targetPremium | 94.7% | 100.0% | +5.3pp | 19 |
| subclass | 88.9% | 100.0% | +11.1pp | 9 |
Live spot-check
Re-runs 5 golden cases through the live pipeline with prompt v3 — real API calls, scored on arrival. Full 27-case runs happen at build time (`pnpm evals`) and are committed.
- 01queued
- 02queued
- 03queued
- 04queued
- 05queued