Model Quality · Internal

Evals

Ground-truth benchmarks across extraction, covenant detection, and memo generation. Updated each model/prompt revision.

Current: v3.2n = 1,101
Extraction Accuracy
94.3%
+13.1 pts vs v1.0
Breach Precision
92.3%
12 TP · 1 FP
Breach Recall
100.0%
12 caught · 0 missed
Memo Hallucination Rate
6.2%
↓ 12.2 pts vs v1.0

Breach Detection · Precision / Recall

n = 13 events
12
True Positive
1
False Positive
0
False Negative

Version Regression

v2.0 → v3.2

Field-Level Accuracy

Held-out set
Field
Accuracy
n
%
Facility Size
42
100.0%
Covenant Threshold
168
97.6%
ARR
288
98.8%
Cash Balance
288
99.3%
Burn Multiple (derived)
288
90.2%
MAC clauses (interpretive)
27
74.1%
Synthetic data, portfolio demonstration only.