Model Quality · Internal
Evals
Ground-truth benchmarks across extraction, covenant detection, and memo generation. Updated each model/prompt revision.
Current: v3.2n = 1,101
Extraction Accuracy
94.3%
+13.1 pts vs v1.0
Breach Precision
92.3%
12 TP · 1 FP
Breach Recall
100.0%
12 caught · 0 missed
Memo Hallucination Rate
6.2%
↓ 12.2 pts vs v1.0
Breach Detection · Precision / Recall
n = 13 events12
True Positive
1
False Positive
0
False Negative
Version Regression
v2.0 → v3.2Field-Level Accuracy
Held-out setField
Accuracy
n
%
Facility Size
42
100.0%
Covenant Threshold
168
97.6%
ARR
288
98.8%
Cash Balance
288
99.3%
Burn Multiple (derived)
288
90.2%
MAC clauses (interpretive)
27
74.1%
Synthetic data, portfolio demonstration only.