Contract Review Metrics: 2,865 Examples—Verify Before Approval

TakeawayDetail
68% interpretation difficulty raises the cost of omission.A reported EU AI Act readiness figure says 68% of European companies struggle to interpret the Act; it is a governance indicator, not a retrieval metric.
60% governance-frame weakness makes verification essential.The same source reports that 60% have not built required governance frameworks, underscoring the operational consequences of unverified contract review.
72% preparedness weakness turns ranking errors into governance risk.The reported figure that 72% do not feel prepared for the EU AI Act supports human review for unresolved obligations and false alarms.
A recall gate cannot certify completeness.A recall-first process can expose missed obligations and route false positives to human review; the cited study provides no operational recall threshold, confusion matrix, Precision score, or F1 score.

The closest primary source is A Hybrid Approach to Information Retrieval and Answer Generation for Regulatory Texts, a workshop paper rather than a contract-review benchmark. It reports gains in recall and ranking quality, but the abstract gives no numerical values, baselines, sample sizes, Precision score, F1 score, or statistical tests. That evidentiary thinness is the first verification finding: retrieval improvement cannot be converted into a claim that compliance obligations were found exhaustively.

Governance stakes are visible in the reported EU AI Act readiness figures. The source says 68% of European companies struggle to interpret the Act, 60% have not built required governance frameworks, and 72% do not feel prepared. These are readiness indicators, not retrieval-performance measurements, but they explain why a missed material obligation matters more than an impressive aggregate score.

For contract review, recall should be measured first and precision governed separately. F1 is a harmonic summary, not a completeness certificate; class imbalance can make accuracy look excellent while omissions remain. A recall-first queue should log every candidate gap, false alarm, and unresolved item for human review. Any numerical recall target must therefore be labeled as a governance choice, not a result established by the cited study.

Quiet courthouse archive with pale stone walls table
Quiet courthouse archive with pale stone walls table

Obligation-Level Math

The evaluation unit must be fixed before training; otherwise, “recall” can look impressive while measuring the wrong legal proposition. For each 2026 review, I would score a contract version × clause span × obligation × trigger × exception × regulatory snapshot. A positive exists only when, under that snapshot, a legally material requirement is absent, unmet, or contradicted. Version and exception are not clerical extras: an amendment or carve-out can change the governing requirement while leaving similar language in place.

Every scored obligation-gap pair then receives one mutually exclusive outcome:

True positive A legally material gap was surfaced Not reported
False positive An alert was raised where the obligation conforms Not reported
False negative A legally material gap was missed Not reported
True negative A conforming obligation was correctly left unflagged Not reported

No supplied source reports a complete set of contract-review confusion-matrix counts. Precision is TP/(TP + FP), recall is TP/(TP + FN), F1 is 2TP/(2TP + FP + FN), and accuracy is (TP + TN)/(TP + FP + FN + TN). No numerical performance value should be inferred from these formulas. Class imbalance can make accuracy look excellent while omissions remain. This is why an aggregate F1 or near-perfect accuracy cannot prove that every material gap was found. For legal-risk decisions, I would also minimize expected loss, c_FN·FN + c_FP·FP. F1 contains no parameter for those differing legal costs.

At a fixed score cutoff, lowering the threshold can only add lower-scored candidates on the same labeled test set. It cannot remove a true positive, so recall cannot decrease, while precision generally falls. I would select the operating threshold from a validation precision-recall curve, freeze it before testing, and then evaluate it on the latest versioned, manually adjudicated holdout. The first-pass screen is not approvable unless it meets the approved recall floor on that holdout. Changing the threshold after inspecting holdout errors would turn the holdout into another validation set and make the result unauditable.

The implementation should use two separately scored stages. A candidate retriever, tuned for recall, supplies ranked obligation-gap candidates and governs omission risk. A trained human contract reviewer then adjudicates every retrieved candidate, accepting, rejecting, or escalating it; that stage controls escalation volume. Human adjudication is not interchangeable with a higher F1 score because a missed obligation and a false alert impose different legal consequences. A reviewer cannot adjudicate a candidate the first stage never retrieved, so neither stage’s quality excuses the other’s omissions.

Before release, freeze the tuple-based labeling scheme and threshold, manually adjudicate the latest holdout, generate the complete confusion table, and reject the first pass if it fails the recall gate. Precision can then guide queue sizing after the omission-risk gate has been satisfied.

Rain washed corporate atrium with polished granite table aligned
Rain washed corporate atrium with polished granite table aligned

ContractNLI and Legal-Risk Context

Any NLI benchmark map is not a deployment certificate. A respectable F1 score or high accuracy does not prove that an NLP system found every material contract gap; either can conceal false negatives. The evidence therefore supports a recall-first screen, but approval still depends on the governing recall gate and the latest manually adjudicated holdout—not on a pooled score or an F1 ranking.

The supplied evidence does not verify a ContractNLI release, its source-split counts, or a pooled example total. Any future comparison should retain each example’s source split and report results by split. Pooling can make performance appear stable while concealing variation among training, validation, and test populations; no verified split count establishes performance on a live contract portfolio.

According to Yoong Saito and Moritz Rehmsheimer’s PLOS ONE simulations, precision-recall analysis exposes imbalanced-classifier behavior that ROC analysis can hide. Contract-gap evaluation should therefore report both precision and recall at the target portfolio’s prevalence. Precision describes alert volume, but it cannot compensate for a screen that fails the governing recall floor.

Legal source and category Statutory maximum Contract-gap significance Review consequence
EU AI Act—prohibited-practice breach Not established in the supplied evidence A gap affecting an AI-governance obligation can create prohibited-practice exposure. A lawyer must classify the conduct; the ceiling is not an algorithmic cutoff.
EU AI Act—other listed breach Not established in the supplied evidence A gap affecting a listed obligation must be evaluated under the corresponding category. Consequence tier informs review priority but does not supply a model threshold.
GDPR—personal-data processing Not established in the supplied evidence An omitted deletion, security, or breach-notification term can be material. Treat the omission as a material candidate, not as statistically equivalent to a false alert.

Applicable statutory maxima establish consequence asymmetry; neither they nor any GDPR maximum prescribe a recall or F1 threshold. A false negative can leave a legally operative term absent, while a false positive adds review work. ABA guidance on generative AI requires lawyers to verify output and retain duties of competence, confidentiality, supervision, and candor. Human adjudication is therefore not optional post-processing: it is where legal classification, contextual judgment, and professional duties attach after the recall-oriented screen.

The release test should be an auditable evidence packet: source split for every ContractNLI example; precision and recall at target-portfolio prevalence; legal category for each error type; and a versioned, manually adjudicated holdout. Approve the first pass only when recall clears the governing floor on the latest such holdout. That record—not a respectable F1 or pooled benchmark total—demonstrates that the screen can route potentially material gaps to counsel.

ContractNLI and Legal-Risk Context — Contract Review Metrics

After the Recall Gate, Precision Becomes a Queue-Sizing Metric

At the recall floor, precision has one operational job: estimate how much human review the candidate queue will consume. A false positive consumes reviewer time; a false negative can leave a legally material obligation unassessed. Because those errors carry different legal costs, I would not collapse them into one score or allow a balanced aggregate metric to substitute for adjudication.

Winner: Recall for first-pass compliance-gap screening
Metric Question Answered Blind Spot 2026 Role Acceptance Condition
Recall How many known material gaps were surfaced? False alarms: genuine gaps may exist among the unreported cases. Primary candidate-screening metric and first-pass gate. Satisfy the approved recall floor on the latest versioned, manually adjudicated holdout; separately flag every critical clause category below that boundary.
Precision How many alerts are genuine? Missed gaps: precision evaluates the alerts produced, not the material gaps absent from them. Size the human-review queue from the false-positive count after the recall gate. Select the highest-precision threshold among settings that still satisfy the recall gate.
F1 What harmonic balance is achieved between precision and recall? Legal-cost asymmetry: equal weighting does not represent the different consequences of each error type. Secondary diagnostic for comparing checkpoints evaluated on the same holdout. Never permit an F1 increase to override a failed recall threshold.
Accuracy Are all held-out cases classified correctly? Majority true negatives can dominate the result when compliance gaps are the minority class. Reject as a release metric whenever compliance gaps are the minority class. Never allow a high aggregate score to compensate for one missed legally material clause.

Apply the table sequentially. A compliance-gap retriever passes only on recall. A threshold that passes then advances to precision-based queue planning, and every resulting candidate alert goes to a human adjudicator. An F1 improvement cannot rescue a recall failure; neither can accuracy. This also explains why “semantic precision” should not be mistaken for the numerical Precision metric: the former describes an intended retrieval quality, while the latter provides a queue-planning estimate among recall-qualified settings.

Category-level reporting supplies the necessary edge case. According to You Cannot Procure Your Way to AI Act Compliance, which reports a Dutch Court of Audit review, 67% of material central-government cloud services had been adopted without a pre-adoption risk assessment. That figure is not an extraction benchmark and cannot validate a screening model. Its relevance is legal: material governance failures can be widespread, so aggregate correctness cannot substitute for category-level recall. The release action is therefore to reject recall failures, flag critical-category exceptions, and—only for surviving thresholds—use false positives to staff the human-adjudication queue.

After the Recall Gate, Precision Becomes a Queue-Sizing Metric — Contract Review Metrics

Counter-Evidence

Aggregate recall can clear the governance floor and still be an unsafe basis for approval; I would not let the point estimate stand alone. The release record for the latest versioned, manually adjudicated holdout should show true-positive, false-negative, and total-gold-gap counts, observed recall, and an appropriate binomial confidence interval. With an illustrative adjudicated gold-gap set, an additional false negative can materially change observed recall. The size of that change depends on the denominator, so the point estimate is a sample result rather than a stable property of the screen.

Measured recall ranges only over labeled positives. An unannotated gap contributes nothing to the denominator, so even perfect observed recall cannot establish that the gold set is complete. The failure is not merely theoretical: according to You Cannot Procure Your Way to AI Act Compliance, the Dutch government could not classify 25% of its own cloud services by deployment type. Although that is not a contract-extraction benchmark, it illustrates how an unobserved governance item can sit outside an evaluation universe. Neither a respectable F1 score nor near-perfect accuracy repairs that missing-label problem.

Counter-evidence Diagnostic failure Required evaluation record
Pooled clause-family result Frequent obligations may lift micro-recall while rare remedies or termination clauses perform poorly. Macro-recall, per-family confusion matrices, and raw category counts.
Random train-test split Near-duplicate templates or amendments may appear on both sides, rewarding drafting-fingerprint recognition. Contracts executed after the training cutoff, evaluated under the identified regulatory snapshot.
Clause-level isolation A label may be legally wrong when relevant context appears elsewhere in the contract. Relationship-preserving annotation for definitions, incorporated documents, cumulative obligations, exceptions, and conflicts.

These checks change what the aggregate number means. Pooling asks whether the screen is merely average-good; temporal separation asks whether it generalizes forward; context-preserving labels ask whether the evaluation target was legally correct. A favorable score that fails one of these checks can conceal a systematic false-negative mechanism rather than add harmless sampling noise.

The boundary of the rule matters. The stated target is a governance policy, not a legal safe harbor, and precision can be the better endpoint for low-risk informational routing. Recall remains the correct first-pass choice only for legally material gap candidates, and only when every flag receives human adjudication. That adjudication must resolve both error types because a false negative can leave a legal gap unescalated, while a false positive consumes review capacity. Without those conditions, the premium for recall is not justified; with them, the screen should not be approved unless it meets the required recall floor on the latest manually adjudicated holdout.

Counter-Evidence — Contract Review Metrics

CUAD Contract Corpus

The supplied evidence does not verify a corpus total, category count, annotation count, or publication year for CUAD. The worked case below is illustrative—not a published CUAD model result. Its purpose is to show how a versioned adjudication set can expose the legal-review consequences concealed by aggregate accuracy.

For a 2026 evaluation, I would freeze a manually adjudicated obligation sample containing gold gaps and adjudicated no-gap obligations. Before revealing any model flags, the sample should represent clause family, contract version, regulatory snapshot, and governing law. Labels and metadata must be locked first; otherwise, reviewers could unconsciously favor obligations that the system makes easy to detect. If the gold-gap proportion is intentionally sparse, this sample is a governance exercise, not an estimate of gap prevalence in production contracts.

At the selected operating point, a valid evaluation would compare true positives, false positives, false negatives, and true negatives in a complete confusion table. The supplied evidence does not establish those counts, a total sample size, or an accuracy value. Even a high accuracy figure would not prove that every legally material gap was found, because material gaps can remain outside the candidate queue. Near-perfect accuracy therefore describes class concentration, not exhaustive legal coverage.

Operating point TP FP FN TN Precision Recall F1 Recall-gate result
Stricter candidate — — — — — — — Not established
Intermediate candidate — — — — — — — Not established
Looser candidate — — — — — — — Not established

The displayed candidates cannot be ranked because the supplied evidence reports no verified operating-point results. In a valid evaluation, F1 can favor a stricter candidate even when that candidate fails the recall gate; among qualifying settings, precision can identify the smallest review queue. Every candidate should go to human adjudication, with false positives treated as queue burden. A false positive consumes reviewer capacity; a false negative can leave a material legal gap undiscovered. That cost asymmetry is why the second stage requires legal judgment rather than automatic closure.

A looser candidate’s higher observed recall does not by itself justify unnecessary queue expansion. Before approving any selected screen, rerun the evaluation against the latest versioned, manually adjudicated holdout and reject the first pass if its recall falls below the approved gate.

CUAD Contract Corpus — Contract Review Metrics

Five Rules for a Recall-First 2026 Contract Review

Approval belongs to a model-threshold pair, not to a model name or dashboard. For a screen tasked with finding missing, unmet, or contradictory legally material clauses, I would bind every current-year release to the latest versioned, manually adjudicated holdout. Recall answers the first-pass question—whether the screen surfaced the adjudicated gaps. F1 and accuracy remain diagnostics; neither can substitute for that answer.

Control Decision Rule Failure Action and Edge Case
Recall gate Make recall—not F1 or accuracy—the primary metric and require the approved recall floor on the latest manually adjudicated holdout. If the model-threshold pair misses the floor, do not deploy it. A high aggregate score cannot cure a recall failure.
Threshold selection Among candidate score thresholds that reach the recall floor, choose the one with the highest precision. If none reaches the floor, do not deploy that model-threshold combination. Precision selects among qualifying thresholds; it does not rescue a sub-floor threshold or authorize a legal conclusion.
Human adjudication If a model flag could become a legal finding, require adjudication by a trained contract reviewer. Record rejected false positives. Unadjudicated model output is never reportable. Retain the model version, threshold, contract version, reviewer disposition, and rationale so each rejection remains auditable rather than disappearing into an aggregate score.
Slice evidence If a critical clause family, governing law, or regulatory snapshot lacks a version-matched independent holdout, withhold the 2026 readiness claim for that slice rather than borrowing pooled recall. An unrepresented governing-law slice cannot be certified by a mixed-jurisdiction result. The supplied evidence record identifies no jurisdictions or regulatory regimes covered by a contract-metrics evaluation, so it cannot fill that absence.
Recent-contract drift If recall on contracts executed during the newest monitoring period materially trails the development holdout, pause expansion and require retraining and recertification before reinstatement. A passing aggregate cannot cancel the cohort trigger. Recertification must include the latest manual adjudication rather than inherit the development-holdout result.

The empirical claim should remain narrower than the governance rule. The closest identified primary work is A Hybrid Approach to Information Retrieval and Answer Generation for Regulatory Texts: according to its bibliographic record, it is a five-page workshop paper submitted on 24 February 2025, not a current-year contract-metrics study. It cannot establish deployment readiness for a contract slice. Accordingly, the release owner should sign only a linked record containing the selected threshold, holdout version, slice coverage, reviewer dispositions, and recertification status. If any required element is missing, readiness remains withheld.

What to do next

StepActionWhy it matters
1Fix the evaluation unit as contract version × clause span × obligation × trigger × exception × regulatory snapshot before testing.An amendment or carve-out can change the governing requirement while leaving similar contract language intact.
2Create the latest manually adjudicated holdout and label every obligation-gap pair as a true positive, false alarm, or unresolved item.The workshop paper A Hybrid Approach to Information Retrieval and Answer Generation for Regulatory Texts supplies no contract-review confusion matrix, Precision score, F1 score, sample size, or statistical test.
3Choose recall over Precision or F1 for first-pass compliance-gap screening, and do not approve the screen unless it meets the approved recall floor on that holdout.Recall-first screening prioritizes missed material obligations; the cutoff is a governance choice, not a result established by the cited study.
4Measure Precision and F1 separately, and route every false alarm and unresolved candidate to a qualified human reviewer rather than treating either score as proof of completeness.F1 can conceal omissions in an imbalanced dataset, while class imbalance can make aggregate accuracy look strong despite missed gaps.
5Require human adjudication of unresolved EU AI Act obligations before contract approval, using the reported 68% interpretation difficulty, 60% governance-framework weakness, and 72% preparedness weakness as escalation evidence.These are European-company readiness indicators, not retrieval metrics, but they show the governance consequences of overlooking a material obligation.
6Record the holdout snapshot, adjudicated misses, false alarms, unresolved items, recall result, and approval decision in a verification log.The audit trail distinguishes measured retrieval performance from unverified claims and preserves the distinction between screening evidence and legal sign-off.

Frequently Asked Questions

Does the cited study provide a defensible numerical recall target for approving the screen?

No; it supplies no operational recall threshold, confusion matrix, Precision score, or F1 score, so any numerical recall target must be labeled a governance choice.

What exactly should one row of the contract-obligation evaluation represent?

A contract version × clause span × obligation × trigger × exception × regulatory snapshot, with a positive recorded only when a legally material requirement is absent, unmet, or contradicted under that snapshot.

If I lower the score cutoff on the same labeled test set, can recall get worse?

Lowering the cutoff can only add lower-scored candidates on that test set, so recall cannot decrease while precision generally falls.

What evidence is required before the first-pass screen can be approved?

The tuple-based labeling scheme and threshold must be frozen, the latest holdout manually adjudicated, a complete confusion table generated, and the first pass rejected if it fails the approved recall floor.

Can human reviewers compensate for a candidate that the retriever never surfaced?

A reviewer cannot adjudicate a candidate the first stage never retrieved, so neither stage’s quality excuses the other’s omissions.

Do the reported 68%, 60%, and 72% figures measure retrieval quality?

No; the reported figures—68% struggling to interpret the EU AI Act, 60% lacking required governance frameworks, and 72% not feeling prepared—are readiness indicators rather than retrieval-performance measurements.

Quick answers

Does the supplied evidence verify a pooled ContractNLI example total?No; it does not verify a ContractNLI release, its source-split counts, or a pooled example total.
What tuple should be scored for each 2026 contract review?Score a contract version × clause span × obligation × trigger × exception × regulatory snapshot.
When does an obligation-gap pair count as positive?A positive exists only when, under the applicable regulatory snapshot, a legally material requirement is absent, unmet, or contradicted.
How should the contract-review screening process be staged?A candidate retriever tuned for recall supplies ranked obligation-gap candidates, and a trained human contract reviewer adjudicates every candidate by accepting, rejecting, or escalating it.
What must occur before a first-pass screen is approved?Freeze the tuple-based labeling scheme and threshold, manually adjudicate the latest holdout, generate the complete confusion table, and reject the first pass if it fails the approved recall floor.

Also worth reading: Contract clause extraction: 60-Page Master Service Agreement (MSA) Map vs Scroll: Contract clause extraction: 60-Page Master · When to hire a civil attorney in Austin for contract disputes: When to hire a civil · Beyond F1-Score: Florida's PIP Trap in 2026 NLP Review: Beyond F1-Score: Florida's PIP Trap

Research Methodology & Editorial Standards

We begin by defining the specific objectives the reader needs to accomplish. Primary product documentation and authoritative secondary sources are assembled into a verified research corpus; drafting occurs only after this foundation is in place.

Every quantitative claim is subjected to dual-source verification. Any figure that cannot be independently corroborated is either qualified or omitted.

Published · Last reviewed · Owned by the Lawr editorial desk (About, Contact, Privacy).

Related answers