Deal Terms Desk

Results

Every number on this page is generated from the project's measured facts. Numbers marked human-labelled (MAUD) are scored against lawyers' labels; numbers marked machine-built come from model passes, not lawyers.

Retrieval

The retrieval ladder on MAUD's questions (report split) human-labelled (MAUD)
RungWhat it addsrecall@5recall@10MRR@10nDCG@10recall@5 change from the rung abovems p50ms p95Context tokens
R1BM25 keyword search over section-aware passages0.4892 (0.4666 to 0.5123)0.59550.40250.3902—9.8349.471769.4
R2Dense only: a local embedding model0.3846 (0.3674 to 0.4029)0.46720.33720.3241-0.1047 (-0.1264 to -0.0826; hurts)23.6252.391659.2
R3Hybrid: R1 and R2 fused by reciprocal rank0.5111 (0.491 to 0.5335)0.61230.42310.41470.1265 (0.1084 to 0.1437; helps)45.34314.871777.3
R4R3 and a cross-encoder reranker0.502 (0.482 to 0.5217)0.61230.40190.3975-0.0091 (-0.0247 to 0.0054; no measurable change)1317.22088.261731.0
R5R4 and a lexicon rewrite of lay words into contract words (the lexicon is machine-built)0.5725 (0.5522 to 0.591)0.6910.45270.46520.0706 (0.0543 to 0.0871; helps)1333.272863.381772.7
R6R5 and each passage's defined terms0.5743 (0.5546 to 0.5942)0.69240.45480.4650.0018 (-0.0093 to 0.0133; no measurable change)1341.982230.93526.8
R6nR6 without the reranker: the path the live desk answers from0.5962 (0.5794 to 0.6116)———0.0218 (0.007 to 0.038; helps)———
R6n, live indexThe same rung over the one index file the live desk reads0.5921 (0.573 to 0.6114)———————

Report split: 1406 questions from 76 agreements. Intervals are bootstrap intervals clustered by agreement; changes are paired. Timings were measured on the development machine; the live path's and the live server's are in the last table.

The retrieval ladder on the tech deals (report split) machine-built
RungWhat it addsrecall@5MRR@10recall@5 change from the rung abovems p95Context tokens
R1BM25 keyword search over section-aware passages0.1672 (0.1391 to 0.1963)0.1356—21.581928.1
R2Dense only: a local embedding model0.6935 (0.6612 to 0.7247)0.53450.5263 (0.489 to 0.561; helps)21.841859.6
R3Hybrid: R1 and R2 fused by reciprocal rank0.6632 (0.6296 to 0.6956)0.405-0.0303 (-0.0562 to -0.0054; hurts)55.621866.8
R4R3 and a cross-encoder reranker0.4893 (0.4524 to 0.5253)0.3858-0.1739 (-0.2088 to -0.1399; hurts)2166.381916.3
R5R4 and a lexicon rewrite of lay words into contract words (the lexicon is machine-built)0.5594 (0.5242 to 0.5953)0.40920.07 (0.0516 to 0.0903; helps)1865.741902.0
R6R5 and each passage's defined terms0.5797 (0.5451 to 0.6142)0.4190.0203 (0.001 to 0.0395; helps)1670.173177.6
R6, all agreementsNo deal scoping: every agreement searched at once0.0229 (0.0134 to 0.0335)0.0176—2214.213028.2
R7, all agreements (first run)R6 inside the deal the question names0.7035 (0.6665 to 0.738)0.59160.6805 (0.6403 to 0.717; helps)1885.463393.7
R7, re-run for the answers, all agreementsThe same R7 as re-run in the answer milestone: the rung the live path's change is paired against0.6974 (0.6602 to 0.7314)————
R7n, all agreementsR7 without the reranker: the live path0.7727 (0.7351 to 0.8071)—0.0753 (0.0518 to 0.1008; helps)——

Report split: 621 machine-built questions about 216 tech agreements, each a lay question naming the company. Across all agreements, a passage from the wrong agreement counts as a miss.

Side comparisons, not rungs (MAUD's questions, report split)
Comparisonrecall@5ChangeAnswer key
Fixed-size chunks instead of section-aware passages, on R30.4977 (0.477 to 0.5185)-0.0134 (-0.0332 to 0.0064; no measurable change)human-labelled (MAUD)
A live model rewriting the question instead of the lexicon, on R50.5365 (0.5186 to 0.5536)-0.036 (-0.0524 to -0.0191; hurts)machine-built

By question family

By MAUD deal-point category (report split) human-labelled (MAUD)
CategoryR6 recall@5R6 change from R1Answer accuracyMost-common-answer baseline
Conditions to closing0.2290.0476 (0.0004 to 0.0936; helps)0.3355 (0.2932 to 0.3764)0.7887 (0.7462 to 0.8289)
Deal protection0.54340.0984 (0.0673 to 0.1277; helps)0.6403 (0.6182 to 0.6629)0.6687 (0.6429 to 0.6931)
General information0.1429-0.1476 (-0.2429 to -0.0571; hurts)0.8158 (0.7237 to 0.8947)0.6053 (0.4868 to 0.7105)
Knowledge0.90540.5405 (0.4189 to 0.6486; helps)0.6599 (0.5913 to 0.7222)0.8071 (0.75 to 0.8599)
Material adverse effect0.9467-0.0533 (-0.1067 to -0.0133; hurts)0.5985 (0.5799 to 0.6178)0.8027 (0.7852 to 0.8193)
Operating and efforts covenants0.71150.1024 (0.0489 to 0.1607; helps)0.6359 (0.5931 to 0.6809)0.7562 (0.7216 to 0.7888)
Remedies0.8867-0.1133 (-0.1867 to -0.0533; hurts)0.8553 (0.7763 to 0.9342)0.8553 (0.7763 to 0.9342)
By lead question family (tech deals, report split) machine-built
FamilyKept itemsTwo-pass agreement rateR6 recall@5 inside the dealR7 recall@5 across all agreementsAnswer agreesAgrees or partly
Employee equity awards3070.99030.50620.59050.07910.8047
Termination (break-up) fee2990.99340.56330.69680.00950.5619
Employees' pay and benefits after the deal2800.98590.67780.83460.02040.8163

Do the two answer keys agree?

The rung order under the lawyers' key and under the machine-built key
Rungrecall@5, lawyers' key human-labelled (MAUD)recall@5, machine-built key machine-built
R10.5080.498
R20.39480.3765
R30.53440.5227
R40.52960.5105
R50.61560.5777
R60.60250.5738

The same two-pass procedure that built the tech-deal key was run on MAUD's own questions, so those questions have two keys.

How far the machine-built key can be trusted machine-built
MeasureValue
The machine span overlaps the lawyers' span0.9789 (0.9615 to 0.9937)
Kendall's tau between the two rung orders1.0 (0.7333 to 1.0)
Items where both machine passes agreed331
Items asked555
Agreements30

This is evidence for or against trusting the machine-built key, not proof.

Answers

Answers to MAUD's questions (report split) human-labelled (MAUD)
MeasureValue
The answerer picks MAUD's answer0.6039 (0.5883 to 0.6181)
The same, counting only picks whose citations survived the gate0.5235 (0.5026 to 0.5415)
Always picking each question's most common answer (baseline)0.7528 (0.7393 to 0.7665)
Answerer minus baseline-0.1489 (-0.1698 to -0.1286; hurts)
Questions scored6307
Questions asked6321

Answer model: claude-haiku-4-5-20251001.

Answers to the tech-deal questions (report split) machine-built
MeasureValue
A judge model finds the answer agrees with the kept machine answers0.037 (0.0226 to 0.0533)
Agrees or partly agrees0.7262 (0.6849 to 0.7681)
The judge declined to rate128
Questions621

Judge model: claude-sonnet-5-5. The judge compares each answer with the two kept machine answers.

Declining when the answer is not there machine-built
GroupQuestionsCorrect declineFalse answer
Lead question, both passes found no such clause360.63890.3056
Earn-out question (one machine pass found none)300.76670.2
Answer sits in an unfiled schedule (weakest key)200.050.95
A company name that matches no deal's aliases301.00.0
A name matching several deals301.00.0

The right answer to each of these is a decline: "not stated in this agreement", "in a schedule that was not filed" or "which agreement?". The two name groups hold by construction: they test the deal resolver, which makes no model call.

Citations

The citation gate
QuestionsClaims returnedClaims keptShare kept
MAUD's questions1143198360.8605
Tech-deal questions156214490.9277

A claim is kept only if its quote occurs word for word in the passage it cites. This check needs no answer key.

A second model tries to refute each kept claim machine-built
MeasureValue
Share of claims it failed to refute0.7421
Claims checked11150
Unreadable replies0

Why things go wrong

Why retrieval misses or answers go wrong
ClassValueWhat it countsAnswer key
All misses (R6, MAUD)1433Report-split questions whose gold span is not in the top five passageshuman-labelled (MAUD)
Wrong section810The gold span sits in another sectionhuman-labelled (MAUD)
Right section, definition missing134The section was found but the definition it depends on was nothuman-labelled (MAUD)
Right section, wrong passage451Another passage of the right section was returnedhuman-labelled (MAUD)
Wrong deal0Tech-deal questions that R7 scoped to the wrong agreementmachine-built
Answer in an unfiled schedule0.95Share answered anyway when the answer sits in a schedule that was not filed (weakest key)machine-built
Label disputed0.0533 (0.0189 to 0.0966)Share of sampled misses where a model judged the first result to answer the questionmachine-built
Superseded textnot classifiedAmended passages are shown with their amendment; misses on them were not counted apart

Against the published baselines

Against LegalBench-RAG's published MAUD baselines (all agreements, character recall, percent)
Methodrecall@1recall@2recall@4recall@8recall@16recall@32recall@64Answer key
LegalBench-RAG: naive fixed-size chunks2.543.124.538.7513.1618.3625.62published
LegalBench-RAG: recursive text splitter1.652.094.596.1812.9321.0428.28published
LegalBench-RAG: recursive splitter and Cohere reranker0.522.484.397.2414.0322.631.46published
This project, R10.190.451.01.73.726.5611.92human-labelled (MAUD)
This project, its best rung0.340.661.372.694.658.1614.87human-labelled (MAUD)

Published figures: Pipitone and Houir Alami, LegalBench-RAG, arXiv:2408.10343v1. The setups differ (chunking, query wording, and k counted in their chunks against our passages), so read this as indicative, not head to head.

The live desk

The live desk
MeasureValue
Answer modelclaude-haiku-4-5-20251001
Price per million input tokens, US dollars1.0
Price per million output tokens, US dollars5.0
Prices checked on2026-10-05
Hosting per month, US dollars7.09
Model budget per month, US dollars2.91
Model budget per day, US dollars0.29
Input tokens per answer, mean4250.3
Output tokens per answer (including thinking), mean1325.7
Cost per new answer, mean, US dollars0.010879
New answers the monthly budget covers267
Live-path retrieval inside one agreement (R6n), development machine, ms p5046.4
Live-path retrieval inside one agreement (R6n), development machine, ms p95365.6
Live-path retrieval across all agreements (R7n), development machine, ms p95516.66
Search on the server, ms p5076.6
Search on the server, ms p95120.9
A new answer on the server, ms p5017225.8
A new answer on the server, ms p9524998.1
A cached answer on the server, ms p5077.0
Peak memory of the service, MB528.7
Index file, bytes634859520
The cap trips: a new question is refusedyes
The cap trips: a cached answer is still shownyes

Token counts come from a calibration sample of 40 questions sent to the API. Server timings leave out the network.