Results
Every number on this page is generated from the project's measured facts. Numbers marked human-labelled (MAUD) are scored against lawyers' labels; numbers marked machine-built come from model passes, not lawyers.
Retrieval
The retrieval ladder on MAUD's questions (report split) human-labelled (MAUD)
| Rung | What it adds | recall@5 | recall@10 | MRR@10 | nDCG@10 | recall@5 change from the rung above | ms p50 | ms p95 | Context tokens |
| R1 | BM25 keyword search over section-aware passages | 0.4892 (0.4666 to 0.5123) | 0.5955 | 0.4025 | 0.3902 | — | 9.83 | 49.47 | 1769.4 |
| R2 | Dense only: a local embedding model | 0.3846 (0.3674 to 0.4029) | 0.4672 | 0.3372 | 0.3241 | -0.1047 (-0.1264 to -0.0826; hurts) | 23.6 | 252.39 | 1659.2 |
| R3 | Hybrid: R1 and R2 fused by reciprocal rank | 0.5111 (0.491 to 0.5335) | 0.6123 | 0.4231 | 0.4147 | 0.1265 (0.1084 to 0.1437; helps) | 45.34 | 314.87 | 1777.3 |
| R4 | R3 and a cross-encoder reranker | 0.502 (0.482 to 0.5217) | 0.6123 | 0.4019 | 0.3975 | -0.0091 (-0.0247 to 0.0054; no measurable change) | 1317.2 | 2088.26 | 1731.0 |
| R5 | R4 and a lexicon rewrite of lay words into contract words (the lexicon is machine-built) | 0.5725 (0.5522 to 0.591) | 0.691 | 0.4527 | 0.4652 | 0.0706 (0.0543 to 0.0871; helps) | 1333.27 | 2863.38 | 1772.7 |
| R6 | R5 and each passage's defined terms | 0.5743 (0.5546 to 0.5942) | 0.6924 | 0.4548 | 0.465 | 0.0018 (-0.0093 to 0.0133; no measurable change) | 1341.98 | 2230.9 | 3526.8 |
| R6n | R6 without the reranker: the path the live desk answers from | 0.5962 (0.5794 to 0.6116) | — | — | — | 0.0218 (0.007 to 0.038; helps) | — | — | — |
| R6n, live index | The same rung over the one index file the live desk reads | 0.5921 (0.573 to 0.6114) | — | — | — | — | — | — | — |
Report split: 1406 questions from 76 agreements. Intervals are bootstrap intervals clustered by agreement; changes are paired. Timings were measured on the development machine; the live path's and the live server's are in the last table.
The retrieval ladder on the tech deals (report split) machine-built
| Rung | What it adds | recall@5 | MRR@10 | recall@5 change from the rung above | ms p95 | Context tokens |
| R1 | BM25 keyword search over section-aware passages | 0.1672 (0.1391 to 0.1963) | 0.1356 | — | 21.58 | 1928.1 |
| R2 | Dense only: a local embedding model | 0.6935 (0.6612 to 0.7247) | 0.5345 | 0.5263 (0.489 to 0.561; helps) | 21.84 | 1859.6 |
| R3 | Hybrid: R1 and R2 fused by reciprocal rank | 0.6632 (0.6296 to 0.6956) | 0.405 | -0.0303 (-0.0562 to -0.0054; hurts) | 55.62 | 1866.8 |
| R4 | R3 and a cross-encoder reranker | 0.4893 (0.4524 to 0.5253) | 0.3858 | -0.1739 (-0.2088 to -0.1399; hurts) | 2166.38 | 1916.3 |
| R5 | R4 and a lexicon rewrite of lay words into contract words (the lexicon is machine-built) | 0.5594 (0.5242 to 0.5953) | 0.4092 | 0.07 (0.0516 to 0.0903; helps) | 1865.74 | 1902.0 |
| R6 | R5 and each passage's defined terms | 0.5797 (0.5451 to 0.6142) | 0.419 | 0.0203 (0.001 to 0.0395; helps) | 1670.17 | 3177.6 |
| R6, all agreements | No deal scoping: every agreement searched at once | 0.0229 (0.0134 to 0.0335) | 0.0176 | — | 2214.21 | 3028.2 |
| R7, all agreements (first run) | R6 inside the deal the question names | 0.7035 (0.6665 to 0.738) | 0.5916 | 0.6805 (0.6403 to 0.717; helps) | 1885.46 | 3393.7 |
| R7, re-run for the answers, all agreements | The same R7 as re-run in the answer milestone: the rung the live path's change is paired against | 0.6974 (0.6602 to 0.7314) | — | — | — | — |
| R7n, all agreements | R7 without the reranker: the live path | 0.7727 (0.7351 to 0.8071) | — | 0.0753 (0.0518 to 0.1008; helps) | — | — |
Report split: 621 machine-built questions about 216 tech agreements, each a lay question naming the company. Across all agreements, a passage from the wrong agreement counts as a miss.
Side comparisons, not rungs (MAUD's questions, report split)
| Comparison | recall@5 | Change | Answer key |
| Fixed-size chunks instead of section-aware passages, on R3 | 0.4977 (0.477 to 0.5185) | -0.0134 (-0.0332 to 0.0064; no measurable change) | human-labelled (MAUD) |
| A live model rewriting the question instead of the lexicon, on R5 | 0.5365 (0.5186 to 0.5536) | -0.036 (-0.0524 to -0.0191; hurts) | machine-built |
By question family
By MAUD deal-point category (report split) human-labelled (MAUD)
| Category | R6 recall@5 | R6 change from R1 | Answer accuracy | Most-common-answer baseline |
| Conditions to closing | 0.229 | 0.0476 (0.0004 to 0.0936; helps) | 0.3355 (0.2932 to 0.3764) | 0.7887 (0.7462 to 0.8289) |
| Deal protection | 0.5434 | 0.0984 (0.0673 to 0.1277; helps) | 0.6403 (0.6182 to 0.6629) | 0.6687 (0.6429 to 0.6931) |
| General information | 0.1429 | -0.1476 (-0.2429 to -0.0571; hurts) | 0.8158 (0.7237 to 0.8947) | 0.6053 (0.4868 to 0.7105) |
| Knowledge | 0.9054 | 0.5405 (0.4189 to 0.6486; helps) | 0.6599 (0.5913 to 0.7222) | 0.8071 (0.75 to 0.8599) |
| Material adverse effect | 0.9467 | -0.0533 (-0.1067 to -0.0133; hurts) | 0.5985 (0.5799 to 0.6178) | 0.8027 (0.7852 to 0.8193) |
| Operating and efforts covenants | 0.7115 | 0.1024 (0.0489 to 0.1607; helps) | 0.6359 (0.5931 to 0.6809) | 0.7562 (0.7216 to 0.7888) |
| Remedies | 0.8867 | -0.1133 (-0.1867 to -0.0533; hurts) | 0.8553 (0.7763 to 0.9342) | 0.8553 (0.7763 to 0.9342) |
By lead question family (tech deals, report split) machine-built
| Family | Kept items | Two-pass agreement rate | R6 recall@5 inside the deal | R7 recall@5 across all agreements | Answer agrees | Agrees or partly |
| Employee equity awards | 307 | 0.9903 | 0.5062 | 0.5905 | 0.0791 | 0.8047 |
| Termination (break-up) fee | 299 | 0.9934 | 0.5633 | 0.6968 | 0.0095 | 0.5619 |
| Employees' pay and benefits after the deal | 280 | 0.9859 | 0.6778 | 0.8346 | 0.0204 | 0.8163 |
Do the two answer keys agree?
The rung order under the lawyers' key and under the machine-built key
| Rung | recall@5, lawyers' key human-labelled (MAUD) | recall@5, machine-built key machine-built |
| R1 | 0.508 | 0.498 |
| R2 | 0.3948 | 0.3765 |
| R3 | 0.5344 | 0.5227 |
| R4 | 0.5296 | 0.5105 |
| R5 | 0.6156 | 0.5777 |
| R6 | 0.6025 | 0.5738 |
The same two-pass procedure that built the tech-deal key was run on MAUD's own questions, so those questions have two keys.
How far the machine-built key can be trusted machine-built
| Measure | Value |
| The machine span overlaps the lawyers' span | 0.9789 (0.9615 to 0.9937) |
| Kendall's tau between the two rung orders | 1.0 (0.7333 to 1.0) |
| Items where both machine passes agreed | 331 |
| Items asked | 555 |
| Agreements | 30 |
This is evidence for or against trusting the machine-built key, not proof.
Answers
Answers to MAUD's questions (report split) human-labelled (MAUD)
| Measure | Value |
| The answerer picks MAUD's answer | 0.6039 (0.5883 to 0.6181) |
| The same, counting only picks whose citations survived the gate | 0.5235 (0.5026 to 0.5415) |
| Always picking each question's most common answer (baseline) | 0.7528 (0.7393 to 0.7665) |
| Answerer minus baseline | -0.1489 (-0.1698 to -0.1286; hurts) |
| Questions scored | 6307 |
| Questions asked | 6321 |
Answer model: claude-haiku-4-5-20251001.
Answers to the tech-deal questions (report split) machine-built
| Measure | Value |
| A judge model finds the answer agrees with the kept machine answers | 0.037 (0.0226 to 0.0533) |
| Agrees or partly agrees | 0.7262 (0.6849 to 0.7681) |
| The judge declined to rate | 128 |
| Questions | 621 |
Judge model: claude-sonnet-5-5. The judge compares each answer with the two kept machine answers.
Declining when the answer is not there machine-built
| Group | Questions | Correct decline | False answer |
| Lead question, both passes found no such clause | 36 | 0.6389 | 0.3056 |
| Earn-out question (one machine pass found none) | 30 | 0.7667 | 0.2 |
| Answer sits in an unfiled schedule (weakest key) | 20 | 0.05 | 0.95 |
| A company name that matches no deal's aliases | 30 | 1.0 | 0.0 |
| A name matching several deals | 30 | 1.0 | 0.0 |
The right answer to each of these is a decline: "not stated in this agreement", "in a schedule that was not filed" or "which agreement?". The two name groups hold by construction: they test the deal resolver, which makes no model call.
Citations
The citation gate
| Questions | Claims returned | Claims kept | Share kept |
| MAUD's questions | 11431 | 9836 | 0.8605 |
| Tech-deal questions | 1562 | 1449 | 0.9277 |
A claim is kept only if its quote occurs word for word in the passage it cites. This check needs no answer key.
A second model tries to refute each kept claim machine-built
| Measure | Value |
| Share of claims it failed to refute | 0.7421 |
| Claims checked | 11150 |
| Unreadable replies | 0 |
Why things go wrong
Why retrieval misses or answers go wrong
| Class | Value | What it counts | Answer key |
| All misses (R6, MAUD) | 1433 | Report-split questions whose gold span is not in the top five passages | human-labelled (MAUD) |
| Wrong section | 810 | The gold span sits in another section | human-labelled (MAUD) |
| Right section, definition missing | 134 | The section was found but the definition it depends on was not | human-labelled (MAUD) |
| Right section, wrong passage | 451 | Another passage of the right section was returned | human-labelled (MAUD) |
| Wrong deal | 0 | Tech-deal questions that R7 scoped to the wrong agreement | machine-built |
| Answer in an unfiled schedule | 0.95 | Share answered anyway when the answer sits in a schedule that was not filed (weakest key) | machine-built |
| Label disputed | 0.0533 (0.0189 to 0.0966) | Share of sampled misses where a model judged the first result to answer the question | machine-built |
| Superseded text | not classified | Amended passages are shown with their amendment; misses on them were not counted apart | |
Against the published baselines
Against LegalBench-RAG's published MAUD baselines (all agreements, character recall, percent)
| Method | recall@1 | recall@2 | recall@4 | recall@8 | recall@16 | recall@32 | recall@64 | Answer key |
| LegalBench-RAG: naive fixed-size chunks | 2.54 | 3.12 | 4.53 | 8.75 | 13.16 | 18.36 | 25.62 | published |
| LegalBench-RAG: recursive text splitter | 1.65 | 2.09 | 4.59 | 6.18 | 12.93 | 21.04 | 28.28 | published |
| LegalBench-RAG: recursive splitter and Cohere reranker | 0.52 | 2.48 | 4.39 | 7.24 | 14.03 | 22.6 | 31.46 | published |
| This project, R1 | 0.19 | 0.45 | 1.0 | 1.7 | 3.72 | 6.56 | 11.92 | human-labelled (MAUD) |
| This project, its best rung | 0.34 | 0.66 | 1.37 | 2.69 | 4.65 | 8.16 | 14.87 | human-labelled (MAUD) |
Published figures: Pipitone and Houir Alami, LegalBench-RAG, arXiv:2408.10343v1. The setups differ (chunking, query wording, and k counted in their chunks against our passages), so read this as indicative, not head to head.
The live desk
The live desk
| Measure | Value |
| Answer model | claude-haiku-4-5-20251001 |
| Price per million input tokens, US dollars | 1.0 |
| Price per million output tokens, US dollars | 5.0 |
| Prices checked on | 2026-10-05 |
| Hosting per month, US dollars | 7.09 |
| Model budget per month, US dollars | 2.91 |
| Model budget per day, US dollars | 0.29 |
| Input tokens per answer, mean | 4250.3 |
| Output tokens per answer (including thinking), mean | 1325.7 |
| Cost per new answer, mean, US dollars | 0.010879 |
| New answers the monthly budget covers | 267 |
| Live-path retrieval inside one agreement (R6n), development machine, ms p50 | 46.4 |
| Live-path retrieval inside one agreement (R6n), development machine, ms p95 | 365.6 |
| Live-path retrieval across all agreements (R7n), development machine, ms p95 | 516.66 |
| Search on the server, ms p50 | 76.6 |
| Search on the server, ms p95 | 120.9 |
| A new answer on the server, ms p50 | 17225.8 |
| A new answer on the server, ms p95 | 24998.1 |
| A cached answer on the server, ms p50 | 77.0 |
| Peak memory of the service, MB | 528.7 |
| Index file, bytes | 634859520 |
| The cap trips: a new question is refused | yes |
| The cap trips: a cached answer is still shown | yes |
Token counts come from a calibration sample of 40 questions sent to the API. Server timings leave out the network.