Contract Review Comparison

Contract Review Comparison

An independent evaluation of AI-powered contract review tools

Introduction

Contract review dominates the workload of in-house legal departments — enterprise teams process thousands per year. Legal leaders know that AI could help ease the burden. The AI-powered contract analysis market is now worth $4.3 billion and projected to reach $12.06 billion by 2030. But the category is now so crowded that evaluating vendors has become difficult. Our research found that over 50% of respondents struggle to objectively compare AI tools across vendors.

We created an independent benchmarking study to answer three questions:

We assembled a panel of three senior attorneys to evaluate contract review output from Ivo, Claude for Word (Opus 4.6), and a Practicing Special Counsel at an Am Law 25 firm.

Ivo outperformed Claude by a significant margin, and its performance was comparable to the accomplished human attorney.

Methodology

Ivo conducted a benchmarking study comparing a purpose-built legal AI tool to a general AI tool and a human for redlining tasks.

The Participants

The Contracts

19 real, anonymized contracts were reviewed, spanning NDAs, MSAs, and DPAs.

The Judges

The outputs were judged by three attorneys with recent experience either working with an Am Law 100 firm or serving as in-house counsel at technology companies. Every output was stripped of identifying information and scored blind across five different criteria for contract review.

The Scoring

Outputs were judged on a scale of 1-10. Final scores represent the mean across all three judges.

Download the report

Overall scores

Ivo’s performance was nearly indistinguishable from a human lawyer and significantly outperformed Claude on the five judging criteria.

Tool Score
Ivo 4.52
Claude for Word 3.50
Human Attorney 4.56

Key learnings

Purpose-built legal AI cannot be replicated by general AI.

Specialist teams have spent years crafting the prompts, logic, and outputs that create redlines comparable to top performing senior lawyers. General AI tools cannot yet compare to the performance.

The largest score delta was in Surgical Redlining and Judgment.

Ivo’s team has spent a great deal of time perfecting its surgical redlining abilities. In addition, Ivo achieved the highest scores in legal judgment, even higher than the human attorney. Finally, Ivo excelled at analyzing complex contracts with complicated transactions.

Ivo’s output is broadly comparable to a senior practicing attorney at a highly regarded law firm.

The human attorney and Ivo had very similar scores, suggesting Ivo’s output was comparable to a high-performing senior lawyer. However, the attorney completed their redlining tasks in 10 hours, whereas Ivo’s average performance was 2 minutes and 45 seconds and Claude for Word’s was 4 minutes and 53 seconds.

Average time for each participant to finish reviewing a document

Tool Time
Ivo 2m 45s
Claude for Word 4m 52s
Human Attorney ~32m

Issue Spotting analysis

Did the author spot all the issues within the scope of the playbook? Did the author over-spot issues that may not apply to this contract?

Playbook position

The playbook prescribes California as the preferred governing law, with Delaware and New York as acceptable alternatives. Binding arbitration (JAMS or AAA) is the preferred dispute mechanism, preceded by a 15–30 day good-faith escalation step. Governing law should never be silent, and litigation in the counterparty's home jurisdiction should be avoided.

Tool Score
Ivo 8
Claude 6
Human Attorney 5

Evaluation

This redline kept Delaware law but replaced litigation with JAMS arbitration in Wilmington, Delaware, added a 15-day good-faith negotiation step, and preserved an equitable relief carveout. This most precisely matches the playbook's preferred structure of binding arbitration with a senior-leadership escalation step before arbitration.

Surgical Redlining analysis

Did the author make minimal, precise changes to fix the issue, or did they aggressively rewrite the whole paragraph?

Playbook position

Neither party should assign without consent, except to affiliates or in connection with M&A. Require 30-day advance written notice. Restrict assignment to direct competitors.

Tool Score
Ivo 7
Claude 2
Human Attorney 5

Evaluation

Made two precise insertions into the existing sentence without deleting or restructuring any original text: (1) inserted the affiliate assignment right the playbook calls for, and (2) appended the written notice requirement and the direct competitor restriction as a trailing proviso. The original sentence structure, the M&A carve-out language, and all existing terms were preserved exactly as written.

Formatting Retention analysis

When the author makes changes, did they maintain font, spacing, paragraph numbering, and cross-references flawlessly in Word? Did they respect and correctly capitalize defined terms specific to this document?

Evaluation

This redline inserts the CCPA provision as a standalone clause, structured with sub-clauses. It preserves the document's formatting conventions such as the section heading being bold and in all caps, and the body text being left-aligned. However, it omits a section or sub-section number.

Commenting analysis

Did the author adhere to the playbook rules and add all the approved comments necessary?

Evaluation

Ivo attached the standard language for external comment to the appropriate location. The tone is professional and the message is clear.

Judgment analysis

Did the author pick the right position for each issue when there are different fallback options? When some playbook rules are ambiguous and strict compliance may harm the party’s interest, did they make judgments that are the best for the party?

Evaluation

Ivo left the Wisconsin governing law clause entirely untouched. This reflects correct application of the playbook, which permits any reasonable US state where the counterparty has a nexus, and the counterparty in this agreement is a Wisconsin company.

Conclusion

Ivo performed comparably with a very accomplished human lawyer and outperformed Claude for Word on every metric. Underneath the headline are some interesting patterns. It’s exciting that the scores achieved by both Ivo and the human are so close. It took the human lawyer approximately 10 hours to complete the review of all the documents whereas Ivo’s average time to review a contract was 2 minutes and 45 seconds. Ivo performs a deeply manual and lengthy task nearly identically to a human lawyer in a fraction of the time.