# What Should Legal Teams Require When Buying AI in 2026?

Natalie Fletcher · September 26, 2026

> Direct Answer to the AI Procurement Question A legal AI procurement checklist should test whether a product can be governed, not merely whether it...

## Direct Answer to the AI Procurement Question

A legal AI procurement checklist should test whether a product can be governed, not merely whether it performs well in a demonstration. At minimum, a buyer should evaluate the vendor’s data rights, security controls, confidentiality protections, audit evidence, model-development practices, incident duties, subcontractor chain, output ownership, termination rights, and regulatory responsibilities. The same review must account for the legal team’s intended use: a system that summarizes public legislation carries a different risk profile from one that recommends contract language, reviews privileged documents, or participates in a client’s vendor-selection process.

**Also worth reading:** [What Does an Enterprise Legal AI Compliance Framework Actually Require in 2026?](https://lawr.io/knowledge/what_does_an_enterprise_legal_ai_compliance_framework_actually_require_in_2026.php) · [How Do Organizations Evaluate Legal AI Vendors Without Buying the Wrong Tool?](https://lawr.io/knowledge/how_do_organizations_evaluate_legal_ai_vendors_without_buying_the_wrong_tool.php) · [How Should Companies Diligence an AI Legal Services Broker Before Buying AI Deal Workflows?](https://lawr.io/knowledge/how_should_companies_diligence_an_ai_legal_services_broker_before_buying_ai_deal_workflows.php)

A defensible process generally requires four approvals: business ownership, information security, privacy, and legal or compliance. Higher-risk uses may also need model-risk, records-management, ethics, procurement, and executive review. Rather than assigning arbitrary scores, decision-makers should document each material risk and require an accountable owner to accept it. “The vendor says it is enterprise-ready” is not an answer; the relevant question is what evidence demonstrates that claim and what happens when the evidence stops applying.

The governing standard should be proportionality. A 5-user tool summarizing public regulations does not need the same review as an enterprise system processing 500,000 documents or interacting with customers. Nevertheless, a low-risk designation should never conceal use-case expansion. If users can connect email, customer files, or other systems after approval, the contract and review should expressly address those new connectors. The best checklist therefore functions as a repeatable control framework, not as a one-time form completed immediately before signature.

## Establishing Scope, Risk, and Decision Rights

Begin by naming the exact use case in ordinary operational language. “Legal AI” can include legal research, due diligence, contract review, matter management, timekeeping, document translation, internal knowledge search, compliance monitoring, and autonomous workflow execution. Each category creates different exposure. Legal research may introduce erroneous authority or outdated law; contract review may alter obligations; document translation may reveal confidential information; and autonomous execution may bind the company without meaningful human review.

Risk should then be classified across several dimensions rather than reduced to one label. Consider the sensitivity of the data, the degree of human supervision, the scale of deployment, the number of affected people, the possibility of legal decisions, and the severity of foreseeable harm. A relevant threshold is whether the system can influence a contract, payment, employment, regulatory, litigation, or public-benefit decision. Public-sector buyers must additionally follow applicable procurement rules, while regulated industries may face sector-specific controls.

Decision rights need explicit names and backup owners. The business sponsor should own the benefits and ongoing use; security should approve technical controls; privacy should approve personal-data processing; and legal should review contractual and regulatory terms. A risk committee can resolve disagreements, but it should not become a place where unresolved concerns are merely noted. Approval should expire or be revisited after a material change, such as a new model family, acquisition, processing location, subprocess, use case, or connection to an operational system. A common practical trigger is review at least every 12 months, with immediate review after a serious incident.

## Data Rights, Confidentiality, and Legal Holds

The data provision should define what enters, what is retained, and what can be learned from that data. Buyers should reject ambiguous assurances that customer information is merely “used to improve services.” Permitted uses need a clear boundary, including whether prompts, retrieved documents, feedback, telemetry, embeddings, and derived outputs can train or evaluate any model. A contractual prohibition should also cover selling data, combining it with other customers’ information, or using it to develop models for unrelated products without specific permission.

The vendor should disclose the full data path: hosting regions, encryption methods, tenant separation, backup locations, support-access arrangements, and material subprocessors. The procurement team should ask for independent assurance reports where available, such as SOC 2 or an equivalent control framework, ISO 27001 certification, penetration-test summaries, and vulnerability-management information. These artifacts are evidence, not substitutes for contract rights. Their scope, period, exceptions, and reliance limitations should be checked before they are accepted.

Privilege and work-product protection require specific contractual language. A vendor may provide a product while retaining no ownership of customer data, but that alone does not resolve waiver, confidentiality, or compelled-disclosure questions. The agreement should limit use for service delivery, impose confidentiality duties on personnel and subprocessors, define retention and deletion periods, and provide for legally required holds. Searchability matters as well: a system that indexes matter files should support ethical walls, permission inheritance, access revocation, and export where the organization needs to retrieve information.

| Feature | Conventional Legal AI Subscription | Custom or Model-Assisted System | Broker-Led Procurement Support |
| --- | --- | --- | --- |
| Typical ownership | Legal, security, and privacy approve the vendor and configuration | Business unit funds development while legal, security, architecture, and procurement share risk | Independent broker structures requirements, demos, evidence requests, and negotiations |
| Best fit | Standard research, summarization, drafting, or intake workflow | Organization-specific models, integrations, or high-volume document processes | Multi-vendor evaluation, unclear use cases, or regulated enterprise adoption |
| Cost structure | Usually per user, tier, or usage band, often requiring annual commitment | Upfront design, data preparation, integration, testing, hosting, and ongoing operation | Often paid as a fixed project, retainer, success fee, or combination; scope must be agreed in writing |
| Primary advantage | Fast deployment and mature vendor support | Greater customization and potential workflow fit | Comparability and reduced dependence on vendor sales claims |
| Main limitation | Shared-platform features may not match legal workflows | Long delivery times, specialist talent needs, and uncertain performance | Adds a procurement layer and does not replace the buyer’s legal accountability |

## Security, Model Behavior, and Human Oversight
Security diligence should follow the actual data architecture, not a generic feature list. Review single sign-on, role-based access, multifactor authentication, encryption in transit and at rest, key management, logging, disaster recovery, backup testing, tenant isolation, and secure development practices. Ask how quickly the vendor remediates critical vulnerabilities, how customers are notified, and whether the vendor maintains a coordinated vulnerability-disclosure program. Contract language should connect reasonable security commitments to breach notification, investigation, evidence preservation, cooperation, and liability.

AI-specific diligence requires details about the model and retrieval system. Buyers should learn which models answer prompts, whether those models change without notice, how retrieval citations are generated, and what prevents one customer’s information from affecting another customer. The vendor should explain its approach to hallucination, stale sources, conflicting authorities, prompt injection, malicious documents, sensitive-data extraction, and unauthorized tool calls. If the system can send email, modify documents, execute code, or call external services, the review must treat those as high-risk capabilities rather than convenience features.

Human oversight should be designed into the workflow. A reviewer needs enough time, authority, and information to challenge the result; a generic statement that every output is subject to “human review” is weak if reviewers routinely accept suggestions at scale. Policies should specify when double review is required, which outputs cannot be sent without verification, and how user overrides are recorded. For legal citations, the team should test whether links open, quotations match the source, later authority is identified, and jurisdiction is correct. Accuracy claims should be measured on representative tasks, with a defined sample size, baseline, success rate, and owner for remediation.

No universal accuracy percentage is credible without a defined test. A vendor claiming 95% accuracy on document classification may have used a balanced, clean sample, while 85% on a consequential recommendation may still be inadequate. Procurement should therefore require a trial using the buyer’s documents and scenarios, subject to confidentiality safeguards. The pass threshold should reflect harm and reversibility: a low-reward classification task may tolerate more errors than a workflow that files evidence or changes contract positions.

## Contract Clauses, Liability, and Exit

The contract must allocate responsibility for decisions made with AI assistance. The vendor should warrant applicable law, authorized use of materials, security controls, and its contractual obligations, while the customer should remain responsible for approved deployment, instructions, supervision, and downstream use. Warranties should not be undermined by an acceptance clause that transfers every model error to the buyer. If the vendor supplies a draft, it should stand behind the promised service, support, and remedies even if a human later edits the result.

Indemnities deserve careful treatment because AI liability terms are not uniform. A buyer may seek protection for third-party claims involving intellectual-property infringement, confidentiality breaches, data misuse, or vendor-caused security incidents, subject to exclusions and procedural conditions. The parties should address whether the provider is liable for regulatory fines, remediation costs, recall or correction expenses, professional fees, and business interruption. Liability caps should be evaluated against plausible exposure, not only annual fees; a low subscription price does not make severe contractual exposure harmless.

Change control is essential because vendors may alter models, infrastructure, or subprocessors. The agreement should identify material changes, provide advance notice where possible, give the customer termination or fee-adjustment rights, and prohibit degradation of security or service. Renewal should require current evidence rather than allowing indefinite auto-renewal without review. Auto-renewal deadlines, notice periods, price escalators, minimum commitments, and overage rates should all be recorded in the procurement record.

Exit planning should cover data portability and deletion. The contract should define export formats, retrieval responsibilities, assistance during transition, deletion verification, backup expiration, and treatment of derived data such as embeddings. Customers should know whether they can retain search indexes or evaluation artifacts after termination. A tested export is more valuable than a promise that the vendor will “help” without specifying timing, format, scope, or cost. For mission-critical workflows, exit readiness should be tested before production, not after a dispute or acquisition.

## Regulatory, Ethical, and Sector-Specific Review

The compliance owner should map the system to every applicable legal regime as of the actual launch date. For the European Union, the AI Act entered into force on 2 August 2024, with obligations for general-purpose AI models applying from 2 August 2025 and most remaining provisions applying from 2 August 2026, subject to the statute’s detailed provisions and guidance. This timetable makes 2026 a significant review point, but classification should not be based on the word “assistant” alone. A buyer should assess provider and deployer roles, prohibited practices, transparency, human oversight, data governance, and any sectoral law.

In the United States, the federal regulatory position remained subject to policy, litigation, and agency action in 2026, so buyers should avoid relying on a static list of rules. They should track executive orders, agency guidance, sectoral requirements, state laws, and procurement terms, and should obtain current advice for high-impact deployments. Public procurement may also involve accessibility, records, security, competition, and transparency rules. In other jurisdictions, including China, Indonesia, and the United Kingdom, requirements differ, and cross-border data or service delivery can add diligence rather than remove it.

High-impact uses require an escalation path. Employment, housing, credit, insurance, healthcare, education, legal-services, public-benefits, and regulatory decisions may receive heightened scrutiny where AI is used to recommend outcomes or allocate opportunities. The review should consider disparate impact, accessibility, notice, explanation, contestability, and the effect on vulnerable groups. The 18 July 2024 European Union AI Act prohibition on certain uses, for example, is a reminder that general enterprise value does not override use-case restrictions. A system initially limited to assisting employees can still create risk if employees use it to make consequential decisions outside the approved scope.

## How to Run a Practical Procurement Process

A workable process usually runs through six stages. First, the business prepares a one-page use-case and risk description. Second, security and privacy issue requirements rather than merely reviewing a vendor questionnaire. Third, legal defines evidence, contract, and exit criteria. Fourth, a controlled proof of concept compares at least two credible approaches against a fixed test set. Fifth, decision-makers document unresolved exceptions and accept them through named owners. Sixth, legal, procurement, security, and the business sign the final approval before production data is connected.

The comparison should include status quo, an off-the-shelf product, a larger suite with broader functionality, and a custom or broker-supported route where appropriate. The status quo may be less expensive but can leave cycle time and error exposure unchanged. A broad suite may reduce integration work while adding irrelevant features and vendor concentration. A custom system can improve workflow fit but introduce development, maintenance, and model-drift obligations. An independent broker can help structure the evaluation and negotiate across vendors, but it adds a fee and cannot accept ultimate legal responsibility.

Timeline expectations should be realistic. A low-risk, standard product may complete an accelerated review in roughly 4 to 8 weeks if evidence is current and no sensitive data is connected. A sensitive enterprise deployment may take 3 to 6 months, and custom development can take 6 to 18 months or longer. These ranges are planning assumptions, not guarantees. Jurisdictional legal review, security assessment, integration, and vendor due diligence can each become the critical path, so a proposed launch date should not be treated as a reason to omit required approvals.

For a pilot, restrict access, use de-identified or synthetic data where possible, disable unapproved connectors, and define automatic stop conditions. Evaluate factuality, citation quality, latency, accessibility, administrator effort, user behavior, and total operating cost. Record adverse cases rather than only headline averages. The pilot should have a written exit decision: proceed, proceed with specified controls, repeat testing, or stop. Ending a weak pilot early is a sign of sound procurement, not procurement failure.

## Pricing, Common Mistakes, and When to Act Immediately

Pricing varies too much for one responsible market range. Individual legal research or drafting subscriptions may run from tens to several hundred US dollars per user per month, while enterprise platforms can cost from thousands to hundreds of thousands of dollars annually, with implementation and integration added. High-volume extraction, premium models, private hosting, or custom development can increase total cost substantially. Buyers should compare at least 3 years of cost when possible, including setup, data cleaning, connectors, usage overages, training, support, assurance reviews, and exit. Cheapest per seat is not necessarily cheapest per completed legal task.

The most common mistake is purchasing before defining the problem. Another is accepting a polished demonstration that uses the vendor’s preferred documents instead of the buyer’s real work. Teams also fail by treating security certification as proof of model accuracy, by allowing a free pilot to ingest privileged material, by relying on a click-through agreement, and by postponing exit testing until renewal. Executive pressure can be especially harmful when the desired date precedes the evidence; an AI feature that creates a reportable security, privilege, or regulatory event can be more expensive than a delayed launch.

Immediate escalation is appropriate when an incident is suspected, a vendor announces unauthorized training, a material model change is imminent, or users begin using unapproved tools. The same response is needed if the tool can execute external actions, access production records, affect individual rights, or enter material into a legally privileged matter. A buyer should contain access, preserve logs, notify responsible teams, and avoid deleting evidence before the legal-hold position is decided. Escalation should be fast, but conclusions should not be invented: the facts and contract should determine notification, disclosure, and remediation.

## The Procurement Decision Standard

The definitive standard is not whether a product is called “legal AI,” “agentic,” or “enterprise-grade.” It is whether the organization can explain the system’s purpose, evidence its controls, supervise its use, measure its performance, and exit safely. A sound purchase connects each risk to a technical control, contractual right, accountable owner, and verification method. Where the vendor cannot answer material questions, the correct decision may be to delay, narrow the deployment, select another product, or continue with the current process.

For most legal teams, the best first action is to approve a controlled 4-to-8-week evaluation using representative but protected data, two or more credible options, and predefined failure thresholds. High-risk or highly customized systems deserve a longer, multidisciplinary review. The 2026 environment rewards organizations that treat procurement as ongoing governance rather than a vendor-comparison event, because a system that passed review for one model, tenant, or jurisdiction can fail after a technical or legal change.

## Quick answers

### What is the fastest safe way to evaluate legal AI vendors?

Use a controlled 4-to-8-week pilot with representative, protected data and a fixed test set. Compare at least two credible options and predefined thresholds for accuracy, security, user supervision, cost, and exit readiness. Do not connect production systems until the designated business, security, privacy, and legal owners approve the result.

### Does a SOC 2 report prove that a legal AI system is accurate?

No. A SOC 2 report primarily addresses the design and operation of specified controls, not whether legal outputs are factually correct or useful. Buyers still need model-specific testing, retrieval testing, human-oversight design, and contractual commitments covering security, confidentiality, service changes, and remediation.

### Can a legal team permit vendors to train models on client data?

It may do so only if the business, privacy, security, and legal owners knowingly approve the use and the contract provides enforceable restrictions. Silence or a broad improvement-of-services clause is not sufficient. Regulated, confidential, or privileged information often requires a stronger prohibition than ordinary business information.

### What makes an AI legal vendor high risk?

Risk rises when the tool can access confidential records, make recommendations affecting individual rights, execute external actions, or change legal or financial outcomes. Unapproved model changes, weak audit rights, and no deletion or export process also increase exposure. The appropriate review depends on the actual use case rather than the product’s marketing category.

### Should legal teams use an independent broker to select legal AI?

A broker can be useful for multi-vendor comparisons, requirement design, evidence collection, pilots, and negotiation, especially when the buyer lacks in-depth AI or procurement expertise. The buyer remains responsible for legal acceptance and should define fees, deliverables, conflicts, confidentiality, and access to the underlying evaluation evidence before engagement.

Canonical: https://lawr.io/knowledge/what_should_legal_teams_require_when_buying_ai_in_2026.php
Markdown: https://lawr.io/knowledge/what_should_legal_teams_require_when_buying_ai_in_2026.php/index.md
