What AI Vendor Due Diligence Actually Means
AI vendor due diligence is the process of deciding whether a third-party artificial intelligence product, model, service, or data supplier is suitable for a defined business purpose. It is not a single security questionnaire or a general review of the vendor’s marketing claims. The buyer should examine the system’s intended uses, training and retrieval data, subprocessors, deployment method, monitoring, human oversight, incident reporting, contractual remedies, and exit options. A tool that performs well in a demonstration can still create material risk if it makes decisions about customers, employees, credit, compliance, health, or public benefits. The appropriate depth depends on the consequence of error, the sensitivity of the data, the degree of automation, and the organization’s regulatory obligations. In 2026, due diligence should therefore be treated as a documented risk decision, not procurement paperwork.
Also worth reading: How do I conduct an AI legal services broker comparison to find the right platform for my business? · What is an AI vendor due diligence checklist and how should legal teams use it in 2026? · How Should Legal AI Risk Tiers Guide Law Firm and Business Decisions in 2026?
Regulators increasingly expect companies to understand how AI is used in their own operations, even when a vendor supplies the model. Financial institutions, for example, may face expectations associated with model risk management, third-party risk management, consumer protection, privacy, cybersecurity, and anti-money-laundering controls. The European Union’s AI Act also classifies uses by risk rather than merely by the technology used, making a human-resources application more sensitive than an internal writing assistant. Vendors may provide evidence and controls, but the customer normally remains responsible for selecting the use case, configuring the system, supervising its outputs, and deciding what to do when results are wrong. Due diligence is consequently both a vendor-selection exercise and the beginning of ongoing governance.
A Risk-Based Due Diligence Method
A sound process begins by defining the proposed use before reviewing product features. The business should identify who will use the system, what decisions it will influence, whether a person can meaningfully review those decisions, and what happens when the system is unavailable or produces unreliable output. Data classification should follow: public information, confidential business information, personal information, regulated information, and information subject to contractual restrictions may require different contractual and technical protections. Buyers should also map the full service chain, including cloud hosts, model providers, software components, retrieval databases, monitoring tools, and any human-review providers. A locally operated interface does not necessarily make the entire service local if it sends prompts to a remote API, retains telemetry, or depends on externally managed models.
The evaluation should then test the vendor’s claims through evidence rather than assurances. Depending on the use, this may include penetration-test summaries, independent assurance reports, model cards, system cards, data provenance records, evaluation results, bias testing, disaster-recovery evidence, and incident metrics. A questionnaire asking whether a product is “secure” or “fair” is too vague to support a decision. More useful questions specify which frameworks or controls apply, when they were tested, what population was assessed, which languages and demographic groups were included, and what limitations were found. Organizations should require vendors to distinguish between pre-deployment testing and continuing production monitoring because performance can change after integrations, data updates, model changes, or changes in user behavior.
Legal, Regulatory, and Ethical Review
Contract review is central because technical controls can be changed, suspended, or made ineffective by the vendor’s legal position. The agreement should identify the controller, processor, or other legal roles; state the permitted purposes; prohibit model training on customer inputs without a separate lawful and informed authorization; and define retention, deletion, location, and government-request policies. Liability provisions should address security breaches, intellectual-property claims, confidentiality failures, regulatory penalties, incorrect outputs, and third-party claims. The customer should also have audit or assurance rights, notice of material changes, incident deadlines, cooperation obligations, business-continuity support, and a workable termination and data-export process. For higher-risk systems, the contract should connect the vendor’s obligations to the customer’s actual governance process rather than promising generic compliance.
AI-specific legal and ethical concerns require separate analysis. A vendor may claim that its system is unbiased, but testing may cover only one population, language, or decision threshold. If the system screens applications, detects fraud, predicts credit risk, or recommends employee actions, the buyer must assess disparate impact, explainability, contestability, and the availability of human review. Human review is useful only when reviewers have authority, training, time, and information to depart from the recommendation. The Palantir dispute referenced in current due-diligence discussions illustrates that criticism can extend beyond privacy to human-rights impacts associated with a vendor’s customers and contracts. Similarly, AI notetakers can create confidentiality, recording-consent, privilege, and sensitive-processing concerns even when the transcription itself is accurate. The key question is not whether AI is ethical in the abstract, but whether the vendor’s practices and the proposed deployment can be governed responsibly.
Comparing the Main Vendor Models
AI vendors can be evaluated as managed cloud services, enterprise-controlled deployments, or narrower specialist tools, and the risk profile of each option is different. Buyers should compare commercial models, operating models, data handling, assurance, and exit arrangements instead of treating them as interchangeable products. The table below highlights the main operational questions that should inform a pilot and contracting decision.
| Feature | Option A: Managed AI cloud service | Option B: Enterprise-controlled deployment |
|---|---|---|
| Operations | Vendor hosts and updates the service | Customer or specialist firm controls infrastructure and updates |
| Data exposure | Inputs and telemetry may leave the customer environment | Sensitive workloads can remain under customer control |
| Administrative burden | Lower for the buyer | Higher because the customer operates controls and integrations |
| Customization | Usually limited to configuration and approved interfaces | Greater control over models, retrieval, logging, and deployment |
| Evidence | Vendor documentation and assurance reports support review | Customer must evaluate more dependencies itself |
| Typical pricing | Subscription, API, or per-seat usage fees | Setup, infrastructure, engineering, and support costs, sometimes plus vendor fees |
| Main concern | Provider dependency, data use, and changing vendor controls | Internal capability burden and configuration errors |
| Best fit | Lower-risk, fast deployments with acceptable contractual protections | Sensitive, regulated, or high-consequence workloads |
Turning Findings Into a Procurement Decision
The practical process is to create an evidence file that another reviewer could understand without relying on sales conversations. The file should contain the intended use, risk classification, data-flow diagram, vendor and subprocessor inventory, security and privacy materials, model or system documentation, test results, contractual positions, identified gaps, remediation commitments, and the accountable business owner. A pilot should use representative data and realistic workflows, not only a curated demonstration. The team should measure accuracy, false-positive and false-negative rates, latency, uptime, explainability, reviewer override rates, and performance across relevant populations. It should also simulate unavailable systems, incorrect outputs, vendor personnel changes, data-export requests, and contractual termination.
Findings should be graded by both likelihood and business consequence. A missing assurance report may be acceptable for a low-impact use if the system is isolated and the contract imposes clear controls; the same gap may justify rejection for a system involved in credit, employment, or regulatory reporting. Organizations can use red, amber, and green status categories, but each amber item should have a named owner and deadline. High-severity issues should be closed before production, formally rejected as outside scope, or accepted by an authorized executive with a documented rationale. This prevents a long due-diligence process from ending without a decision and makes it harder for attractive product features to compensate for unmanageable legal or operational exposure.
Common Mistakes and Cost Considerations
One common mistake is asking for an AI certification and treating the badge as proof that the product is fit for every purpose. Certification or assurance can support a decision, but no single label substitutes for testing the intended use, reviewing the contract, and understanding the data. Another mistake is assuming a vendor’s compliance with privacy or security law establishes that its model is accurate, fair, or legally usable for a particular decision. Buyers also frequently overlook subcontractors, prompt retention, human-review services, model updates, and the fact that a vendor may change its processing practices after the initial review. Finally, many organizations conduct extensive diligence but fail to assign ongoing ownership after contract signature.
Costs vary widely because the relevant expense is not limited to the vendor’s quote. A managed chatbot may cost a few dollars to a few hundred dollars per user per month, while API systems are often priced by tokens, calls, documents, audio minutes, or consumed resources. Enterprise deployments can add implementation, infrastructure, security review, integration, evaluation, training, and support costs; specialized assurance or legal work may also be required. The low sticker price of an AI tool does not make it economical if it creates manual review, compliance work, incident response, data loss, or erroneous decisions. Conversely, an expensive enterprise contract does not justify weak controls. Before selecting a model, buyers should calculate total cost of ownership over at least the initial contract term and include expected review labor, storage, monitoring, model upgrades, and exit expenses.
When to Act and What to Ask Next
A business should begin before signing a pilot agreement if the system will process confidential information, interact with customers, influence employment or financial decisions, or become part of a regulated activity. It should also act when an existing vendor materially changes its model, training practices, subprocessors, infrastructure location, or incident history. A short reassessment is insufficient where the change affects risk, so the organization should determine whether revalidation, revised contract terms, customer notice, or a new approval is required. The review cadence should be risk-based: high-impact systems may need quarterly performance and incident reviews, while lower-risk tools may be reviewed annually, with event-driven reviews after material changes. The key point is that due diligence does not end when the purchase order is signed.
For mid-market companies, the most efficient approach is often a staged engagement. Start with a limited pilot, use minimized and representative data, require contractual protections before uploading sensitive material, and define measurable exit criteria. Escalate the review when a system moves from drafting or search into decisions about people, money, safety, or legal rights. If internal expertise is limited, an independent legal, privacy, cybersecurity, or model-risk specialist can help test assumptions, but management must retain responsibility for the decision. An AI legal services broker may be useful for comparing vendors and coordinating specialist review, though it should disclose compensation and conflicts and should not replace the customer’s own governance. Due diligence is best treated as a repeatable business discipline that combines technical evidence, legal accountability, commercial judgment, and a clear decision record.