What AI Vendor Due Diligence Actually Means

AI vendor due diligence is the documented process of deciding whether an external provider of an AI system, model, software platform, data service, or related infrastructure is suitable for a particular business use. It is not a paper exercise or a generic questionnaire about security. The buyer must examine the vendor’s legal ownership of its technology, training and input data, model behavior, output reliability, security controls, subcontractors, hosting arrangements, regulatory duties, insurance, and exit options. The correct standard depends on the use: an internal drafting assistant presents a different risk from a credit-scoring model, healthcare triage tool, automated customer-service system, or AI component used in a safety-critical product. The same vendor may therefore pass one review and fail another. As of 29 September 2026, the prudent practical question is not simply “Does this company use AI?” but “What exactly is being sold, what can it do, what data does it process, and what evidence would allow us to control it?” That framing turns an abstract AI governance discussion into an accountable procurement decision.

Also worth reading: How should businesses conduct an AI agent risk assessment in 2026 to comply with emerging regulations and prevent autonomous failures? · What are the best EU AI Act vendor compliance tools for businesses facing the August 2026 deadline? · What are the key legal AI vendor contract red flags businesses should watch for in 2026?

A useful definition begins with the system’s role, not the vendor’s marketing label. A buyer should distinguish between a model provider, application vendor, cloud host, data supplier, orchestration platform, and reseller. Each party may have different control over model changes, incident reporting, retention, localization, and customer remedies. Contract language matters because the organization that instructs a chatbot, for example, may not be the party operating the underlying model. Regulators and customers increasingly expect buyers to understand that chain. This matters most where personal data enters prompts, retrieval databases, logs, support tickets, telemetry, or improvement datasets. It also matters where an output can affect a person’s employment, credit, insurance, health care, education, or access to services. The review should produce evidence, assigned owners, and remediation decisions rather than an unverified assurance copied from a vendor website.

Why Conventional Vendor Reviews Often Miss AI Risks

Traditional procurement commonly evaluates uptime, price, support, financial stability, and familiar security controls. Those remain necessary, but they do not reveal whether a model can reproduce protected material, infer sensitive attributes, create unlawful decisions, or change materially after deployment. AI systems may also rely on external datasets and APIs that do not appear in the main contract. Open-source components, embedded models, annotation providers, and cloud-computing subcontractors can create dependencies that a customer never directly selects. The risk is therefore operational as well as contractual. A vendor can provide excellent conventional security while still offering weak documentation about model testing, data provenance, bias evaluation, or customer-controlled deletion.

The hidden third-party dimension is particularly important. RSM US LLP has described hidden third-party risks associated with AI, while Finosec has promoted an AI governance module directed at community banks and other financial institutions. NCUA materials on artificial intelligence emphasize that regulated institutions must connect technology adoption with existing risk-management and oversight duties. These sources illustrate a broader point: adopting AI does not transfer the customer’s responsibility to the supplier. A financial institution cannot answer a supervisory concern merely by pointing to a vendor’s marketing claims. It must show which model it uses, how it tested the system, who can override results, what happens when performance declines, and whether records can support examination or customer disputes. The organization remains accountable for how the tool is deployed, even where the vendor controls the code.

AI-specific review also addresses behavior that conventional questionnaires do not measure. A secure system can still generate defamatory statements, expose another customer’s information, violate copyright restrictions, or make inconsistent decisions across demographic groups. Accuracy tests must be connected to the intended purpose and population rather than relying on one vendor-wide percentage. For lower-impact applications, sampling and escalation may be reasonable. For consequential uses, buyers may need independent testing, documented thresholds, human review, periodic recertification, and a right to suspend the service. The cost of that work must be compared with the harm that a weak system could create. “The technology is innovative” is not evidence, and neither is a benchmark produced under conditions unlike the buyer’s own.

The Due Diligence Process: From Use Case to Evidence

The first stage is to define the proposed use in ordinary operational language. The business should identify the decision being supported, the people affected, the data entering the system, the model or provider involved, and the consequences of error. It should also determine whether the AI is advisory, generates content, ranks options, detects fraud, predicts behavior, or makes a final decision. That classification sets the required depth of review. A 5% error rate might be unacceptable in a payment-fraud workflow but tolerable in an optional internal brainstorming tool, provided staff understand the limitation. A procurement scorecard cannot substitute for this analysis. Every material claim should be translated into a testable control, an evidence request, a contractual obligation, or an explicit risk acceptance.

The second stage is to map the technology supply chain. Reviewers should obtain architecture diagrams, lists of material subprocessors, hosting locations, model-development practices, data sources, retention schedules, and descriptions of human review. They should determine whether customer data trains shared or customer-specific models, how long prompts and outputs are retained, and whether administrators can disable those uses. For retrieval systems, the organization must identify the source repositories, ingestion dates, access permissions, and deletion behavior. For third-party models inside a purchased application, the vendor should disclose which components are used and how updates are tested. Many contracts permit service changes without project-specific consent, so change-management terms deserve as much attention as the initial security exhibit. The buyer should require advance notice and, for material changes, a documented right to test or terminate.

The third stage is independent enough to be meaningful. Buyers can request assurance reports, penetration-test summaries, model cards, system cards, data-processing agreements, business-continuity plans, and incident records, but they should verify scope and age. “SOC 2 Type II” is not an AI certification, and a report covering one product or period may say little about the service being purchased. A report that is over 12 months old may also fail to reflect recent changes. Where stakes justify it, the buyer can run prompt-injection tests, privacy tests, accuracy studies, subgroup comparisons, and adversarial scenarios using representative data. Results should be recorded with the test date, model version, configuration, sample size, and known limitations. A 95% result from a 20-case sample is not equivalent to 95% accuracy across 20,000 representative cases.

Legal, Regulatory, and Contractual Review

Legal review should cover more than signature authority. The team must establish whether the vendor warrants ownership or sufficient rights in its code, model weights, documentation, and outputs. Copyright uncertainty can arise from training material, generated text, voices, likenesses, or a vendor’s promise that customers will “own everything” despite undisclosed third-party rights. Warranty language should address accuracy only to the extent that accuracy can be promised, while indemnities may appropriately focus on IP infringement, confidentiality breaches, and vendor-caused unlawful processing. Organizations should avoid assuming that an AI product creates legally binding rights in every output. A disclaimer can allocate some risk, but it does not make an inaccurate or unlawful decision defensible under applicable law.

Data-protection obligations require a separate analysis. Buyers should identify controller, processor, or other roles under applicable privacy law and document the lawful basis for relevant processing. The vendor must provide necessary assistance with access, correction, deletion, portability, and incident response where applicable. International transfers may require a valid transfer mechanism and supplementary safeguards. Sector rules can add stricter duties: financial institutions must consider consumer protection, fair-lending, anti-money-laundering, recordkeeping, and model-risk requirements; healthcare organizations must address protected health information and clinical safety; public bodies may face procurement, equality, transparency, and human-review rules. The report “AI Third-Party Risk Management Under Global AI Regulations,” discussed in JD Supra commentary, reflects the growing view that vendor governance is global rather than confined to one privacy checklist. The exact obligations depend on jurisdiction and use, so legal conclusions should be made by qualified counsel rather than inferred from a general AI policy.

Contract terms should operationalize the review. A strong agreement normally identifies permitted data uses, retention and deletion periods, security requirements, subprocessor notice, audit rights, incident deadlines, service levels, model-change controls, IP allocation, confidentiality, regulatory cooperation, and termination assistance. Buyers may need to prohibit using their data to train a shared foundation model unless expressly approved. They may also need to require disclosure of government demands, export restrictions, model deprecation, and planned migrations. Exit assistance is not theoretical: prompts, embeddings, logs, configuration files, and validation records may be needed to move the system without losing institutional knowledge. A fixed 30-day termination right is of limited value if the vendor retains the customer’s data or cannot export it in a usable format.

Comparison of Due Diligence Options

FeatureFull institutional reviewTargeted application reviewVendor questionnaire only
Best suited forHigh-impact or regulated AIModerate-risk business toolsLow-risk, reversible productivity use
TestingRepresentative validation, subgroup analysis, adversarial testing, security testingFocused accuracy, privacy, and prompt testingReliance on vendor evidence and limited sampling
EvidenceIndependent tests, architecture records, metrics, regulator-ready documentationApplication-specific test results and configuration recordsCertifications, policies, attestations, and sales answers
Human oversightNamed approver, escalation route, appeal or correction processStaff training and case escalationInformal user caution
Contract depthDetailed rights, audit, change control, incident and exit dutiesMaterial risks documented in the agreementStandard terms supplemented only for known concerns
Indicative costRoughly $25,000-$200,000+ per material deploymentRoughly $5,000-$40,000Roughly $0-$10,000 of internal time
Ongoing reviewAt least quarterly for changing systems, or more often by riskMonthly or quarterly monitoring after launchAnnual questionnaire reassessment
These ranges are planning estimates, not market standards, and they vary by model complexity, data sensitivity, integration work, and whether the buyer uses external testers. Full institutional review may cost more than a conventional software assessment, but it can prevent a far larger loss in a consequential deployment. Targeted review is often the sensible middle path when a company embeds a low-impact feature in an established product. Questionnaire-only review is defensible mainly where data is synthetic or non-sensitive, outputs are reviewed before external use, and the application can be switched off quickly. Even there, relying only on unchecked vendor claims is weak. A short walkthrough, basic privacy test, and documented owner can materially improve confidence at little cost.

Common Mistakes and Warning Signs

A frequent mistake is treating model accuracy as the only quality measure. Accuracy can conceal class imbalance: a fraud model predicting “no fraud” in 99% of cases may score 99% accuracy while detecting none of the rare events that matter. Buyers should consider precision, recall, false-positive rates, calibration, drift, and business impact. Another mistake is using one aggregate score for all demographic or operational groups. Algorithmic bias cases have shown that apparently neutral data and objectives can produce discriminatory effects. Where relevant, the organization should compare error rates and outcomes across affected groups, examine proxy variables, and test whether human reviewers can effectively challenge the tool. Fairness tests are not a substitute for lawful decision-making, but they can expose failures before deployment.

Another common error is allowing a vendor to answer away a known limitation. Statements such as “the model is not legally liable,” “the customer is responsible for outputs,” or “hallucinations are inherent” may be true in a contractual sense without resolving the buyer’s exposure. Warnings also arise when the vendor refuses to identify subprocessors, cannot state whether customer data trains models, provides no incident-notification period, or changes models without notice. Buyers should be cautious where a pilot is described as production-ready despite lacking access controls, monitoring, deletion, or documented test results. Pressure to sign a non-standard agreement before a launch date is itself a governance issue. A rushed deployment may have a legitimate business deadline, but risk acceptance should be made by a named person with authority to accept it and should include compensating controls.

The most damaging mistake can be treating pilot performance as permanent. Models, prompts, retrieval sources, user behavior, and data distributions change. A system that performs well in June may degrade after an API update, a new customer segment arrives, or an upstream dataset changes. Continuous monitoring should compare actual results with approved test conditions and trigger investigation when predefined thresholds are crossed. Example thresholds might include a 3-percentage-point decline in a critical recall measure, any confirmed cross-tenant data exposure, a 10% increase in escalation rates, or a 24-hour period with repeated material hallucinations. Thresholds should be calibrated to the use rather than copied mechanically. Monitoring without an owner and response procedure merely creates more reports.

When to Act and How to Budget

Due diligence should begin before contract negotiation, not after a free trial has become embedded in operations. The first trigger is any proposal to send confidential, personal, regulated, or customer-derived data to an external AI service. The second is use in a decision that materially affects an individual, even if a person is nominally “in the loop.” The third is a contract, acquisition, or product launch that embeds AI supplied by another party. Organizations should also act when a vendor announces a major model change, new subprocessor, new hosting region, or changed data-retention policy. Waiting for a public enforcement action or customer complaint is too late. A short internal deadline—often 10 to 20 business days for a low-risk tool—is useful, but high-risk deployments need deeper testing and may require 60 to 120 days.

Budgeting should cover more than license fees. Costs include integration, data preparation, legal review, security testing, model validation, employee training, monitoring, insurance, procurement effort, and eventual migration. A low per-seat price can be poor value if staff must manually correct outputs or if sensitive data must be isolated in a separate environment. Conversely, a costly platform may still be economical if it supplies auditable logs, configurable retention, role-based access, evaluation tools, and contractual change control. Before purchase, the organization should calculate a one-year total cost of ownership and a three-year cost including plausible usage growth. It should also price failure: manual review, incident response, customer remediation, regulatory defense, and replacement. As of 29 September 2026, AI legal-services brokers can help compare contract positions, supplier claims, and risk allocations, but they should remain independent of vendor commissions and disclose those relationships.

Legal-services procurement is one category, not the whole control. Buyers may need a security architect, privacy counsel, data scientist, compliance specialist, domain owner, and business sponsor. A broker or law firm can organize requests and negotiations, particularly where several vendors offer overlapping platforms, but it should not certify technical performance without qualified testing. The organization should decide who owns the risk, who approves the deployment, and who can stop it. Contract review alone cannot determine whether retrieved data is accurate or whether an AI answer is reliable in context. Conversely, a technically successful pilot does not solve defective rights, unclear warranties, or unlawful processing. Effective AI vendor due diligence combines legal, security, scientific, operational, and commercial evidence.

A Defensible Decision Standard

The strongest outcome is a written decision package showing that the intended use was classified, the supply chain was mapped, material claims were tested, legal duties were allocated, and residual risk was accepted by an authorized owner. The package should name the exact product, model version where available, configuration, data categories, evaluation dataset, dates, metrics, limitations, and review frequency. It should explain why a 90% score, zero reported incidents, or vendor certification was sufficient—or insufficient—for that particular use. A complete process may reject the tool, approve a limited pilot, require contractual changes, impose stronger access controls, add human review, narrow the use, or defer deployment. There is no virtue in approving every AI purchase, just as there is no virtue in rejecting useful technology without analysis.

Organizations should update the process at least annually and after material changes, with more frequent monitoring for production systems. Evidence can become stale quickly. A supplier’s financial condition, security posture, ownership, regulatory position, and model behavior may all change. Standards such as the NIST AI Risk Management Framework provide a useful structure for governance, mapping, measurement, and management, while sector guidance and applicable law define mandatory duties. The framework does not replace those duties, but it can help a company document them consistently. For mid-market companies, the key is proportionality: a 10-person business does not need the same testing volume as a national bank, but it still needs clear ownership, factual vendor evidence, and controls proportionate to the data and decisions involved. That is the defensible standard for AI vendor due diligence in 2026.