You Are Accountable for AI You Did Not Build
Assuring Third-Party and Vendor AI Systems
There is a governance gap sitting at the center of most enterprise AI programs, and the majority of organizations have not fully reckoned with it.
When an AI system makes a consequential error, the organization that deployed it is held accountable. Not the vendor. Not the foundation model provider. Not the API platform. The institution whose name is on the decision, the customer interaction, or the regulatory filing is the one that answers to affected parties, regulators, and the public. Yet most organizations are deploying AI capabilities they did not design, cannot fully inspect, and have only a limited ability to monitor in production.
This is the third-party AI problem, and it is fundamentally different from traditional vendor risk. When you outsource payroll processing or cloud storage, the vendor’s system either works or it does not. The failure modes are largely predictable, and the accountability is relatively clear. AI is different. A vendor’s model can work correctly on average while producing systematically biased outcomes for specific populations. It can perform well during evaluation and degrade after deployment as data distributions shift. It can generate outputs that are plausible, fluent, and completely wrong. And when any of these things happen in your deployment, you are responsible.
The stakes are substantial and growing. Many third-party vendors, such as providers of off-the-shelf software, have begun embedding AI into their products, often without full visibility into or understanding of their customers. Service providers are incorporating AI into their workflows whether you have asked them to or not. When third-party AI tools are introduced, you are extending your risk exposure deep in the supply chain. On the surface, it may look like you have one product and therefore one vendor, but there are typically multiple parties behind the scenes providing capabilities or data.
This post addresses the practical challenge of governing AI you do not control. That means vendor questionnaires that actually surface meaningful information, contract clauses that protect your organization and create real accountability, shared responsibility frameworks that close gaps rather than document them, and evaluation approaches that give you genuine insight into external models and APIs before and after you deploy them.
Why Traditional Vendor Risk Management Falls Short
Most organizations have mature third-party risk management (TPRM) programs. They issue security questionnaires, review SOC 2 reports, conduct periodic audits, and monitor contractual compliance. These programs were designed for a world in which third-party risk focused primarily on data security and operational continuity.
Traditional tools for managing vendors were not built to address the challenges AI can raise, such as model training, bias mitigation, and data lineage controls. Without updated controls and AI-specific visibility into these vendors, enterprises risk falling behind emerging regulations and stakeholder expectations.
The failure modes of AI systems do not map well onto traditional risk frameworks. Consider what you need to know about a vendor’s AI system that a standard security questionnaire cannot tell you. How was the training data collected, and does it represent the populations your deployment will affect? What bias testing was performed, using which fairness metrics, and against which demographic groups? How does the model perform on edge cases that differ from its training distribution? Has the model been red-teamed for adversarial inputs? What happens to your data after it enters the vendor’s system? Is it being used to improve the vendor’s model? These questions require entirely different assessment instruments and different contractual protections.
According to Prevalent’s 2024 Third Party Risk Management study, 60% of organizations experienced a third-party breach in the past year, highlighting the growing complexity of third-party risks. Despite this, many still rely on outdated tools, such as spreadsheets, for management. If organizations are still catching up on traditional third-party risk, most are even further behind on AI-specific risk.
The regulatory landscape is adding urgency to this challenge. The EU AI Act establishes what it calls a “deployer” category, which is the organization that puts an AI system to use in its operations, and imposes direct compliance obligations on deployers regardless of whether they built the underlying model. Specifically, a deployer may want the developer to warrant that it responsibly developed the AI system, including through appropriate data governance, risk mitigation measures, documentation and instructions for use, transparency notices, cybersecurity, and testing for algorithmic discrimination, accuracy, and robustness. You cannot outsource compliance accountability to your vendor. You can only negotiate for the contractual protections and information rights that allow you to discharge your own obligations.
Building an AI-Specific Vendor Questionnaire
A well-designed AI vendor questionnaire does more than collect documentation. It signals to vendors that you are a sophisticated buyer with robust governance standards, which, in turn, influences vendor behavior. It surfaces the information you need to make risk-based procurement decisions. And it creates the baseline against which you can measure vendor compliance throughout the relationship.
Your questionnaire should be tiered based on AI risk level. A vendor providing a recommendation engine for internal knowledge management warrants different scrutiny than a vendor providing an AI system that will influence credit decisions or clinical recommendations. The key is to apply proportionate diligence.
Model provenance and training data
Begin by understanding what the model is built from. Ask vendors to describe their training data sources, including the approximate volume and vintage of data, the domains they represent, and any known limitations. Specifically, ask whether the training data included any data from your industry, your company, or your customers. Understand whether the vendor built their own foundation model, fine-tuned an existing model, or is essentially a wrapper around a third-party foundation model. If the latter, you have a fourth-party risk situation that requires its own assessment.
Bias and fairness testing
Ask vendors to describe the fairness metrics they evaluated during development, the protected characteristics they tested, and the disparity thresholds they used. Critically, ask for evidence, not just assertions. A vendor that says it conducts bias testing should be able to provide documentation of the testing methodology, test results, and descriptions of any issues found and how they were addressed. Ask whether they have conducted third-party bias audits and whether those reports are available for customer review.
Performance documentation
Request model cards or equivalent documentation describing intended use cases, performance characteristics across different population segments and input types, known failure modes, and specific use cases that the vendor does not recommend. The absence of this documentation is itself an important signal. Responsible AI vendors document their models. Vendors who cannot provide this information have likely not given governance much thought.
Data handling
Ask explicitly whether the vendor retains your inputs, including prompts and query data. Ask whether your data is used to train or fine-tune models, either the vendor’s foundation model or models shared with other customers. Ask about data residency and whether your data may be processed in jurisdictions with different regulatory requirements. Understand subprocessor arrangements, including which third-party infrastructure and model providers the vendor relies upon.
Security and adversarial robustness
Traditional security questions remain relevant, but they should be supplemented with AI-specific inquiries about adversarial testing, prompt-injection defenses for generative AI systems, and procedures for handling jailbreaking attempts. Ask whether the vendor has tested for data poisoning vulnerabilities and model extraction attacks.
Governance and incident response
Ask who in the vendor organization owns AI risk, what their governance structure looks like, and how AI-related incidents are defined, escalated, and communicated to customers. Understand the vendor’s model update policy, including how they notify customers of significant model changes and what testing is performed before updates are pushed to production.
To address the challenge of gaining visibility into when and how third parties are using AI, some organizations use tools that analyze DNS traffic and web data to flag potential GenAI use. However, many enterprises rely on manual outreach, which can create friction and delays in the onboarding process for new vendors. Standardizing your questionnaire process and integrating it into procurement workflows reduces this friction while maintaining rigor.
Contract Clauses That Create Real Accountability
The most sophisticated AI vendor questionnaire is worth little if your contract does not create enforceable obligations. AI vendor agreements have historically favored providers, often dramatically. Contracts studied favor providers by limiting liability and shifting compliance burdens onto customers, but this approach is arguably unsustainable as AI becomes deeply embedded in regulated industries like finance, healthcare, and legal services.
Negotiating position depends on your scale and leverage. Still, the following categories of protection should be pursued in every AI vendor agreement, with the effort calibrated to the application's risk level.
Data use and training restrictions
This is the area where default vendor terms most frequently conflict with enterprise interests. Most AI vendors default to using customer data to improve their models unless the contract says otherwise. Your contract should explicitly restrict the vendor from using your inputs, outputs, or interaction data to train, fine-tune, or otherwise improve any model, whether proprietary or shared with other customers, without your explicit consent. Include prohibitions on data commingling, meaning your data should not be pooled with other customers’ data in ways that could enable cross-contamination. For sensitive data, include no-retention provisions requiring the vendor to delete your data after processing rather than caching it for any purpose.
Transparency and disclosure obligations
Require the vendor to disclose material changes to the underlying model, including retraining events, significant architecture changes, and changes to training data. This matters because a model update can meaningfully change system behavior, potentially affecting your compliance posture. The vendor’s obligation should include advance notice before updates, where feasible, and retrospective notification where advance notice is not possible. Require maintenance of model documentation, and specify that updated documentation accompanies each material model change.
Compliance representations and warranties
Generic compliance-with-applicable-laws language is insufficient. For high-risk applications, require the vendor to represent that the AI system was developed in accordance with named frameworks, such as NIST AI RMF or ISO/IEC 42001, and that the vendor maintains a documented AI governance program. Require specific warranties regarding the bias-testing methodology and results, and tie them to indemnification provisions. Deployers may need to rely on developers to test and validate that the AI product or service does not create unlawful discrimination. In that event, the deployer should contractually require the developer to represent and warrant that the AI product or service does not create unlawful bias and link that representation to the defense and indemnity provisions.
Audit rights
Negotiate for the right to conduct or commission AI-specific audits, including fairness audits, security audits, and assessments of the vendor’s AI governance practices. Establish specific audit rights over the information that underpins compliance representations, even if access to the model itself is restricted for intellectual property reasons. Third-party attestations, such as AI-focused addenda to SOC 2 reports or independent bias audit reports, can partially substitute for direct audit access when vendors resist broader rights.
Incident notification
Define what constitutes an AI-related incident in your contract, including model performance degradation below defined thresholds, discovered bias issues, security events affecting the model, and significant unexpected behavioral changes. Specify notification timelines that allow you to respond before customer or regulatory harm materializes. Seventy-two hours is a common threshold for security incidents, but you may want shorter windows for high-risk AI applications where operational impact is immediate.
Performance standards and service levels
Traditional SLAs cover availability and response time. AI-specific performance standards should also address accuracy and output quality metrics. This is difficult to specify in absolute terms given the probabilistic nature of AI, but you can specify evaluation benchmarks, testing frequency, and remediation obligations when performance falls below agreed thresholds. For generative AI systems, consider specifying limits on hallucination rates and toxicity thresholds relevant to your use case.
Regulatory adaptation
The AI regulatory environment is changing rapidly across jurisdictions. AI systems are dynamic, as models drift, data shifts, and new legal standards emerge. Laws like Colorado’s AI Act and the EU AI Act explicitly contemplate ongoing monitoring and risk management for high-risk systems, not just one-time assessments. Include provisions allowing contract amendment to address new regulatory requirements, and specify the allocation of compliance costs between the parties as requirements evolve.
Portability and exit rights
Vendor lock-in is a significant risk in AI deployments. Negotiate for data portability rights that allow you to extract your data in usable formats. For fine-tuned or custom models developed during the relationship, clearly specify ownership of those model artifacts. Include transition assistance obligations that require the vendor to support migration if you terminate the relationship.
Shared Responsibility: Closing the Gaps, Not Just Documenting Them
One of the most dangerous assumptions in AI governance is that contractual assignment of responsibility equals actual management of risk. It does not. Your vendor can agree in writing to maintain bias testing practices while those practices are inadequate or inconsistently applied. You can have iron-clad contractual protections while having no operational visibility into whether they are being honored.
Effective third-party AI governance requires a shared responsibility model that is operational, not just documentary. This means understanding what the vendor is accountable for, what you are accountable for, and where the interfaces between those responsibilities require active coordination.
What remains your responsibility regardless of vendor obligations
Even when your vendor bears responsibility for model development, fairness testing, and technical security, you retain accountability for the deployment decisions you make. You are responsible for defining the use case and determining whether the AI system is appropriate for it. You are responsible for the human oversight mechanisms you build around the system. You are responsible for communicating to affected individuals when AI is influencing consequential decisions about them. You are responsible for the monitoring you conduct on your end of the deployment.
This point is critical in regulated industries. FINRA, for example, reminds firms of their regulatory obligations when using generative AI and large language models, noting that rules continue to apply when firms use Gen AI or similar technologies in the course of their businesses, just as they do when using any other technology or tool. As with any technology or tool, a firm should evaluate Gen AI tools before deploying them and ensure it can continue to comply with the existing rules applicable to the business use of those tools. Your vendor’s compliance with their own governance standards does not satisfy your regulatory obligations. Those obligations require your independent evaluation and ongoing monitoring.
Defining the interface
For high-risk AI deployments, consider building explicit responsibility matrices that assign clear ownership to every governance activity across the AI system lifecycle. This is not the same as a standard vendor risk assessment. It should specify who is responsible for input data quality before it reaches the vendor’s API, who monitors output quality after the vendor returns results, who conducts periodic revalidation as deployment conditions evolve, who is the point of contact for regulatory inquiries, and how the parties coordinate when an incident occurs.
Preferred vendor programs as a governance lever
One way to unlock value is to standardize the organization on preferred, vetted providers whose AI practices align with the organization’s Responsible AI standards. While total standardization may not be realistic, identifying and vetting the likely small group of vendors that represent the majority of the third-party usage within the organization can significantly streamline governance. When you concentrate AI spend with a smaller number of thoroughly vetted providers, you gain more negotiating leverage on contract terms, more operational familiarity with their systems, and more efficient ongoing governance. You also build relationships that make collaborative problem-solving feasible when issues arise.
The shadow AI problem
A particular challenge in shared responsibility is the proliferation of AI tools adopted outside formal procurement processes. Employees are using consumer AI tools, connecting work accounts to third-party AI services, and embedding AI capabilities into their workflows in ways that bypass vendor assessment entirely. Gaining visibility into when and how third parties are using AI is a growing challenge for enterprises. Your shared responsibility framework only works for vendor relationships you know about. Shadow AI creates ungoverned third-party risk that contractual protections cannot address. This requires both policy controls specifying which AI tools are authorized and technical controls that provide visibility into actual usage patterns.
Evaluating External Models and APIs You Do Not Fully Control
Beyond questionnaires and contracts, you need technical and operational approaches to evaluate the AI systems you are deploying. The fact that you cannot fully inspect a vendor’s model does not mean you cannot develop meaningful insight into how it behaves in your context.
Pre-deployment evaluation
Before integrating any external AI system into a consequential workflow, conduct a structured evaluation using test cases that represent your actual deployment conditions. Do not rely exclusively on vendor-provided benchmarks, which are typically constructed to show the system in its best light. Design your own evaluation dataset that reflects the distribution of inputs your system will actually encounter, including edge cases, adversarial inputs, and the demographic profiles of affected populations.
For generative AI systems, red-team the deployment before it goes live. This means systematically attempting to produce harmful, inaccurate, or policy-violating outputs through prompt injection, jailbreaking attempts, and edge-case exploration. What you discover in controlled testing is far less costly than what users or adversaries discover in production.
For predictive AI systems making consequential decisions, conduct structured fairness testing across protected characteristics. Measure disparate impact across demographic groups. If the vendor provides documentation of their bias testing, validate those claims by testing your implementation against comparable methodology. Vendors test under their conditions, but you need to test in your environment.
The model card gap
Model cards and system cards represent the AI industry’s current best practices for model transparency documentation. A well-constructed model card describes intended uses and limitations, provides performance metrics across different population segments, identifies known failure modes, and specifies ethical considerations. However, coverage is uneven. Request this documentation as a condition of procurement, and assess the quality of what you receive. Vague or incomplete model cards are a signal worth weighing in your vendor selection decision.
Continuous monitoring as third-party governance
Your monitoring obligations do not end at the vendor boundary. You should be tracking the performance of AI outputs you receive from external systems with the same rigor you would apply to models you built yourself. This means establishing baseline performance metrics at deployment, tracking them over time, and setting alert thresholds that trigger review when performance deviates from the baseline.
For high-risk applications, consider independent model validation, where a team separate from the deployment team periodically re-evaluates the external system against your governance standards. This is standard practice in financial services for credit models, and the logic applies equally to any high-risk AI deployment, whether the model is internally or externally sourced.
API-specific risks
When you consume AI capabilities through an API, you are exposed to a category of risk that has no parallel in traditional software procurement. The vendor can change the underlying model without changing the API interface. A generative AI API can return qualitatively different outputs after an unannounced model update. A predictive API can drift as the vendor updates training data. Your monitoring must account for these shifts.
Implement canary testing for API-delivered AI capabilities. Maintain a set of reference inputs with known expected outputs, and run these against the API on a regular schedule. Deviations from expected outputs can indicate model changes that require your assessment before they fully propagate into your production workflow. This is a lightweight but valuable form of ongoing model validation that catches silent updates before they cause downstream harm.
Managing the foundation model layer
Many AI vendors are themselves building on foundation models from a small number of providers. When you evaluate a vendor’s AI product, you are implicitly evaluating the governance of the underlying foundation model. Ask vendors what foundation models they build upon. Evaluate the governance practices and trust-layer characteristics of each foundation model independently. The Anthropic and Salesforce partnership in regulated industries, for example, reflects an explicit recognition that foundation model governance is a prerequisite for responsible deployment in high-stakes contexts. Your evaluation framework should extend to this level.
Operationalizing Third-Party AI Governance
Components, such as questionnaires, contract protections, shared responsibility frameworks, and evaluation practices, are individually necessary but collectively insufficient unless embedded in repeatable operational processes. Third-party AI governance needs a home in your organization.
Several practical steps are worth prioritizing.
First, update your existing TPRM program to include AI-specific triggers and assessment instruments. Not every vendor relationship requires full AI due diligence, but any vendor that uses AI in delivering services to your organization, whether explicitly as an AI product or embedded in their underlying operations, should trigger an AI-specific review. Develop criteria for tiering vendors by AI risk level and match assessment depth to risk level.
Second, integrate AI governance requirements into your standard contract templates and negotiation playbooks. Legal teams should not need to invent AI governance provisions from scratch for each new vendor engagement. Develop baseline language for each protective category described above, with escalation guidance specifying which terms are essential versus negotiable and under what conditions.
Third, establish a vendor AI inventory. You need to know which vendors are using AI in their service delivery, what AI systems they are using, and how those systems interact with your data and workflows. This sounds basic, but many organizations lack this visibility. Consider requiring vendors to disclose AI use as part of annual renewal processes, and conduct periodic scanning to identify AI capabilities embedded in vendor platforms that may not have been disclosed at initial procurement.
Fourth, build AI performance monitoring into your vendor governance cadences. Periodic business reviews with high-risk AI vendors should include a review of performance metrics, incident history, and updates to governance practices. This is not just a contract compliance review. It is an ongoing dialogue about whether the AI system continues to meet your standards as both the technology and the regulatory environment evolve.
Finally, recognize that third-party AI governance is a shared responsibility within your own organization. Procurement, legal, risk management, technology, and the business functions deploying AI capabilities all have roles to play. The governance vacuum that characterizes too many AI programs, where no one is clearly in charge, is as likely to occur in third-party AI governance as in internal AI governance. Assign clear ownership, create cross-functional working groups for high-risk vendor relationships, and ensure that AI risk is a standing agenda item in the forums that manage vendor relationships.
Accountability Cannot Be Outsourced
The growth of the AI vendor ecosystem is genuinely enabling. Organizations can access capabilities that would take years and hundreds of millions of dollars to develop internally. APIs and platforms have democratized access to state-of-the-art AI, driving enormous enterprise value.
But the governance implications of that access are not democratized. When you deploy an AI system in a consequential context, the accountability is yours. Regulators do not accept “my vendor did it” as a defense. Affected individuals do not distinguish between an AI decision made by your model and one made by your vendor’s model running in your system. Your customers and stakeholders hold you to account for outcomes, not for architectures.
By applying a holistic AI vendor assessment approach, organizations gain deeper visibility into AI vendor risk, reduce operational surprises, streamline contractual and governance alignment, support regulatory compliance, and enhance trust with stakeholders by demonstrating proactive vendor oversight.
The organizations that are getting third-party AI governance right are not treating it as an extension of traditional vendor management. They are treating it as a core dimension of their overall AI governance program. They have updated their TPRM frameworks, negotiated meaningful protections, built operational monitoring that extends to the vendor boundary, and created accountability structures that close gaps rather than document them.
That is what it means to own the risk that comes with the capability. There is no doubt that AI capability is worth owning, but the governance discipline to match it is not optional.


