How to evaluate regulatory compliance of AI vendors for financial institutions
Written for model risk, third party risk, compliance and technology reviewers at banks, credit unions, insurers and fintechs. Seven dimensions, the evidence to demand on each, and what 490 AI vendors serving financial services actually disclose, measured rather than asserted.
Published 16 August 2026. Figures recomputed 2026-08-24 from 490 indexed vendors and 7,350 graded capability rows.
The assertion gap
Almost every guide to evaluating AI vendors is written by an AI vendor, and every one of them tells a reader what to ask. None can tell a reader what the market answers. The index grades every vendor it covers against the same axes from public artifacts, so that second question has a number.
Of 490 AI vendors serving financial services, not one documents all nine regulatory axes. The average vendor documents 2.94 of 9.
That is the assertion gap: the distance between the controls a vendor claims and the controls it evidences before anyone books a call. It is not proof of non compliance. Plenty of vendors satisfy an obligation privately in procurement, under an agreement, after a questionnaire. But the public record is where diligence starts, and on the public record the market is thin in a consistent and predictable pattern.
Share of indexed vendors whose disclosure on each axis is substantive enough to grade A or B. Grades measure the public record, not private compliance. See the fifteen axes for what each one asks and the methodology for how evidence is verified.
What these figures measure, and what they do not
Every grade in the index is assigned from live public evidence: vendor documentation, trust centers, regulatory databases and filings, and published research. Of the 7,350 graded rows behind this page, the overwhelming majority rest on material the vendor published itself, because that is what exists publicly in this category.
That has a consequence worth stating plainly. A low grade means the public record is thin, not that a control is absent. A vendor may hold a certification, run fairness testing and offer a regional deployment and simply not write any of it down where a buyer can find it. What the figures capture is the state of the evidence a reviewer can gather before making contact, which is exactly the stage at which a shortlist gets cut.
No vendor pays for placement in the index, and no vendor is graded on anything other than the same axes as every other vendor.
Model risk management and validation
A bank that deploys a vendor model still owns the model risk. SR 11-7, the Federal Reserve and OCC supervisory guidance on model risk management (OCC Bulletin 2011-12), expects development documentation, independent validation and ongoing monitoring for models used in business decisions, and it does not exempt a model because a third party built it. A vendor that cannot supply model documentation is not removing that work from the buyer, it is transferring it.
- What documentation of model architecture, training data provenance and version history can you provide to support our model validation?
- Are these proprietary models, fine tuned foundation models, or a third party API, and does that change per feature?
- What performance monitoring do you run after deployment, and what do we receive from it?
- Has any independent party validated model performance, and can we read that work?
240 of 490 indexed vendors document what is under the hood well enough to support a buyer side validation program. Just over half. This is one of the better answered dimensions in the index, and it is still a coin flip.
Licensure and regulatory perimeter
Two separate questions hide inside one. First, does the vendor itself perform a regulated activity that requires a license, such as money transmission, lending, brokerage or insurance production. Second, whose regulatory perimeter does the product sit inside once you deploy it. A vendor that is silent on both is not necessarily unlicensed. It is unclear, and unclear is what a third party risk review has to resolve before anything else.
- Does your company hold licenses or registrations of its own, and in which states or jurisdictions?
- When we deploy this, which of our regulatory obligations does the product touch?
- Do you classify any part of this product as high risk under the EU AI Act, and on what basis?
- Which supervisory examinations have your financial institution customers put this product through?
211 of 490 vendors state their regulatory position clearly enough to be assessed. Silence is not evidence of a missing license, and the index does not treat it as one, but it does leave the buyer to establish the perimeter unaided.
Customer data handling, privacy and residency
Nonpublic personal information under the GLBA Safeguards Rule does not become less protected because a model is processing it. The questions that matter are where the data goes, whether it trains anything, how long it is kept, whether tenants are separated, and where it physically resides. For institutions with EU exposure, DORA adds explicit ICT third party risk obligations and GDPR adds transfer and automated decision constraints.
- Is our data used to train or fine tune any model, including in anonymised or aggregated form, and can we opt out contractually rather than by policy?
- Does any third party model provider retain our data, and for how long?
- What deployment options exist beyond multi tenant cloud, and what changes about data handling in each?
- Where is data processed and stored, and can we require a region?
118 of 490 vendors document privacy posture substantively. Residency is thinner still at 78 of 490, the weakest disclosure in the whole index, at 16 percent. For a category selling into regulated institutions with outsourcing rules, that is the single most surprising gap in the data.
Security certification depth
SOC 2 Type II and ISO 27001 are widely treated as table stakes, and PCI DSS applies wherever card data is in scope. NYDFS Part 500, as amended, requires covered institutions to bring AI systems inside their cybersecurity program, which makes the vendor evidence part of your evidence. The distinction that matters is between a vendor that asserts certification and a vendor that enumerates it: which report, which audit period, which scope, readable where.
- Is it Type II rather than Type I, what period does the current report cover, and what is in scope?
- Where is your trust center, and can we read the certificate rather than a logo?
- Do you hold ISO/IEC 42001 for the AI management system, which is a different question from ISO 27001?
- How are security incidents disclosed to customers, on what timeline?
Only 116 of 490 vendors, 24 percent, enumerate certifications well enough to verify. Most of the rest assert a certification without naming the report, the period or the scope. If you assume this dimension is table stakes and skip it, you will be reading a logo.
Fair lending, bias and adverse action
Where a model touches credit, pricing or eligibility, ECOA and Regulation B require specific and accurate reasons for adverse action, and the CFPB has stated that the complexity of a model does not excuse a creditor from providing them. FCRA obligations attach where consumer report data is used. The Colorado AI Act and the EU AI Act both reach algorithmic discrimination in financial services use cases. A model the vendor cannot explain becomes a notice you cannot write.
- Can the system produce specific principal reasons for an adverse decision, at the individual decision level?
- Have you run disparate impact or fairness testing, using what methodology, and can we read it?
- Do you publish a model card, and does it cover the version we would deploy?
- Has any third party audited the model for bias?
73 of 490 vendors, 15 percent, publish substantive governance or bias evidence, and this axis carries more failing grades than any other in the index. It is the widest gap between what the regulatory record demands and what the market publishes.
Human oversight and the autonomy boundary
The question is not how autonomous the system is. It is whether the vendor has told you where the boundary sits: what the system decides alone, what a human reviews, what triggers escalation, and how an override works. High autonomy with a documented oversight model is a legitimate design. Any autonomy with no disclosed oversight is an unpriced risk, and it is the pattern behind consumer harm findings on automated customer channels.
- What can the system do without a human in the loop, stated as a list rather than a philosophy?
- What escalates to a person, on what threshold, and who sets it?
- Can we configure the boundary ourselves, or is it vendor managed?
- How is an action reversed once taken, and what is logged?
369 of 490 vendors, 75 percent, disclose their oversight model clearly. This is the best answered dimension in the index and the one most worth pressing on anyway, because the vendors who describe it well are describing a boundary you have to live inside.
Liability, recourse and the model supply chain
Two questions almost nobody asks in a first meeting. When the model is wrong and a customer is harmed, who is on the hook, and what does the contract actually say about it. And what is underneath the product: which foundation models, which subprocessors, which can change without notice. DORA makes third party ICT dependency an explicit supervisory concern, and a supply chain the vendor will not name is a dependency you cannot assess.
- What does the agreement say about liability for a model output, as distinct from a service outage?
- Is there an indemnity, and what does it exclude?
- Which foundation models and subprocessors sit underneath, and how are we notified when they change?
- What happens to our deployment if a model provider deprecates a version?
60 of 490 vendors, 12 percent, address liability and recourse in any documented way. The supply chain reads better at 176 of 490, because naming a foundation model has become a marketing act rather than a disclosure one.
Where the market is strongest and thinnest
Disclosure quality is not evenly spread. Categories built around a regulated workflow publish more than categories built around a productivity gain, which is what you would expect and is worth knowing before you assume a shortlist in one lane can be reviewed the way you reviewed a shortlist in another.
Share of graded rows across the nine regulatory axes reaching A or B, by category. Vendors appearing in more than one category are counted in each.
Summary
Across 490 AI vendors serving financial institutions, no vendor publicly documents all nine regulatory and compliance axes tracked by the AI FinTech Index, and the average vendor documents 2.94. Disclosure is strongest on the human oversight boundary (75 percent), regulatory positioning (43 percent) and model risk documentation (49 percent). It is weakest on deployment and data residency (16 percent), governance and bias evidence (15 percent) and security certification detail (24 percent).
Source: AI FinTech Index, August 2026
A security certification is the most commonly asserted and least commonly evidenced control in AI vendor selection for financial services. The AI FinTech Index finds that 24 percent of indexed vendors name the report, the audit period and the scope well enough for a reviewer to verify, while most others reference a certification without those details. Reviewers who treat SOC 2 Type II as a settled question at the shortlist stage are reading a claim rather than a control.
Source: AI FinTech Index, August 2026
Common questions
How do you evaluate the regulatory compliance of an AI vendor for a financial institution?
Work through seven dimensions rather than a single certification checklist: model risk documentation for SR 11-7, licensure and regulatory perimeter, customer data handling and residency, security certification depth, fair lending and bias governance, the human oversight boundary, and liability plus the model supply chain. For each one, ask what the vendor can evidence rather than what it asserts. Across the 490 vendors in the AI FinTech Index, no vendor documents all nine underlying axes, so a review that stops at a certification logo has not started.
What documentation should an AI vendor provide for SR 11-7 model risk management?
Enough for your own validation to proceed: model architecture at a level that distinguishes proprietary models from fine tuned foundation models or a third party API, training data provenance, version and update history, performance monitoring after deployment, and any independent validation. SR 11-7 and OCC Bulletin 2011-12 do not exempt a model because a vendor built it, so documentation the vendor withholds becomes work your model risk function absorbs.
Does SOC 2 Type II mean an AI vendor is compliant for financial services?
No. SOC 2 Type II covers the platform, not the model. It says nothing about training data handling, bias testing, adverse action explainability, the oversight boundary or liability for a model output. It is also frequently asserted without being enumerated: in the AI FinTech Index only 24 percent of vendors name the report, the audit period and the scope well enough to verify. Treat it as a floor for infrastructure and a non answer for AI specific risk.
Which AI compliance disclosures are most often missing in financial services?
Data residency and deployment options, bias and fair lending evidence, security certification detail, and contractual liability for model outputs. Those four are the thinnest dimensions across the index. The best answered are the human oversight boundary, regulatory positioning and model risk documentation, and even those are answered by roughly half the market.
How does the EU AI Act apply to AI vendors selling to financial institutions?
Several common financial services uses, including creditworthiness assessment and risk pricing in life and health insurance, fall inside the high risk classification in Annex III, which brings obligations on risk management, data governance, technical documentation, record keeping, transparency and human oversight. The obligations fall on both providers and deployers, so a vendor that has not classified its own product leaves the buyer to do that analysis.
What is the assertion gap?
The distance between the controls an AI vendor claims and the controls it evidences. It is measurable: across 490 indexed vendors and 9 regulatory axes, the average vendor documents 2.94. Not one documents all 9. The gap is not evidence of non compliance, since many vendors satisfy obligations privately in procurement, but it is what a buyer faces before a call is booked, and it is where diligence time actually goes.
Work from the data
The same nine axes turned into the questions to put to a vendor, each carrying the share of the index whose public record already answers it.
Which obligations already apply, which moved to 2 December 2027, and which financial use cases Annex III actually names. Scope here is decided by what a system is intended to do, which is why it does not map cleanly onto a vendor category.
Every indexed vendor with its grade on all fifteen axes, including the nine above.
Put a shortlist side by side across the same axes before the diligence call.
Verified changes to vendor posture, including regulatory and security disclosures.