Xapien
Xapien automates the research step of due diligence and returns a document rather than a score. The product takes the name of a person or a company and produces a fully sourced background report, typically in minutes, drawing on corporate registries, sanctions and watchlist data, politically exposed person data, court records and the indexed open internet. The vendor puts the average at 7.5 minutes across a sample of a thousand reports, against a legacy baseline it constructs at 8.5 hours.
The design premise is stated openly and it is unusual: the vendor publishes a written argument that generative AI on its own cannot be used for due diligence, naming hallucination, inconsistency between similar prompts, and the absence of structured compliance data as the three reasons. Xapien positions itself as the answer to that, triangulating structured screening data against unstructured web material and citing every finding to a source.
The engineering centre of the product is an identity resolution engine that decides which of several similarly named subjects the report is actually about. The company markets false positive suppression as the primary benefit, describing a representative report as confirming 111 relevant sources and filtering out 297.
Financial services is one of six industry lines, alongside legal, enterprise, risk consultancies, education and non profit, and only the financial lines are graded in this record. Those lines are named and separately developed: know your customer and customer due diligence for banks, client onboarding for wealth managers and private banks, investor and director checks and target company diligence for private equity, and underwriting research and claims investigation for insurers. Named financial customers include Griffin, a United Kingdom bank, and Dow Jones Risk and Compliance, which sells a product called Integrity Check built on Xapien.
The workflow runs identify, run report, review and edit, share, then decide and approve. The vendor states that the review step can be manual or automated with configurable rules and that approval can be fully configured to the customer's own policy, which is the autonomy position and is graded as such.
Sold as annual subscription in three tiers set by report volume, with unlimited users. The legal entity is Digital Insight Technologies Ltd, founded in London in 2018 by two former BAE Systems engineers, and Xapien is its registered trademark.
Capability Axes
Capability grades
15 of 15 axes rated · 6 graded A or B
The removal test leaves nothing behind. Strip the models and what remains is a list of source links, which is the manual process the product exists to replace. Language models do the load bearing work at three separate points: parsing unstructured material in its original script, resolving which similarly named subject the report concerns, and writing the risk summary a compliance officer reads.
The founder and chief operating officer describes a deliberate multi model architecture on the cloud provider's own surface, routing high volume parsing tasks such as double barrelled family names and post nominal honours to a faster model and the reasoning that summarises risk to a stronger one. Report accuracy is attributed to that architecture combined with entity resolution.
The company was founded by two engineers from a defence contractor's cyber and financial crime division and has described itself as a deeptech company throughout, so this is not a screening database with a summariser bolted on.
The vendor publishes the option to remove the human and publishes no floor beneath it. Its own five step workflow states that the review and edit step can be manual or automated with configurable rules adaptable to the customer's risk policy, and that the decide and approve step can be fully configured to customer policy.
Read plainly, a customer may configure a path in which a machine written dossier on a named person produces an approval with no person reading it, and the vendor does not describe any decision it refuses to let the customer automate. That is a permissive position rather than a one, and the honesty of stating it is worth something.
Elsewhere the vendor frames the product as freeing analysts for higher value work and clearing low risk cases so teams concentrate on complex ones, which describes triage rather than replacement, and the fully sourced citable report is built for a human to interrogate. The two framings are not reconciled anywhere public. The question to put to this vendor is which approval decisions, if any, its configuration will not allow a customer to automate.
A headline accuracy figure with no method underneath it. The vendor publishes 97 percent report accuracy and a 30 percent reduction in generation time, and separately footnotes a 7.5 minute average to a sample of a thousand reports, which is more methodological disclosure than most of this category offers. What is not published is what accuracy means here, how ground truth was established, who measured it, over what sample, or how the figure decomposes.
For a research product the decomposition is the whole question, because the failure that ends a transaction is the risk the report never surfaced, and a single accuracy percentage does not separate a wrong statement from a missing one. No false negative rate, no recall measure against a known set of adverse subjects, no drift or retraining disclosure, and no per language or per jurisdiction breakdown was located across two passes.
The vendor's own surfaces are also inconsistent on coverage, stating support for more than 200 languages on the plans page, more than 130 in its published questions, and 40 or more on a partner's description of the same engine.
Named customers with attributed quotes, quantified outcomes, and independent analyst recognition, weakened by where the named references sit. On the record by name and role: the global head of due diligence at Dow Jones Risk and Compliance, the money laundering reporting officer at Griffin Financial Technology, the general counsel at Pinsent Masons, and staff at the University of Cambridge, ClientEarth and Dartmouth College.
Quantified claims include 97 percent report accuracy and 30 percent faster generation on the cloud provider's surface, a 7.5 minute average footnoted to a sample of a thousand reports, and a client expectation that low risk clients clear in hours rather than days. Third party standing is real: a Chartis FCC50 placement winning the entity management category, a KYC data and solutions category leader designation, and RiskTech100 recognition.
The limit is that the financial services references are anonymised by role, appearing as a head of compliance at an unnamed private equity firm, while the strongest named references sit in legal, education and philanthropy, outside the segments this record grades. The legacy comparison figures the vendor publishes alongside its own are a vendor constructed baseline rather than a measured control.
One genuine safety position and one empty headline. The genuine part: the vendor publishes a written argument that generative technology on its own is unfit for due diligence, naming hallucination and inconsistency between subtly different prompts as two of the four reasons, and builds the product around triangulating structured screening data against unstructured material with every finding cited to its source.
Citing every claim to a retrievable source is the strongest practical safety control available to a research product, because it makes a fabricated finding checkable by the person reading it. The empty part sits on the platform page, which carries a heading claiming industry leading safeguards with nothing published beneath it. Across two passes no red team result, evaluation method, incident disclosure, accuracy monitoring regime or acceptable use boundary was located.
The product's own failure mode also goes unaddressed in public: the vendor markets the suppression of roughly three quarters of candidate sources as a benefit, and what stops a genuinely adverse source being filtered out with the noise is not described anywhere.
A clear European posture with an American gap and a subject shaped hole. The vendor states it collects and processes only publicly available online sources, encrypts personal data in transit and at rest, aggregates searches so that the identity of the searching customer is not disclosed to the sources being queried, and commits that customer data is never used to train its models. That last commitment is specific and is the kind of statement most of this category leaves implicit.
What is absent is the American frame: the Gramm Leach Bliley Act is not referenced on any vendor surface, and no United States privacy programme, safeguards rule position or state privacy framework was located across two passes, despite the vendor reporting that a large share of revenue now comes from the United States. The larger gap is structural to the product rather than to the company.
Every report is a compiled dossier on a named individual who did not choose to be researched and will not see the output, and the public surface says nothing about how that person exercises access, correction or objection rights, or how long the dossier is retained.
An actual certificate, published as a document rather than asserted as a badge, which places this above the norm for the segment. The vendor links a downloadable ISO 27001 certificate naming the certified legal entity, states that it undergoes regular audits by independent external parties, and publishes baseline controls: encryption in transit and at rest, multi factor authentication at every tier, and single sign on from the middle tier upward. Three things keep it out of the top band.
There is no trust centre, so no penetration test summary, no subprocessor list and no policy set is obtainable without contact. There is no Statement of Applicability on the public surface, which means the scope of the certificate, the boundary that determines whether the research pipeline itself sits inside the management system, cannot be read from outside.
And there is no service organisation control report, which is the credential North American financial buyers ask for first and which the vendor will meet in procurement anyway. One drafting point a buyer should not read past: the same page names the standard as ISO 27000 twice alongside the correct reference, and ISO 27000 is the vocabulary document rather than a certifiable standard.
An unregulated software supplier to regulated buyers, which it neither claims otherwise nor addresses. The company is a United Kingdom private company, Digital Insight Technologies Ltd, trading under a registered trademark, and it sells a research tool that feeds obligations its customers hold rather than holding any itself.
It does not claim authorisation from a financial regulator, does not present itself as a credit reference agency or a regulated data provider, and does not overstate any of this, which is a point in its favour against the pattern of implied standing common in this category. What is missing is the positive statement a compliance buyer needs.
Across two passes no published position was located on its data protection registration, its controller or processor role for the subjects of its reports, its stance under the European Union AI Act given that it profiles named individuals for regulated decisions, or any regulatory examination or supervisory review it has been through with a customer. The analyst placements it publishes are commercial recognition and carry no supervisory weight.
The exposure is specific and the disclosure is absent. Two passes located no bias statement, fairness testing, model card, governance page or published evaluation. The exposure follows from what the product does rather than from a general concern about models.
Adverse media and politically exposed person research on named individuals depends on what has been written about a person and indexed, and that varies sharply by country, language, press freedom and name commonality, so a subject from a jurisdiction with a thin indexed press can return a clean report for reasons that have nothing to do with their conduct.
The vendor's own marketing sharpens this: the identity resolution engine decides which of several similarly named people the report concerns, and name disambiguation is precisely where transliterated and non Western naming conventions fail, while the source filter that discards roughly three quarters of candidate material is unaudited judgement applied to a real person. The vendor states the models handle non Latin scripts better than earlier approaches, which is a claim about capability improvement and not a measurement of differential performance.
A published remedy that is real but small, which is still more than most of this segment offers. The vendor states across several surfaces that a customer does not pay for any report they are not happy with, and pairs it with a named account manager in United Kingdom working hours and a feedback control on the report itself for enquiries outside them. That is a stated commercial remedy with a named trigger, and it is unusual in a category where the common position is silence.
What it is not is liability. The remedy returns the price of one report out of an annual commitment of at least 250, and the loss that matters from a due diligence product is never the report fee: it is the onboarded client who should have been refused, or the transaction abandoned over a finding that was wrong.
Across two passes no public terms of service, warranty, indemnity, service level with credits, liability cap or professional indemnity position was located, and the satisfaction remedy turns on customer dissatisfaction rather than on a demonstrated error, so it is discretionary in a way a contractual warranty would not be.
Disclosure at a resolution almost nothing else in this index reaches. The founder and chief operating officer names the provider, the specific model versions, the task each is assigned and the reason for the split: a lighter model for high volume parsing work such as separating double barrelled family names from post nominal honours, a stronger model for the reasoning that summarises risk and checks report accuracy, a newer model under evaluation, and a coding assistant used internally for development.
The infrastructure is named alongside it, down to the managed container service the platform runs on, and the executive gives cost and geographic data placement as stated reasons for the architecture. A buyer can therefore identify exactly whose model reads a dossier on their client and which one wrote the summary, which is the question this axis exists to ask. One qualification belongs on the record rather than in the grade.
All of this is published on the model provider's marketing surface rather than the vendor's own, and no model provider is named anywhere on the vendor's website, so a buyer who never leaves xapien.com learns none of it, and marketing pages are withdrawn without notice in a way that documentation is not.
Verifiable integration presence outside the vendor's own marketing, which is the test that matters here. A scoped Xapien application is listed on a major workflow platform's public store, providing a dedicated screening workspace for organisations and individuals, initiating requests, reviewing identity candidates, generating reports and storing risk scores inside that platform once configured with an interface credential.
A governance and risk platform publishes its own embedding of the research engine. An onboarding platform in the anti money laundering space runs a joint workflow. A market data and risk business ships a product built on the engine under its own brand, which is deeper than an integration and closer to an original equipment relationship. The vendor also publishes bulk upload for high volume processing and a planned cloud marketplace listing aimed at procurement in large banks.
Two things hold this below the top band. Interface access is gated to the highest tier, so the buyer who most wants to embed the product pays for the largest volume commitment to do it. The vendor's own published questions describe the product as accessed via the web with easy integration into any existing process, which understates the estate above and gives a buyer no connector catalogue to check against.
Two vendor sourced surfaces name two different clouds and only one of them can be current. The information security page states that the entirety of the data processing platform is hosted with one hyperscaler, inside a virtual private cloud dedicated to the vendor's exclusive use, describing that isolation as a wall around all stored information.
A later case study on a second hyperscaler's own site, quoting the founder and chief operating officer directly, describes the platform rebuilt on that second provider's managed container service, running its models, and cites load testing at five times the previous concurrent report volume. The same executive gives data residency as one reason for the choice, saying data can be kept in specific geographical regions.
Both are vendor sourced and the security page is the older of the two, so the likely reading is a migration the security page has not caught up with, but a buyer running a supplier assessment cannot resolve that from the public record and should ask directly. Beyond the contradiction the picture is ordinary: multi tenant hosted software only, no self hosted or private deployment option published, and no named list of available regions on either surface.
Structure without a number. Three tiers are named on a public plans page, set by annual report volume rather than seats, with a twelve month term and unlimited users and teams at every tier. The billing unit is defined more carefully than most: one report is one unit, a company report and a person report count the same, and the vendor states there are no extras. Discounted pricing is published as available to university and non profit buyers.
Against that, no figure appears in any currency on any vendor surface across two passes, and the aggregator listings carry the price field empty. Two of the vendor's own pages also disagree on the ladder itself. The plans page puts the middle tier at 650 reports a year and renders every feature under all three tiers; the platform page puts the same tier at 600 and shows a real progression in which single sign on arrives at the middle tier and interface access, configurable templates and the configurable risk framework are held for the top one. A buyer reading the plans page would not learn that interface access is gated.
Four financial segments are addressed separately rather than as one vertical page, each with its own stated job. Banking is know your customer and customer due diligence feeding financial crime compliance. Wealth management is onboarding of complex private clients where the vendor frames the risk as reputational rather than procedural.
Private equity is target company diligence early in a deal plus investor and director checks at scale, with a claim of clarity on 90 percent of low risk cases. Insurance is split into underwriting research on corporate structures and leadership backgrounds, and claims investigation on claimants. Financial customers are named: a United Kingdom bank and a market data and risk business that resells a product built on the platform. The ceiling is breadth of the wider company.
Financial services is one of six industry lines the vendor sells to, sitting beside legal, enterprise, risk consultancies, education and non profit, and the vendor's most detailed public case studies are concentrated outside finance.
Pricing
Vendor-published figures are labeled as such. Figures labeled “Estimated” are derived from third-party sources and have not been confirmed by the vendor.
| Entry Price | Pricing Basis | Data Protection Terms | Implementation | Source |
|---|---|---|---|---|
|
Not published. Entry commitment is stated as a volume rather than a figure, at 250 reports a year on a twelve month term
|
Annual subscription across three tiers set by report volume rather than by seat, with unlimited users and teams at every tier. The entry tier starts from 250 reports a year, the middle tier from 650 on the plans page and from 600 on the platform page, and the top tier from 3000. One report is one billable unit and the vendor states that a company report and a person report count identically, with no separate charge for either. The vendor describes the model as scaling with customer usage requirements and states there are no extras. It sells no single reports and no consulting engagements, so the annual commitment is the only route to the product. Coverage of corporate records, sanctions, watchlist and politically exposed person screening and multilingual research is present at every tier. The ladder above that is disputed between the vendor's own two pages: the platform page places single sign on and bulk upload at the middle tier and holds interface access, configurable report templates, the configurable risk framework and custom branding for the top tier, while the plans page renders every feature under all three. Discounted pricing is published as available to university and non profit buyers, and a separate not for profit plan is offered to qualifying organisations worldwide. | Not published at tier level. Data handling commitments are made uniformly on the information security page rather than sold as a tier: publicly available sources only, encryption in transit and at rest, search aggregation so the enquiring customer is not disclosed to the queried sources, and a commitment that customer data is never used to train the vendor's models. No data processing agreement, subprocessor list or retention schedule is obtainable without contacting the vendor. | No separate implementation, onboarding or setup fee appears on any published surface, and the vendor markets deployment as launching in days with no heavy integration work. Support is included rather than charged: every tier carries access to the help centre and the user community, the middle tier adds a named customer success manager, and the top tier adds premium support and solutions team access. The vendor publishes a satisfaction remedy under which a customer does not pay for a report they are not happy with, and offers a co branded or fully white labelled reporting interface through its own design team, with the white label option stated as available to top tier customers. Whether the branding work carries a charge is not stated. A thirty minute demonstration is the published entry route and no free tier or self serve trial exists. | Vendor Published |
Two passes across the vendor's plans page, platform page, published questions, regional landing pages and the major software aggregator listings produced no figure in any currency. The aggregator entries carry the price field empty and route to a quote request. What the vendor does disclose is the structure and the unit, and the unit definition is unusually precise for this category, which is why this sits above a bare contact sales posture without reaching a rate card.
Two inconsistencies belong on the buyer's question list: the middle tier volume floor differs between the vendor's own two pages, at 650 against 600, and the two pages disagree on which tier unlocks interface access, which matters because a buyer intending to embed the product may not learn from the plans page that it is gated to the top tier.