Gradient Labs
Gradient Labs builds Otto, an autonomous customer operations agent for regulated financial services, which does not deflect to a help centre but reasons through a query and takes the action, freezing a card, filing a dispute or running a customer due diligence check, with more than twenty compliance guardrails applied on every turn. It follows procedures the institution writes in plain English, deliberately exposes only the one or two tools a given procedure needs, and works across voice, text and email frontline support as well as back office investigations, sitting over existing helpdesks and connecting into the financial systems where the work completes.
Capability Axes
Capability grades
15 of 15 axes rated · 11 graded A or B
The product is an autonomous agent that reasons through a customer's problem and completes it, and the founders are explicit that they were not building an assistant. Otto interprets a query, follows a procedure written in natural language, calls the institution's systems and takes the action, whether that is freezing a card, opening a dispute or running a due diligence check. Apply the removal test and nothing remains, since a deflection bot pointing at a help centre is precisely the incumbent approach the company was founded to move past.
This is the most autonomy forward product in its lane and the company does not pretend otherwise, describing fully autonomous agents rather than assistants and aiming for minimal human intervention while the agent freezes cards, files disputes and runs due diligence checks.
What keeps the grade respectable is that the control architecture is real and specific: procedures are authored by the institution in plain English, guardrails execute on every turn and are counted, tool exposure is deliberately constrained per procedure, and the platform is described as fully auditable.
What is missing is the human boundary, with no stated escalation threshold, no confidence level that forces referral, and no described route by which a person reviews an action before it lands on a customer's account.
Two publishing habits sit above the category norm. Performance is stated as a curve rather than a headline, roughly 60 percent resolution on day one rising to 80 to 90 percent as a deployment matures, which is an honest admission that the system starts imperfect and a number a buyer can hold it to during a pilot. And guardrail execution volume with an associated quality score gives an operational measure of how often controls engaged.
Procedures written in plain English are readable by a compliance officer without vendor assistance. Absent are model documentation, error analysis, independent evaluation and any stated support for a customer's own validation.
Customers are named across several financial models, including a major money transfer business, digital banks, a savings and investing platform, an insurance provider and a lender, with more than twenty active customers and European bank deployments.
Performance is published as a maturity curve rather than a best case, at roughly 60 percent resolution on day one rising to 80 to 90 percent in mature deployments, with satisfaction above 80 percent across deployments and 98 percent at the top, and the company states its agents beat human teams on both satisfaction and quality scores.
One deployment is quantified in unusual operational detail: more than 500,000 customers served at a European digital bank with nine million guardrail executions and a 98 percent quality score. Revenue reached a million pounds within five months of launch, and a European bank's venture arm co led the funding.
The safety architecture is described concretely and, unusually, measured. More than twenty compliance guardrails run on every turn, and one deployment reports nine million guardrail executions alongside a 98 percent quality score, which is the first time this index has seen a vendor publish how often its guardrails actually fired.
A second design choice is explained rather than assumed: each procedure exposes only the one or two tools it needs rather than the full toolset, deliberately narrowing the agent's option space to raise accuracy. The company also delayed launch by fourteen months to build against internal quality benchmarks. What is absent is independent validation of any of it, and any statement on whether customer conversations inform model training.
The agent holds live customer conversations, reads account data and acts on accounts, which places it inside the most sensitive part of a bank's customer relationship, and the company discloses that it fails over across multiple cloud and model providers, meaning conversation content can traverse several third parties by design. A service organisation control type two certification is stated.
What is not published is the framework beneath it: no privacy policy detail on conversation handling, no retention schedule, no subprocessor list and no statement on how customer data is bounded when it reaches an external model provider.
A service organisation control type two certification is stated explicitly on the product site and framed as banking grade data handling, which names a standard and a level rather than gesturing at unspecified certifications, and that places this ahead of most of the index.
What is absent is the surrounding surface: no trust centre, no report request path, no stated audit period or scope, and no second framework such as an international information security standard, which European bank customers would ordinarily expect alongside it.
Gradient Labs supplies technology and holds no authorisation, the expected posture, and its domain credibility is unusually direct since the founding team led data science and machine learning across customer operations and financial crime at a licensed digital bank before starting the company.
The agent operates inside regulated processes, running customer due diligence and handling payment disputes, both of which carry defined obligations, and the platform is described as pre configured for consumer protection laws. The framing stays at the level of global regulations generally rather than naming the specific instruments it was built against, which is what the strongest vendors in this index now do.
Satisfaction and quality scores are aggregate measures of whether customers were pleased, not measures of whether they were treated evenly. The agent operates across voice, text and email in the United States and Europe, so speech recognition carries the accent, dialect and impairment variance seen throughout this index, and the consequences here are decisions rather than routing: a dispute filed or refused, a card frozen or left active. Nothing public breaks resolution or escalation rates down by language, accent or customer segment, and no fairness testing or accessibility analysis was located.
Outcome based pricing creates alignment rather than liability: the vendor earns on resolutions, so unresolved or reopened cases cost it revenue, which is a real incentive but not a commitment to bear the consequences of being wrong. Guardrails and auditability let an institution reconstruct what the agent did. Neither reaches the customer.
No accuracy guarantee, no remediation term, and no described route for someone whose card was frozen in error or whose dispute the agent declined, which matters more here than in advisory products because the agent takes the action rather than recommending it.
Gradient Labs admits something most vendors in this index avoid, that multiple external model providers sit in the path, disclosed through a failover design spanning several cloud and model suppliers. Acknowledging the dependency at all is more than most manage, and the durable workflow engine underneath is named openly through a published case study.
What is not disclosed is which providers, under what terms, or what they retain, and no subprocessor list exists, so a bank knows its customer conversations reach several third parties without being able to identify them.
The integration position is the product's central claim and it is correct: most agents stop at the help desk, and this one also connects into the financial systems where the work actually completes, which is what allows it to freeze a card rather than describe how to.
It runs as a layer over the major helpdesk platforms or an in house one, connects by interface, file upload or inside existing tools, and is built on a named durable workflow engine so a conversation can persist across hours or days and retry reliably. Failover spans multiple cloud and model providers, which is an availability design most vendors here do not describe at all.
Delivery is cloud hosted with deliberate failover across multiple cloud providers, serving customers in both the United States and Europe, which means conversation data moves between regimes and between providers as a matter of architecture rather than exception.
No hosting regions, residency options, transfer mechanisms or subprocessor locations were located, and the multi provider design makes that omission wider than usual since the set of systems touching a customer conversation is not fixed.
Rates are not published, but the pricing model is, and it is the unusual part: charging is outcome based on resolutions rather than seats or conversations, which means the vendor is paid when the work actually completes. That tells a buyer the shape of the commitment and aligns the commercial interest with the operational one. Independent reviewers note the trade offs plainly, that there is no self serve trial, the process is entirely sales led, and a resolution based bill can be hard to forecast as volumes move.
Coverage spans banks, fintechs, lenders, insurers, wealth management and crypto businesses across the United States and Europe, and the functional reach is wider than customer service alone, taking in back office investigations such as disputes and due diligence plus proactive outreach. Named customers cover digital banking, money transfer, savings and investing, insurance and lending.
The boundary is the function rather than the sector: this is customer and back office operations, with nothing addressing underwriting, capital markets, treasury or credit decisioning, and no credit union material.
What Changed
Material product, regulatory, evidence and commercial changes at Gradient Labs, each verified against a live source and tagged to the capability axis it bears on. Funding rounds and awards are not product changes and are not logged.
Gradient Labs launched Collaborate, which lets operators and engineers edit and improve AI agents using software development practice. The release includes version control, built in evaluations and continual learning.
Compared With
Most editorial comparisons pair two vendors the index assesses as direct competitors for the same buyer. Some pair vendors that are adjacent rather than rival, where the useful question is where one ends and the other begins. Each carries a verdict, the buyer conditions that favor each vendor, and a graded side by side.
Alternatives to Gradient Labs
The closest documented capability profiles to Gradient Labs in the same categories, ordered by similarity across the same fifteen axes the index grades every vendor on. Closest documented profile, not a claim that either product does the same job. No vendor pays for placement.
A lighter documented profile than Gradient Labs
Documents GLBA and Data Privacy Posture where Gradient Labs does not
Documents AI Liability and Recourse where Gradient Labs does not
Stronger documented coverage on Institution and Segment Coverage and Autonomy and Oversight Model
Documents GLBA and Data Privacy Posture and Deployment Model and Data Residency where Gradient Labs does not
Documents AI Liability and Recourse where Gradient Labs does not
Similarity is computed axis by axis from published grades, not from a composite score. The index does not aggregate grades into a total. See the fifteen axes and the methodology.
Pricing
Vendor-published figures are labeled as such. Figures labeled “Estimated” are derived from third-party sources and have not been confirmed by the vendor.
No pricing data has been verified for this vendor. Pricing information will be published here once confirmed through vendor disclosure or third-party estimation.