Capital Markets & Research AI
R

RavenPack

RavenPack is one of the oldest language analytics businesses in finance, founded in 2003 and headquartered in Malaga, Spain with a New York presence, and it now trades under two live names. The company name carries the original business, which reads unstructured financial text at scale and turns it into structured signal: entity resolved sentiment and event detection across more than 40,000 news sources, filings and earnings transcripts, negative news monitoring across more than 7 million companies, executive sentiment drawn from earnings calls, and packaged quantitative factors sold to hedge funds, banks and asset managers. The platform name is Bigdata.com, launched on 22 October 2024 and presented as Bigdata.com by RavenPack, which is the current delivery surface and the one a new buyer meets first. This index has no aliases field, so both names are recorded here deliberately: a reader searching for Bigdata.com should land on this record.

The two are graded as one record rather than two because, unlike a conglomerate with unrelated divisions, everything this company sells is the same capability. There is no second business to dilute a grade. What did change is the architecture. Bigdata.com is a retrieval augmented generation platform combining vector search, run on Vespa Cloud at billion document scale, with the entity knowledge graph the company spent two decades building, exposed through an API, a research assistant, desktop and mobile applications and autonomous research agents that run tasks and produce daily pre market notes. That is a re platforming of delivery on top of a continuous data and entity resolution asset rather than a legacy vendor bolting on artificial intelligence, and the distinction matters when reading the centrality grade.

The content estate is the other half of the business and is unusually well disclosed. More than 170 content providers are licensed, with the Financial Times, the Economist Intelligence Unit and Preqin named individually, alongside a podcast corpus of roughly 20,000 shows and 5 million episodes, fund holdings, jobs data, environmental and governance scores, corporate fundamentals and regulatory filings. In July 2026 the company launched what it calls the tokenisation of content, a marketplace in which agents retrieve and pay for licensed premium content per token rather than per document, priced publicly.

Funding includes a 20 million dollar round led by GP Bullhound with participation from the European Investment Bank Group in 2024, and a subsequent investment from FT Ventures alongside the content agreement. Vendor material states more than 100 global financial institutions use the platform, without naming them.

Last VerifiedAugust 21, 2026
Compare RavenPack with other vendors
Founded
2003
Headquarters
Malaga, Spain
Categories
capital-markets-ai, alternative-data
Assessment

Capability Axes

Capability grades

15 of 15 axes rated · 8 graded A or B

AI Capability
AI Centrality
AA on AI CentralityThe artificial intelligence is the product. Remove the models and there is nothing left to sell.
Vendor Published

The removal test leaves a warehouse of text nobody can read. Every output this company sells is produced by language processing over unstructured sources: sentiment and event detection resolved to named entities across more than 40,000 news outlets, filings and transcripts, negative news screening across more than 7 million companies, executive sentiment extracted from earnings calls, and on the current platform a retrieval augmented generation layer that answers questions against billions of documents with inline citations. Nothing here is a rules engine with a model attached.

The more interesting question is whether this is one product story or two, because the company predates the current era by two decades and the answer is not the usual one. It is not a legacy platform that bolted artificial intelligence on. Nor is it simply continuous. The evidence points to a re platforming: vector search at billion document scale on a third party search cloud, a generative research assistant and autonomous agents are architecturally different from what the company shipped for most of its life. What carried across is the asset rather than the code, specifically the entity resolution and knowledge graph built over twenty years, which is what makes retrieval land on the right company. That is a genuine moat and it is also the part a buyer cannot inspect.

Autonomy and Oversight Model
BB on Autonomy and Oversight ModelA written commitment that the models work alongside human judgment, with real review surfaces, short of the full control structure: commonly the threshold at which the system stops or what happens after it is wrong.
Vendor Published

Provenance is genuinely engineered here rather than asserted, which is the strongest oversight property on the record. The platform is described as grounded by design with a search first architecture, every answer linked to its source and inline citations attached to every claim, and precision retrieval that injects only the ranked passages carrying the answer rather than whole documents. The company states this makes output less prone to hallucination, which is the honest form of the claim and deserves explicit credit: it describes a control that reduces the risk rather than asserting an absolute that no vendor can support.

What holds it at this band is confidence and the agent path. Provenance is surfaced at every decision but confidence is not, and nothing published describes what the system does when retrieval finds no supporting passage, whether it abstains, hedges or answers anyway. Separately, the platform now ships autonomous research agents described as solving complex tasks independently and producing daily pre market notes, and no gate, review step or scope constraint on those agents was located. The provenance control sits on the retrieval path; the autonomy sits on the agent path, and the second is where an unreviewed output reaches a decision.

Ask what the agents do when no supporting content is retrieved, whether a confidence or coverage measure is exposed per claim, and what constrains an autonomous agent run before its output reaches an analyst.

Model Risk Management and Transparency
CC on Model Risk Management and TransparencyTransparency is claimed in general terms with no mechanism a model validator could interrogate.
Third Party Estimated

Architecture is described more openly than most of this lane manages: retrieval augmented generation combining vector search with an entity knowledge graph, the search layer named as a third party vendor running at billion document scale, and precision retrieval that returns ranked passages rather than documents. Two quantified claims are published and they are different in quality. Reducing context consumed per query by up to 100 times is bounded, measurable in principle and about an engineering property, though no methodology, baseline or workload is stated. A tenfold increase in research efficiency is unit free and measures nothing a buyer can check.

What is missing is the model risk substance a regulated institution needs. No accuracy measure, no evaluation methodology, no retrieval precision or recall figure, no hallucination rate despite the claim of reduced propensity, no identification of the language models used, no versioning and no independent validation was located.

One check was run specifically and belongs on the record so a future reader does not credit it. Vendors in this segment circulate placements on financial language model leaderboards, and the most cited evaluator in this market, the S&P benchmark suite operated by Kensho, was sunset in May 2026 with its public leaderboards frozen and no longer updated. No placement claim by this vendor was located, and if one surfaces from that evaluator it cannot move this grade, because a frozen leaderboard evidences a moment rather than a current system.

Ask for retrieval precision and recall on your own document classes, for the models used and their versions, and for any independent evaluation dated after May 2026.

Operational and Outcome Evidence
CC on Operational and Outcome EvidenceUnnamed case studies, customer logos, or claims without numbers. Prestige is not measurement: the calibre of the client list describes the buyer rather than the product, and coverage statistics are not adoption statistics.
Third Party Estimated

Institutional validation is real and it is the wrong kind for this axis, which is a distinction worth making carefully because the logo surface here is large. The named organisations are overwhelmingly suppliers and partners rather than disclosed customers: the Financial Times, the Economist Intelligence Unit and Preqin license content in, a search vendor and a cloud data platform supply infrastructure, and a global quantitative manager ran a six week research competition on the data with its own consultant community. Those relationships each required diligence by a serious counterparty and they say nothing about what a customer got.

What is absent is any attributed outcome. No named client is quoted, no performance of a signal is published, no backtest, no accuracy measure and no case study with a result. The headline efficiency claim, a tenfold increase in research efficiency, is unit free: tenfold against which baseline, measured how, on what task. Vendor material states more than 100 global financial institutions use the platform and names none of them. Twenty two years of continuous operation selling to hedge funds and banks is itself meaningful evidence of fitness, and it is the kind a buyer infers rather than reads.

Ask for two or three referenceable clients in your own segment, and for any published or shareable measurement of signal performance or research time saved with the baseline and the task stated.

AI Safety and Data Stewardship
CC on AI Safety and Data StewardshipGeneral assurances that do not answer the question this axis asks, which is whether one customer’s data trains models serving its competitors. Unbounded cross client learning stated with no boundary grades here too.
Vendor Published

The exposure on this platform is what the query reveals, and the commercial model sharpens it. A portfolio manager researching a name through the research assistant discloses live interest in that name, and the platform serves competing funds in the same market. Per token metering makes that record finer grained than a flat subscription would, because every retrieval is individually priced and therefore individually logged, which is excellent for cost transparency and creates a detailed picture of what each client is investigating and when.

Nothing published resolves it. No statement was located on whether query text or retrieval history is retained, for how long, whether it informs ranking or product development, or how one client's research is separated from another's. Marketing language on the developer surface describes building agents on systems the customer fully owns and controls, and that sentence sells rather than binds: it describes where the customer agent runs, not what happens to the queries that agent sends into this platform. The two are different questions and only the second one matters here.

Ask in writing whether query text, entity selections and retrieval logs are retained or used for any purpose beyond serving and billing the request, and require the answer as a contract term rather than a description of the architecture.

Regulatory and Compliance
GLBA and Data Privacy Posture
BB on GLBA and Data Privacy PostureA substantive privacy document that reaches the product itself, short of the subprocessor list or the full data handling detail.
Vendor Published

The governing framework is not the one this axis names, and saying so is more accurate than penalising a fit that does not apply. The subjects here are companies, securities and markets rather than consumers, the inputs are published news, filings, transcripts and licensed datasets, and no consumer financial account data is processed, so the United States financial privacy statute has little purchase. What governs instead is European data protection law, given a Spanish domicile and European operations, together with the terms of more than 170 content licences, which are the instruments that actually constrain what may be retained and redistributed.

Two qualifications keep it from higher. Executive sentiment analysis is explicitly about named individuals, and podcast and news content carries the speech of identifiable people, so personal data is present even though consumers are not the subject. And no data processing agreement, retention schedule or subprocessor list was located in two passes.

Ask for the data processing terms and subprocessor list, and establish what the content licences permit you to retain, redistribute or feed into your own models downstream of a query.

Security Certifications and Trust Center
BB on Security Certifications and Trust CenterA recognised certification named in the vendor’s own material without the artefact, or with a scope or renewal question the buyer has to raise.
Vendor Published

A dedicated trust centre exists on the vendor's own subdomain, running on a recognised compliance automation platform, and two frameworks are named in the vendor's own material, with an enterprise security line citing an examination against the service organisation controls criteria and the international information security management standard alongside single sign on, data residency and audit controls. That is materially more than the records held at the band below, which have no trust centre and no named framework at all.

It is held here rather than higher because the credential does not survive close reading. The frameworks appear as feature bullets on a pricing page with no verb attached, so the claim does not state whether a report exists, what type it is, what period it covers, what systems fall inside the scope boundary, or which body issued the certificate. The service organisation controls examination produces an attestation report over a defined period rather than a certification, and a claim that names neither the type nor the period leaves a buyer unable to tell a point in time design review from a year of tested operation. The trust centre contents themselves were not retrievable in this pass.

The banked check belongs here: a company of this age selling to hedge funds and banks, licensing content from the Financial Times and Preqin, has certainly completed many security reviews and the documentation almost certainly exists behind that portal. The grade records what a buyer can confirm before signing an agreement, not what probably exists.

Ask for the current attestation report with its type, period and scope boundary, the certificate and issuing body for the management standard, and confirmation that the scope covers the platform you are licensing rather than corporate systems only.

Regulatory Status and Licensure
CC on Regulatory Status and LicensureThe regulatory position is unstated. Most vendors in this index are technology suppliers and being unlicensed is the correct posture, so this grade records silence about the posture, not a missing licence.
Vendor Published

No supervisor, licence, registration or statute is named anywhere in vendor material, and for this business model none is obviously required, since selling analytics and licensed content to institutions is not itself a regulated activity. Naming what governs instead is the useful move.

Three frameworks bear on this product and none is addressed. European artificial intelligence legislation applies to a Spanish domiciled provider and its obligations turn on how a system is classified, which is a question a buyer will be asked by its own compliance function. Market abuse rules matter because news derived signals delivered at speed to trading desks raise questions about the handling of information and the timing of dissemination, and the vendor sits directly in that path. And where sustainability scores form part of the content estate, European supervision of sustainability ratings providers has recently tightened, which is a live question for anyone consuming those fields.

The content licences are the fourth instrument and the one most likely to bite operationally, since redistribution terms determine what a customer may do with retrieved material inside its own systems.

Ask how the vendor classifies its systems under applicable artificial intelligence regulation, what its position is on the redistribution of retrieved licensed content, and whether any content class carries its own regulatory conditions.

AI Governance and Bias Disclosure
CC on AI Governance and Bias DisclosureResponsible artificial intelligence committed to in policy language with no evaluation behind it, on a product whose bias surface is modest.
Vendor Published

No model card, governance structure, fairness testing or coverage analysis was located in two passes.

The adapted exposure for a media derived signal is structural and specific, and it runs against the product's own logic. Any measure built on volume and tone of coverage systematically favours companies that are heavily covered, which means large capitalisation, English language and developed market names generate more detectable events than identically behaving companies that are thinly covered. A negative news screen across more than 7 million companies necessarily has very uneven density across that population, and a clean screen on an under covered company means the coverage was absent rather than that the conduct was. Nothing published quantifies that density, by market, language, company size or listing status.

The multilingual dimension compounds it, since extraction and sentiment quality varies by language and no per language performance is disclosed. For a buyer using this to screen counterparties or suppliers, the practical risk is a false sense of clearance rather than a false alarm.

Ask for source density and detection rate broken down by market, language and company size, and establish what a clean screen actually means for a company in a thinly covered jurisdiction.

AI Liability and Recourse
CC on AI Liability and RecourseMechanisms that enable challenge, such as audit trails and source traceability, with nothing standing behind the output and no route for the person affected.
Vendor Published

Terms and conditions and a privacy policy are published, and no warranty, indemnity, accuracy commitment, service level or remediation position attaching to an incorrect output was located in two passes.

The consequence structure is unusual for this index and cuts in the customer's favour on one side and against it on the other. Because the subjects are public companies and securities rather than consumers, the third party who might be harmed by a wrong answer has far less exposure than the individual mis scored by a lending or fraud model. The customer, however, is more exposed than most, because output here feeds investment memoranda, pre market notes and systematic signals where a wrong or mis attributed fact carries money directly and may not surface until after a position is taken.

The mitigation is architectural and real. Inline citation on every claim means a wrong answer is checkable against its source in a way an unsourced summary never is, so the customer holds the means to verify even without a contractual commitment. That is a better position than most records here, and it is a control rather than a remedy.

Ask what the vendor warrants about the accuracy of retrieved content and generated summaries, where liability sits when a cited source is correct but the generated claim misstates it, and what recourse exists if a licensed content set is withdrawn mid term.

Integration and Deployment
Model Supply Chain Disclosure
BB on Model Supply Chain DisclosureSubstantial partial disclosure, or a chain that is structurally short: an explicit in house build, on premise deployment, per customer instances, or zero retention at the model layer.
Vendor Published

The content half of the supply chain is the best disclosed in this lane and possibly in the index. More than 170 content providers are licensed and several are named individually, including the Financial Times, the Economist Intelligence Unit and Preqin, a public data catalogue enumerates what is available, the classes are broken out to the level of individual per token prices, and the podcast corpus is quantified at roughly 20,000 shows and 5 million episodes. A buyer can therefore see what evidence sits behind an answer, who owns it and what it costs, which is exactly what this axis exists to surface. The search infrastructure vendor is named too.

The model half is absent. No language model, model family, provider or hosting arrangement is identified for the research assistant or the autonomous agents, no versioning is stated, and no subprocessor list was located. For a platform whose entire pitch is grounding and citation, the identity of the model doing the generating is a conspicuous omission, and it matters commercially as well as technically, since a customer in a regulated institution will be asked which third party model saw its queries.

Ask which models power generation and where they are hosted, whether query content reaches a third party model provider, and for a written subprocessor list covering both the content and the inference path.

Core Systems and Integration Depth
BB on Core Systems and Integration DepthNamed systems or a documented public API, with the depth or the production evidence left open.
Vendor Published

The developer surface is genuinely deep and is documented publicly rather than described. There is a documented application interface with its own documentation domain, self serve key issuance, a stated intent that customers build and deploy their own agents on the data, desktop and mobile applications, and a research assistant. Distribution reaches named platforms, with the data available through a major cloud data platform and used inside a global quantitative manager's simulation environment for a research competition, which is a real integration into a working quant workflow rather than a logo.

The gap is the systems an institutional buyer actually runs on. No order management, execution management, portfolio management or risk system integration is named, and no market data terminal or compliance case management connection was located. For a research assistant that is tolerable, because the output lands in a document. For the signal business, whose customers consume factors inside quantitative pipelines, the absence of any named portfolio or risk system is a real omission.

Ask which portfolio, risk or order systems the vendor has existing connectors for, and what the delivery mechanism is for systematic signal consumption rather than interactive research.

Deployment Model and Data Residency
CC on Deployment Model and Data ResidencyCloud only with nothing stated, which is the category norm.
Vendor Published

Two infrastructure providers are identifiable, with the vector search layer named as a managed third party cloud and the data separately distributed through a major cloud data platform, which is more supply chain visibility than most records here offer. Data residency appears as an enterprise tier feature alongside single sign on and audit controls, and the enterprise page states that deployment will be tailored to the firm.

That is where it stops, and a two word feature bullet is not a residency posture. No region list, no statement of which jurisdictions data can be pinned to, no commitment on where query text and retrieval logs are processed and stored, no retention schedule, and no confirmation of whether a private or customer hosted deployment genuinely exists or whether tailoring means commercial terms rather than architecture. Naming the search vendor tells a buyer where documents are indexed, not where its own queries land, and for a platform whose queries reveal investment intent those are different questions with different consequences.

Ask which regions can be selected for processing and storage, whether residency is a contractual commitment or a best effort configuration, and whether any deployment option keeps query text inside your own environment.

Commercial
Commercial Transparency
AA on Commercial TransparencyPublished per unit rates a buyer can price against before any conversation.
Vendor Published

The most complete commercial disclosure encountered in this corpus, and it is worth being specific because the grade is unusual. Pricing is usage based and published: charge is per token consumed, described as starting from fractions of a cent, with no seats and no subscription required to begin. A public rate card breaks cost down by content class, with premium news, fund holdings, jobs data, environmental and governance scores, corporate fundamentals, earnings transcripts and regulatory filings each carrying their own figure, and regulatory filings priced roughly forty times below premium news, which tells a buyer exactly where the expensive content sits.

Seven named worked examples are published with their costs, ranging from about 0.57 to 1.82 dollars per run, each labelled with the artefact it produces and its approximate output size, alongside an interactive calculator. The alternative model is stated too: a heavily used content set can be moved to a flat subscription, dropping its per token content charge and leaving only search and retrieval tokens. Volume and committed use discounts, custom dataset licensing and an enterprise tier are named as existing without figures, which is normal and does not undercut the rest. Entry is 25 dollars of free credits with no card.

The honest limit is that the enterprise path, which is where an institutional buyer of the legacy analytics products will actually land, is quote based, and the legacy factor and analytics products carry no published price at all. That is a real gap sitting behind an otherwise exemplary page.

Institution and Segment Coverage
BB on Institution and Segment CoverageNamed segments with dedicated material behind part of the coverage.
Vendor Published

Subject coverage is the strength and is quantified throughout. More than 7 million companies are monitored for adverse news, more than 40,000 news sources are ingested, and the licensed estate spans more than 170 content providers across at least seven distinct classes including news, regulatory filings, earnings transcripts, fund holdings, jobs data, environmental and governance scores and corporate fundamentals, plus a podcast corpus of roughly 20,000 shows and 5 million episodes. Operations reach five continents.

The institutional side is narrower and that is what holds the grade. The documented buyers sit almost entirely in capital markets: hedge funds, banks and asset managers, and now developers building agents on the API. No insurer, payments business, lender, retail bank or credit union use case was located, and this index grades coverage across financial institutions rather than across data. A vendor can be extremely broad in what it reads and still serve one corner of the market.

Ask which of your peer institution types the vendor has live deployments in, and whether the content sets your workflow needs are included in standard licensing or negotiated separately.

Tracked Since Listing

What Changed

Material product, regulatory, evidence and commercial changes at RavenPack, each verified against a live source and tagged to the capability axis it bears on. Funding rounds and awards are not product changes and are not logged.

Jul 20, 2026Pricing / packaging

RavenPack launched what it calls the tokenisation of content on Bigdata.com, a marketplace in which AI agents retrieve, license and pay for premium content per token rather than per document, drawing on more than 170 licensed providers spanning market data, newswires, research, transcription, expert networks and global media. The vendor states that precision retrieval sends a model only the excerpts that carry the answer, reducing context consumed per query by up to 100 times, and that existing provider agreements are honored so a firm bringing its own content licenses pays only for retrieval. Access runs through one connection via MCP or API with entity resolution from the company's knowledge graph built in.

Bears on: Commercial TransparencySource
Our read on this change →Tracked since Jul 2026
Head to Head

Compared With

Most editorial comparisons pair two vendors the index assesses as direct competitors for the same buyer. Some pair vendors that are adjacent rather than rival, where the useful question is where one ends and the other begins. Each carries a verdict, the buyer conditions that favor each vendor, and a graded side by side.

Alternatives to RavenPack

The closest documented capability profiles to RavenPack in the same categories, ordered by similarity across the same fifteen axes the index grades every vendor on. Closest documented profile, not a claim that either product does the same job. No vendor pays for placement.

Documents Model Risk Management and Transparency and AI Liability and Recourse where RavenPack does not

Documents Operational and Outcome Evidence and Model Risk Management and Transparency where RavenPack does not

Documents Model Risk Management and Transparency where RavenPack does not

A lighter documented profile than RavenPack

A lighter documented profile than RavenPack

Documents Operational and Outcome Evidence where RavenPack does not

Similarity is computed axis by axis from published grades, not from a composite score. The index does not aggregate grades into a total. See the fifteen axes and the methodology.

Commercial

Pricing

Vendor-published figures are labeled as such. Figures labeled “Estimated” are derived from third-party sources and have not been confirmed by the vendor.

Entry Price Pricing Basis Data Protection Terms Implementation Source
Free tier with 25 dollars of credits, no card required; thereafter usage based from fractions of a cent per token, with published worked examples at roughly 0.57 to 1.82 dollars per research run
$0 baseline
Usage based per token of content consumed, with an optional flat subscription per content set and a quote based enterprise tier. Not published. Data residency, single sign on and audit controls are named as enterprise tier features without terms. None stated for self serve. Enterprise onboarding is described as including a named team and an agreed service level, without figures. Vendor Published

Exceptional disclosure by the standards of this corpus, on the platform side. Charge is per token of content consumed, published as starting from fractions of a cent, with no seats and no subscription required to begin. A public rate card breaks cost by content class, and the spread is itself informative: regulatory filings sit roughly forty times below premium news, with earnings transcripts, corporate fundamentals, environmental and governance scores, jobs data and fund holdings in between. Seven named worked examples are published with per run costs from about 0.57 to 1.82 dollars, each labelled with the artefact produced and its approximate output size, and an interactive calculator lets a buyer model monthly cost against expected run volume before speaking to anyone.

An alternative model is published rather than hidden: a heavily consumed content set can be switched to a flat subscription, which removes its per token content charge and leaves only the lower search and retrieval tokens. Committed use and volume discounts, custom dataset licensing and tailored deployment are named as available without figures.

Two gaps are worth a buyer's attention. The enterprise path, which is where an institutional buyer of the legacy analytics and factor products will land, is quote only, and those legacy products carry no published price anywhere. And the token unit means cost scales with how much content an agent retrieves rather than with team size, so a poorly scoped agent that retrieves broadly is the main route to an unexpected bill. The published calculator mitigates that considerably.

Contact us

Found a vendor we missed? Have feedback on the index? We’d love to hear from you.

AI FinTech Index

The AI FinTech Index is an independent index that tracks changes to AI vendors in financial services. It holds 489 vendors across banking, lending, insurance, wealth, capital markets and financial crime compliance, each graded on the same 15 capability axes from public sources. No vendor pays for inclusion, placement, or rating.

Index Status
Last index update
September 5, 2026
The AI FinTech Index is an editorial reference, not a regulatory body. Vendor data is verified against published sources and public regulatory filings. Figures labeled “Estimated” have not been confirmed by the vendor. See the Methodology page for evaluation standards and limitations.
© 2026 AI FinTech Index
3801 N Capital of Texas Hwy, Ste E240 · Austin, TX 78746