The definitive private company dataset and virtual analyst

Private company data is fragmented across providers, with no single source of truth. CapData unifies it into one reconciled, AI-ready foundation.

Analytics interfacesIntelligent agentsProprietary systems

How CapData works

From raw sources to one trusted dataset

RAW DATA SOURCES DELIVERED WHEREVER YOU WORK SELECT NORMALIZE RECONCILE ENRICH TRUSTED DATASET Orbis FactSet News Web Platform Web Application AI Assistant Ask Anything Excel Exports </> API REST / GraphQL Your Systems CRM, DWH, Models

Scroll the diagram sideways →

Why CapData

What sets CapData apart

01

Unrivalled coverage and reliability

Starting from best-in-class private companies, deals and financials dataset.

02

AI-based sector intelligence

Proprietary taxonomy, AI-powered sector allocation and intelligent search surface the companies that truly fit.

03

One reconciled dataset

Companies (financials & valuations), deals, advisors, GPs, sectors, news and people. All linked.

04

Cleansing done upfront

No manual cleanup. Your team starts with structured, trusted data.

05

Every value auditable

Trace every figure back to its calculation path and original source.

06

Targeted enrichments

Where data quality is low, we systematically identify and enrich records with better data points.

Coverage

Integration of Moody's best-in-class private companies dataset

The broadest source of truth with the most reliable information on the market.

600M+Private companies tracked
700M+Financial records
200M+Ownership records
400M+Company directors

Coverage figures are directional and rounded.

How It Gets Clean

We reconcile it before you ever run a query

01

Noise removal

~95% of raw entities are subsidiaries or assets, not operating companies. Removed before they clutter search.

02

Normalization

Input errors corrected. Categories aligned across every source.

03

Reconciliation

Duplicates removed within a source, and across sources that name a company differently.

04

Ambiguous datapoints dismissal

Our models score the data, and remove dubious outliers.

Built For AI

An LLM is only as good as the data underneath it

Old Way

Raw, unreconciled data

  • Confusing for LLMs to reason over
  • Hallucinated or ungrounded answers
  • Same question, different answer twice

CapData

Trusted, reconciled dataset

  • Query a reconciled, enriched fact base directly
  • Grounded answers, consistent every time
  • Native to your AI via MCP and Skills

One Dataset, Every Channel

Consume it however fits your stack

Analytics platform

Sector dashboards, league tables and deal screens, built into the CapData platform.

AI integrations

Use CapData's built-in AI assistant, or connect your own AI agents using our MCP- and Skills-native stack.

Your systems

Bulk files via S3, a GraphQL Lookup API, or a deployment matched to your stack.

The full CapData platform view: company screening filters, a sector by geography breakdown of company counts, selection metrics and an exportable results list.

About CapData

Built to power private market intelligence

Private capital markets run on fragmented, manual data. CapData was founded by a former partner at a large-cap private equity firm to fix that.

We combine modern data engineering and AI to turn fragmented information into one trusted foundation for investment teams, analytics platforms and intelligent agents — so leading institutions uncover better opportunities and build conviction faster.

Headquarters
London, UK
Coverage
Companies • Deals • Sponsors • Funds • Advisors • Sectors • People • News
Mission
Better Decisions Through Better Data

Leadership Team

Idriss Soumare

Idriss Soumare

Founder, CEO

LinkedIn profile for Idriss Soumare
Peter Szalai

Peter Szalai

Head of Development

LinkedIn profile for Peter Szalai
Andreas Björklund Nelson

Andreas Björklund Nelson

Head of Operations

LinkedIn profile for Andreas Björklund Nelson
Kostya Kanishchev

Kostya Kanishchev

Head of Data Engineering / ML

LinkedIn profile for Kostya Kanishchev

Let's put CapData to work in your data stack