Real-world data is fragmented across sources that don't connect

Category: Data & AI Infrastructure

Claims, EHR, lab, and registry data each describe a piece of the patient — but they sit in separate systems with separate formats, so no single source answers a research question end to end.

Every real-world data source covers a slice of reality: claims capture billed encounters, EHR data captures what one health system saw, labs and registries capture their own domains. A research or commercial question — how patients with a condition are actually diagnosed, treated, and switched over time — almost always spans several of these slices. Assembling them in-house means licensing multiple sources, reconciling incompatible formats and coding systems, and building linkage infrastructure most teams don't have. Real-world data platforms exist to do that assembly once, at scale, so analysis starts from connected data rather than a data-engineering project.

How data platforms address fragmentation

Platform vendors aggregate many raw sources — claims clearinghouses, EHR networks, laboratory feeds, specialty registries — and do the unglamorous work that makes them usable together: normalizing to common data models, mapping local codes to standard vocabularies, de-duplicating records, and linking events that belong to the same de-identified patient across sources. The output is a longitudinal view that no single source provides, delivered through query tools or as analysis-ready extracts into the buyer's own environment. The honest caveat is that no assembled dataset is complete. Coverage varies by geography, care setting, and therapeutic area, and linkage across sources is probabilistic rather than perfect. Good platforms are explicit about what their data does and does not capture; the evaluation question is whether the coverage matches your specific question, not whether the total patient count is impressive.

What to evaluate before licensing connected data

Coverage in your therapeutic area is the deciding factor, and it has to be tested concretely: define a cohort you know well — from a trial, a registry, or prior research — and ask the vendor to profile it. Check how many patients they find, how deep the longitudinal history goes, and how quickly new data arrives, because data lag determines whether the platform can answer current questions or only historical ones. Then look at the plumbing: which common data model the platform uses, whether extracts land in the tools your analysts already work in, and how the vendor documents its linkage and de-duplication methods. A platform that cannot explain how it decides two records are the same patient is asking you to trust the hardest step blind.

How teams typically get started

A focused pilot beats a broad license: pick one live question your team already needs answered, run it on the platform's data, and compare the result against whatever internal or published benchmark exists. That exposes coverage gaps, data lag, and workflow fit on a real problem before a multi-year commitment. Teams that skip this step tend to discover the limitations after the contract is signed, when the first serious analysis comes back thinner than the sales deck suggested.

AI Use Cases That Address This Problem

  • Patient Identification & Segmentation
  • Patient Recruitment & Enrollment

Frequently asked questions

Why not license the raw sources and integrate them ourselves?

Some large organizations do, but the integration work — normalization, vocabulary mapping, de-duplication, privacy-preserving linkage — is substantial, ongoing, and outside most teams' core mission. The build-versus-buy question usually turns on whether connected real-world data is a strategic asset you will exploit continuously or an input you need for periodic questions.

How do platforms link records without identifying patients?

Most use privacy-preserving tokenization: identifying fields are transformed into irreversible tokens by certified third-party processes, and records sharing a token are linked without anyone handling the underlying identity. Ask any vendor to walk through their specific method, its certification, and its measured accuracy — linkage quality varies and directly affects analysis validity.

What does 'coverage' actually mean in vendor claims?

Headline patient counts usually describe every patient who appears anywhere in the data, however briefly. What matters for research is how many patients have the longitudinal depth your question needs — continuous observation, the right care settings, the right data types. Always evaluate coverage against your specific cohort definition, not the headline number.

Is fragmented data ever the right choice?

Sometimes a single deep source beats several connected shallow ones — a specialty registry may capture clinical detail no claims dataset holds. The right architecture follows the question: connected breadth for journey and utilization questions, single-source depth for clinical nuance, and often both in combination.

AI Vendors for This Problem

Evidence & Outcomes

Intercept used Komodo's real-world data for PBC (Ocaliva) analyses

Intercept Pharmaceuticals worked with Komodo Health to apply its Healthcare Map real-world data to primary biliary cholangitis, supporting patient-journey and treatment-pattern analyses for its Ocaliva franchise.

Vendor: Komodo Health · press release (Komodo Health (Business Wire)) · 2022-11-10 — Partner: Intercept Pharmaceuticals

Launched the Verana Research Network on the AAO IRIS Registry to accelerate clinical research

Verana Health and the American Academy of Ophthalmology launched the Verana Research Network, an IRIS Registry initiative in which academic medical centers and ophthalmology practices use IRIS Registry data and Verana's platform to increase clinical trial access, accelerate recruitment, and advance ophthalmic research.

Vendor: Verana Health · press release (Verana Health) · 2022-12-07 — Partner: American Academy of Ophthalmology