Every regulatory submission and registry analysis depends on the same upstream artifact: an assembled patient timeline. We construct it once, correctly, with provenance, and run real-world data construction as infrastructure rather than a service.
Extraction and normalization have vendors, standards and their own budget lines. Assembly gets absorbed by research coordinators and principal investigators, and there is no published benchmark for it anywhere in the literature.
Pulling structured facts out of clinical notes, scanned documents, lab reports and sponsor PDFs.
Mapping local codes, units and vocabularies onto a shared representation.
Deciding which records belong to the same person, in what order events occurred, which of two conflicting values is right, and when to decline to answer.
Three of the most capable groups in healthcare data converged independently on agentic construction. Each financed it as a proprietary asset, and none of them sells the layer on its own.
Built with capital
raised by Truveta to normalize EHR data across roughly thirty health systems with a proprietary clinical language model.
curated weekly by ConcertAI's agent hierarchy, across 13 million patient records.
of U.S. hospital beds run on Epic, where 85% of customers now use its AI tools, though only on data that has touched an Epic instance.
Everywhere else
hospitals still pay trained clinicians to abstract records by hand, several hours per patient.
rare disease registries catalogued worldwide, fragmented and unable to pool.
sells construction on its own. It is either locked inside a platform or rebuilt by hand for each new study.
Construction is model-agnostic and runs once per patient, while projection is model-specific and runs per request. The same assembled timeline can serve a prediction model, an external control arm and a regulatory submission without being rebuilt for each one.
Runs once per patient, independent of any model. Codes, units and vocabularies are normalized, then events are ordered, deduplicated and resolved into a single canonical stream that every downstream model reuses.
Runs per request, shaped to whichever model is asking. The canonical timeline is reprojected into the schema an analysis requires, with provenance attached to every output.
Synthetic environments are built ground truth first. A clean master record is degraded into source views with every operation logged, so the answer key is generated by diffing. That makes identity resolution, abstention calibration and temporal alignment scoreable instead of asserted.
A pilot is running with a neurodevelopmental rare disease foundation, across registry enrollment, clinic encounters, an external natural history study and contributed trial datasets held under differing consent regimes.

Chief Executive Officer. Computer science, applied mathematics and statistics, and computational bioengineering at Johns Hopkins.

Chief Technology Officer. Postdoctoral fellow at Harvard Medical School, previously at Truveta, with a PhD in biomedical informatics from the University of Washington.
Affiliations