FlagshipIn use with clients

Synthetic cohorts that carry the structure of real histories.

The Patient Synthesizer learns from your longitudinal data and generates synthetic patients that hold up statistically without representing any real person. Training runs inside your environment. What leaves the building is your decision.

What for

Work before data access is settled.

Build and test

Pipelines, cohort logic and models take shape on synthetic data. Access to the real holdings is only needed once the analysis stands.

Share and publish

Cohorts can be passed to partners, contractors and reviewers without information about individuals leaving the building.

Augment

Rare histories are thinly populated in routine data. Synthetic cases widen the basis for method development and validation.

How it runs

The code travels, not the data.

The Synthesizer is installed at the data holder and generates a synthetic cohort there. The data user develops their analysis on that cohort. For the final result, only the finished code runs against the real data — at the data holder, inside their environment.

Path of the synthetic dataPath of code and results
Software1
MedModelsPatient SynthesizerShipped and installed at the data holder.
Holdings2
Data holderReal longitudinal dataRead in as MedRecord. Never leaves the building.
Generated in house3
Data holderSynthetic cohortAs MedRecord, with a utility and a privacy report.
Handover4
Data userSynthetic cohort at the userReleased by the data holder.
Development5
Data userDevelop the analysisCohort logic, models and checks are built on the synthetic cohort.
Run on real data6
Data holderCode runs at the data holderThe same code, unchanged, against the real MedRecord.
Result7
Data userAnalysis returnedOnly aggregated results leave the data holder.

Schematic representation. The roles may coincide within one organisation.

Utility

What the cohort preserves.

Distributions are compared, not people. For every quantity checked, the report states the deviation between the real and the synthetic cohort.

MetricReal cohortSyntheticDeviation
Patients48,21048,210identical
Median age67 years66 years1 year
Hypertension prevalence (I10)61.4%60.8%0.6 pp
Encounters per patient (median)99identical
Days between encounters (median)74713 days
Observation period per patient (median)4.2 years4.1 years0.1 years
First diagnosis to intervention (median)541 days528 days13 days
Inpatient readmission within 30 days12.7%13.4%0.7 pp
Histories with no encounter for 12 months18.3%17.1%1.2 pp
Sequence I10 before I2171.2%68.4%2.8 pp
Green: deviation within the defined tolerance.

Illustrative example of a report. Values differ per data holding.

Reports

Every run is documented.

Each generated cohort comes with two reports: one on utility, one on data protection. Both are versioned and machine-readable, so they can go straight into your documentation.

Request a sample report
Privacy reportData protection
Re-identification of individual historiesno match
Nearest-neighbour distanceabove threshold
Membership inference testnear chance
Rare attribute combinationssuppressed
Operation

Runs where your data lives.

Installation

A container or a Python package, on your own hardware or in your cloud. No outbound traffic, no licence server call at runtime.

Connection

Reads and writes MedRecord. Existing holdings come in through data import, OMOP mapping or the HL7 FHIR Connector.

Use

Versioned software with support and ongoing development. Model weights trained on your data stay with you.

Placeholder

Description to follow.

More in the portfolio

The Synthesizer is one part of it.

Every product reads and writes the same MedRecord: Analytics Engine, Cohort Selector, OMOP Mapping, HL7 FHIR Connector and the EHDS Tool.