Ambry Genetics governs clinical data and builds a PHI-free environment with Collate
Of hereditary cancer data brought under one governance platform
Data assets governed across six PHI-free MySQL databases in the first production environment
Classification tags applied across all source data assets through Collate automations
Healthcare / Clinical Genetic Diagnostics
MySQL, AWS, GCP, clinical reference databases, bioinformatics pipelines
Governing clinical data that never stops changing
In clinical diagnostics, a patient case is rarely static. At Ambry, a case is a composite of biological data (genetic variants, gene-disease relationships) and non-biological commercial data (ordering physician, billing, reimbursement). Nearly every element can change over time. That mutability, layered on top of strict regulatory obligations, made coherent storage, discovery, and compliance a persistent challenge.
Variants are constantly reevaluated as new clinical and academic evidence arrives; gene-disease relationships and underlying reference information (NCBI builds, ClinVar, HGMD, OMIM) also shift. Keeping a single case coherent as its biological facts change is a moving target.
Beyond the biology, ordering information changes too. Physicians in the ordering organization, patient billing, and reimbursement details all evolve. Unifying clinical and commercial data in the same case is already difficult and is compounded further by this change over time.
Clinical data at Ambry and Tempus is subject to CAP and CLIA lab accreditation standards, HIPAA/PHI protections, and FDA electronic-records regulations. Sharing data safely while meeting all of these requirements for protected health information sets a high bar for governance.
Business units had long requested a completely separate, standalone environment stripped of PHI: de-identified data they could use for research, internal metrics, and analysis without touching patient identifiers.
Governed clinical data products with Collate
Ambry adopted Collate as the governance layer for a formal clinical data product lifecycle, running from one-line registration through metadata operations, quality, and consumption. The team leaned on its built-in automations to make each new data product coherent and compliant by default.
Every novel data product starts with a deliberately simple one-line registration, then moves through metadata operations and quality checks. Establishing roles, policies, and stewardship from the outset makes ingestion and creation of subsequent data products easier and lays the groundwork for governance.
Post-registration, Ambry leans on Collate's profiler, lineage, and auto-classification metadata agents. Lineage is essential for showing exactly where a data product's underlying data was sourced, when it was stored, its data types, and any constraints.
Working with the assigned data steward and data product owner, teams fit quality dimensions to each registered data product and incorporate regulatory adherence directly in Collate, supporting the traceability and accuracy of every product against a FAIR-style framework (Findable, Accessible, Interoperable, Reusable).
Using Collate, Ambry defined a de-identified, anonymized environment across six PHI-free MySQL databases and 11,000+ assets. With Collate's tag and lineage functions, and with classification and tag categories defined alongside compliance and legal, the team can track all identified PHI elements and reconcile the PHI-free databases against production.
A trusted, PHI-free foundation for clinical research at scale
Ambry's team describes the clinical data product concept as a significant improvement in how clinical data is stored, shared, and safely transmitted. The PHI-free environment is the clearest evidence of that shift and has changed how quickly researchers can get to the data they need to deliver new medical discoveries.
The new PHI-free environment empowers teams to quickly pull variant-level data for research, internal metrics, and analysis without risk to protected health information. It has been well received for saving people time in their work.
Data products give Ambry a high-level domain map and single source of truth, where internal domains and subdomains reconcile against their corresponding products, with lineage, classification, and version information supplied through the platform.
Reconciling the PHI-free databases against real production improved data quality checks and data fidelity between the two environments, so teams can trust that what they see in the PHI-free environment matches what production holds.
Ambry is standing up a clinical data governing body and unifying the Ambry and Tempus governance approaches, as part of an ongoing effort to strengthen the functioning and storage of clinical data across their combined organizations.
