OMOP-CDM
Category: Observational health data model
OMOP CDM is OHDSI's standardized relational model for observational health data. Convert-Pheno reads SQL or table-file exports rather than connecting to a live OMOP database, and it can write the supported OMOP-CDM tables as CSV.
--stream for large OMOP-to-BFF runsSupported directions are:
- OMOP input: PostgreSQL-style SQL dumps or CSV/TSV table exports supplied as files, one directory, or one ZIP package, with optional gzip compression
- OMOP output: CSV tables written from the normalized BFF representation
The implemented mappings target OMOP CDM 5.3 and 5.4.
For the evaluated OMOP-CDM to Beacon v2 workflow, see the published study and its tutorial repository.
Longitudinal OMOP data
VISIT_OCCURRENCE.visit_occurrence_id is retained under the BFF _visit
extension because Beacon v2 individuals has no direct equivalent for OMOP
visits. The extension is emitted only when VISIT_OCCURRENCE is supplied.
OMOP As Inputβ
Download the tested OMOP CSV inputs: PERSON, CONCEPT, and DRUG_EXPOSURE. Keep table files together and pass them as one OMOP input request.
Accepted source layouts are:
- one
.sqlor.sql.gzdatabase export - one or more
.csv,.tsv, or individually gzipped table exports - one directory containing the table exports
- one
.zippackage containing the table exports
For table packages, the table name is inferred from each filename. Keep only
one file per OMOP table in a package. A practical conversion normally needs
PERSON, CONCEPT or --ohdsi-db, and at least one clinical table such as
CONDITION_OCCURRENCE, DRUG_EXPOSURE, MEASUREMENT, OBSERVATION, or
SPECIMEN.
convert-pheno -iomop omop-export/ -obff individuals.json
convert-pheno -iomop omop-export.zip -obff individuals.json
When producing datasets or cohorts, an optional compact Mapping V2 file can
set project and entity metadata without changing the built-in OMOP field
mapping. See Compact Dataset And Cohort Metadata.
ZIP entries are extracted into a private temporary directory for the duration
of the conversion. This lets --stream use the same bounded-memory table
reader as ordinary files.
Use OMOP input recipes for individuals, biosamples, Phenopackets, and streaming commands.
Input Controlsβ
--max-lines-sqldefaults to500. For SQL dumps it limits rows read per selected table; in non-streaming conversion it also limits emitted individuals. Raise it for complete runs.--omop-tableslimits processing to named clinical tables.PERSONandCONCEPTremain required and do not need to be listed.--ohdsi-dbchecks the Athena-OHDSI lookup database when an OMOP input concept is not available in the suppliedCONCEPTtable. It is required for OMOP output, where it resolves active standard concepts in the expected OMOP domain.--mapping-filecan supply optional dataset and cohort metadata for entity-aware BFF output. It is not required for OMOP conversion.- Gzip input and output reduce storage but add CPU overhead, which can be noticeable during streaming.
SPECIMEN and Biosamplesβ
OMOP SPECIMEN rows can become first-class Beacon biosamples. When
SPECIMEN.quantity is present, it is also emitted as a sample-level
biosamples.measurements entry. The unit is resolved from
unit_concept_id, with unit_source_value as a fallback label.
OMOP SPECIMEN has no measurement-concept field for quantity, so the Beacon
assayCode uses the local CURIE OMOP:SPECIMEN.quantity with label
Specimen quantity.
When an identifier is built from CONCEPT.vocabulary_id and
CONCEPT.concept_code, whitespace in the vocabulary prefix becomes an
underscore. For example, Type Concept becomes
Type_Concept:OMOP4976929. This keeps identifiers usable in APIs and indexes.
Source Provenanceβ
Original rows are retained by default under fields such as
info.PERSON.OMOP_columns and
biosamples.info.SPECIMEN.OMOP_columns. This supports mapping review and
source-level queries. Use --no-source-info to omit these copied rows.
Memory and Streamingβ
The default --no-stream mode groups all rows for each person_id in
memory. This produces consolidated individual records but memory use grows with
the number of rows and tables.
--stream processes clinical rows incrementally and is intended for large
exports. It currently supports individuals-only BFF output. PERSON and
CONCEPT are still loaded into memory, as is VISIT_OCCURRENCE when
provided. Streamed files are not consolidated into one record per person;
linkage is retained through identifiers for downstream storage.
Generated BFF can be checked with Beacon v2 CBI Tools.
OMOP As Outputβ
OMOP output consists of table CSV files such as PERSON.csv,
CONDITION_OCCURRENCE.csv, MEASUREMENT.csv, and other supported tables.
The files are written under --out-dir; --out-name TABLE=FILE can override
an individual table filename.
BFF, PXF, CSV, REDCap, CDISC-ODM, cBioPortal, Dataset-JSON, Dataset-XML, FHIR/mCODE, i2b2, PCORnet, and Sentinel input can reach this output through the supported normalization routes. See Choose a Conversion for commands and BFF to OMOP for the field-level mapping.
Developers using in-memory or HTTP(s) interfaces should use the dedicated Module and API pages rather than format-page payload examples.