Skip to main content

Command-Line Interface

Use convert-pheno to convert local files or run conversions in scripts. For a graphical workflow, use the Desktop App.

Start with your input​

Find a command for your input format, then replace the example paths with your own. New to the tool? The 5-Minute Quickstart includes a downloadable input and expected results.

Check that the CLI is installed:

convert-pheno --version

If the command is unavailable, see Installation.

Read a command​

convert-pheno \
-ipxf phenopacket.json \
-obff individuals.json
  • -ipxf phenopacket.json selects Phenopacket input and its source file.
  • -obff individuals.json selects Beacon output and its destination.
  • Add options only when needed, such as --no-source-info to omit original source-field copies.

This command creates individuals.json in the current directory. Input files are not changed. Use quoted paths when filenames contain spaces.

The compact flags above and the generic form -i pxf phenopacket.json -o bff individuals.json are equivalent.

Default BFF output

For inputs that support BFF, omitting the output option defaults to -obff individuals.json. If you explicitly write -o bff, you must also provide the filename, for example -o bff individuals.json.

Choose the output​

What you wantCommand ending
Beacon individuals in one file-obff individuals.json
Separate Beacon entity files-obff --entities individuals biosamples --out-dir bff-out/
Phenopackets-opxf phenopackets.json
OMOP CSV tables-oomop --out-dir omop-out/ --ohdsi-db
Flattened CSV from BFF or PXF-ocsv records.csv

Choose a target supported by your input; see Supported Formats. For directory output, create the destination first, for example mkdir -p bff-out.

--entities does not replace -obff. It chooses which BFF collections to write. Add datasets or cohorts when needed; biosamples require sample data in the source. Each collection gets its own filename, such as biosamples.json.

OMOP output writes named tables such as PERSON.csv, not a single JSON file. It requires the OHDSI database.

Existing outputs are protected. Add -O only when you intend to overwrite them. Use --out-name biosamples=samples.json to rename an entity output file.

Do you need a mapping file?​

Input or taskWhat to provide
CSV, REDCap, CDISC-ODMA field mapping: --mapping-file mapping.yaml
REDCap CSV or REDCap-origin ODMAlso supply --redcap-dictionary dictionary.csv
Dataset-XMLSupply --define-xml define.xml; optional for Dataset-JSON
BFF to OMOP terminology correctionsAn optional terminology mapping
Dataset ID, name, or cohort metadata for OMOP to BFFA small optional metadata mapping; OMOP field conversion stays built in

See Mapping Files for the appropriate structure. Other built-in routes can also accept optional mappings; their format guides explain the purpose.

Need a dataset ID? Follow the OMOP metadata example. Add --include-dataset-id only when your backend requires top-level datasetId on individuals and biosamples. The ID comes from beacon.datasets.defaults.id in the mapping file, and this also works without generating datasets.json.

Settings worth checking​

NeedOption
Your CSV is comma-separated--separator ','; the .csv default is ;
Smaller BFF without original source copies--no-source-info; mapped fields remain
Review terminology matches and unresolved terms--term-audit review.xlsx (also .tsv or .tsv.gz)
Process all OMOP records rather than the default limit--max-lines-sql 0
Reduce memory use for OMOP-to-BFF--stream; output becomes line-delimited JSON
Set PXF status only when missing from the source--default-vital-status UNKNOWN_STATUS
OMOP input limit

The default --max-lines-sql 500 caps individuals in non-streaming conversions and rows per table in SQL imports. Set it deliberately for your dataset. CSV/TSV streaming is not capped by this option.

Auditing is optional and adds processing time. Start with --search exact (the default), then inspect the report before changing aliases or trying mixed or fuzzy. See Terminology Search for examples and score interpretation.

Streaming is supported for OMOP-to-BFF individuals, biosamples, or both. It does not generate datasets or cohorts. See the OMOP input guide for source layouts and requirements.

More options​

The installed version lists its complete options with:

convert-pheno --help
Terminology and OMOP controls
OptionPurpose
--text-similarity-method cosine|diceToken similarity for mixed/fuzzy; default cosine
--min-text-similarity-score SCOREMinimum mixed/fuzzy score; default 0.8
--levenshtein-weight WEIGHTCharacter-edit contribution to fuzzy scoring; default 0.1
--path-to-ohdsi-db DIRDirectory containing ohdsi.db
--ohdsi-dbUse OHDSI; required for OMOP output, optional for OMOP input
--omop-tables TABLE ...Restrict input tables; PERSON and CONCEPT remain included
--exposures-file FILEOverride the list of OMOP concept IDs treated as exposures
--no-streamExplicitly select in-memory processing
--sql2csvExtract SQL tables rather than convert; incompatible with --stream
Metadata and troubleshooting controls
OptionPurpose
--username NAMEName recorded in conversion metadata
--out-name KEY=FILEOverride an entity or OMOP table filename; repeat as needed
--verboseShow progress
--log [FILE]Save the resolved request/configuration
--debug LEVELPrint diagnostics; level 2 also summarizes SQLite activity
--color / --no-colorEnable or disable terminal colors
--testSuppress time-varying metadata for reproducible tests
--schema-file FILEUse an alternative mapping schema
--self-validate-schemaDeveloper check of the mapping schema itself

If a conversion fails, read the reported error before rerunning. For help, include the tool version, conversion command, and error, using synthetic data and removing private paths or identifiers. See Troubleshooting.