MS MRI Workflow
For a multiple sclerosis MRI dataset, pseudonymize the DICOM files first, then audit their metadata with dicomqc. Keep the original files under restricted access.
Pseudonymize a copy, audit it with dicomqc, and review the results. If problems remain, fix the pseudonymization process and run it again.
Recommended directory layout
project/
raw_mri/ # immutable source copy; restricted access
work/
candidate_release_mri/ # pseudonymized .dcm output
maps/
pseudonym_map.tsv.gpg # protected; never shipped with released data
reports/
dicomqc/
report.json
findings.csv
dicomqc_mqc/
If the provider keeps a file linking patient identifiers to research pseudonyms, the dataset is pseudonymized. Keep that sensitive mapping file with the provider, separate from the shared dataset and public repositories, with restricted access.
Keep the files in DICOM format during pseudonymization and audit. Convert to NIfTI or BIDS later, after reviewing the metadata results.
Step 1: inventory and preserve the raw data
Preserve the original files. You can scan them to find privacy risks before pseudonymization. Keep those reports restricted: paths and tag details may still reveal sensitive information.
dicomqc scan raw_mri/ \
--json reports/dicomqc/raw-risk-inventory.json \
--csv reports/dicomqc/raw-risk-findings.csv
Step 2: pseudonymize and de-identify with an external tool
Create pseudonymized DICOM files in candidate_release_mri/.
Possible options include:
- DCMTK commands such as
dcmodifyfor targeted metadata edits - Orthanc de-identification routes if the site already uses Orthanc
- XNAT anonymization or pseudonymization scripts if the site already manages DICOM through XNAT
- a validated institutional de-identification pipeline
- a project-specific
pydicomscript maintained outside dicomqc
For this project, the pseudonymization/de-identification workflow should at minimum address:
- direct subject identifiers
- accession and clinical workflow identifiers
- dates, according to the release policy
- site and institution identifiers
- operator and physician names
- private/vendor tags
- protocol names that may contain subject, site, or study information
- pseudonym consistency across subject, visit, study, and series
Step 3: audit before sharing
Run dicomqc on the pseudonymized .dcm output:
dicomqc scan work/candidate_release_mri/ \
--json reports/dicomqc/report.json \
--csv reports/dicomqc/findings.csv \
--multiqc reports/dicomqc/dicomqc_mqc
Interpret the result:
- exit code
0: metadata checks passed - exit code
1: warnings need review - exit code
2: errors or unreadable files block release
Step 4: remediate findings and rerun
Use the recommended fixes to update the pseudonymization process. Generate a new copy of the dataset and rerun dicomqc so the fixes apply to every affected file.
The same input and settings should give the same audit result:
same raw input + same pseudonymization policy -> same dicomqc result
MRI-specific privacy risks
dicomqc audits metadata. It does not inspect pixel data and does not deface brain MRI volumes.
For head MRI, de-identification may also require image-level handling such as defacing, skull stripping, or another approved facial-feature protection method, depending on the release policy and downstream analysis needs.
If you convert the dataset to BIDS, audit the DICOM metadata first and run a BIDS validator after conversion. dicomqc does not check BIDS requirements; a BIDS profile is planned.
Release checklist
Before sharing the MS MRI dataset, require:
- raw data preserved under restricted access
- candidate release generated by a reproducible pseudonymization/de-identification workflow
- pseudonym linkage map retained by the provider, stored separately, and protected
- dicomqc JSON, CSV, and MultiQC reports archived
- errors resolved
- warnings reviewed and signed off
- private-tag handling documented
- pixel/facial de-identification handled by an appropriate imaging workflow
- final release reviewed under the institutional data-sharing policy