Skip to main content

Compare source and de-identified files

dicomqc compare compares original source DICOM files with processed candidate DICOM files. It checks that de-identification accounts for every input file and uses patient pseudonyms consistently. It does not compare completed dicomqc runs or their reports.

In the desktop app, select Dataset comparison in Setup and choose the source folder, candidate folder, and pairing CSV. This is separate from Privacy audit; adding a project policy checks candidate metadata only. The walkthrough below uses the CLI and the same synthetic fixtures as the desktop comparison example.

Try the comparison demo​

1. Generate the data​

dicomqc demo --compare --output-dir comparison-demo

The demo creates three source files for two patients, a flawed candidate dataset, and a separate corrected dataset. All names and identifiers below are synthetic. pairs.csv pairs the source and output filenames in this order:

Source fileSource PatientIDOutput fileCandidate PatientIDCorrected PatientID
patient-a-visit-1.dcmLOCAL001image-001.dcmsub-001sub-001
patient-a-visit-2.dcmLOCAL001image-002.dcmsub-099sub-001
patient-b-visit-1.dcmLOCAL002image-003.dcmFile missingsub-002

The command writes these datasets to source/, candidate/, and corrected/ under comparison-demo/. It also runs both comparisons and writes before.html, before.json, before.csv and their after equivalents. The generator creates both examples; the audit itself never edits DICOM files.

2. Inspect the failing comparison​

To rerun the first comparison:

dicomqc compare comparison-demo/source comparison-demo/candidate \
--manifest comparison-demo/pairs.csv \
--html comparison-demo/before.html \
--json comparison-demo/before.json \
--csv comparison-demo/before.csv

This exits 2: only two of the three pairs are readable, and the report contains three error findings in two issue groups:

  • One missing candidate file: manifest row 3.
  • Inconsistent pseudonyms for one source patient: one finding each for rows 1 and 2.

Both existing candidate files pass the built-in metadata checks on their own. The missing output and split patient identity only emerge when comparing the datasets. Open comparison-demo/before.html directly in a browser; MultiQC is not needed. Reports use manifest-row references instead of the synthetic IDs shown in the input table above.

View the comparison report before corrections

HTML comparison report showing three error findings grouped into missing output and inconsistent patient pseudonyms.

JSON and CSV from this comparison​

Example JSON finding — synthetic comparison demo

This is the missing-file entry in before.json's findings array, not the full report. pair-000003/candidate identifies the third manifest row without exposing the source patient ID.

{
"keyword": null,
"message": "The listed candidate file is missing.",
"path": "pair-000003/candidate",
"profile_id": "dataset-comparison-v0.1",
"recommendation": "Check the manifest and regenerate missing output files.",
"rule_id": "dataset-comparison-v0.1.missing_candidate",
"severity": "error",
"standard_refs": [],
"tag": null,
"value_state": "absent"
}
Example CSV rows — synthetic comparison demo

The three findings in before.csv, shown as a table with selected columns for readability. The CSV also includes rule and profile IDs, tag, keyword, value state, and standards references.

pathseveritymessagerecommendation
pair-000003/candidateerrorThe listed candidate file is missing.Check the manifest and regenerate missing output files.
pair-000001/candidateerrorOne source patient maps to multiple PatientIDs.Use the same pseudonym for every file belonging to this source patient.
pair-000002/candidateerrorOne source patient maps to multiple PatientIDs.Use the same pseudonym for every file belonging to this source patient.

3. Inspect the corrected comparison​

The corrected dataset includes the missing file and uses sub-001 for both visits from LOCAL001. Compare it using the same manifest:

dicomqc compare comparison-demo/source comparison-demo/corrected \
--manifest comparison-demo/pairs.csv \
--html comparison-demo/after.html \
--json comparison-demo/after.json \
--csv comparison-demo/after.csv

This exits 0: all three pairs are readable and there are no findings. Open comparison-demo/after.html to see the passing result. This means the checked metadata and pairings passed, not that every privacy risk was assessed. Keep both demo HTML files in the same directory for their before/after links.

View the comparison report after corrections

HTML comparison report after corrections, showing checks passed and no findings.

The demo command itself exits 0 when both examples are generated with their expected results, and 2 otherwise. Use --force to replace a marked demo directory; unmarked directories are refused. Without --output-dir, the demo writes to dicomqc-demo/. See output formats and HTML review controls for details.

Prepare a pairing manifest​

Create a CSV file outside both DICOM directories:

source,candidate
patient-a/visit-1/image.dcm,sub-001/visit-1/renamed.dcm
patient-a/visit-2/image.dcm,sub-001/visit-2/renamed.dcm

Paths are relative to the source and candidate directories respectively. Each file must appear exactly once on its side of the manifest. The directories must be separate, contain only the files being audited, and have no symbolic links. Files do not need a .dcm extension. Keep the manifest restricted: its paths may identify patients.

Have the de-identification process record these pairs as it writes output files. dicomqc trusts that pairing; it cannot prove that two listed files correspond. It does not guess from filenames or UIDs, which may change during de-identification.

Run the comparison​

dicomqc compare raw_mri/ candidate_mri/ \
--manifest pairs.csv \
--json comparison.json \
--csv comparison-findings.csv \
--html comparison.html

Save reports outside both input directories. The command reads metadata and does not modify the DICOM files or the manifest.

It checks for:

  • missing files listed in the manifest;
  • source or candidate files omitted from the manifest;
  • listed files that cannot be read as DICOM;
  • missing or unchanged PatientID values;
  • one source patient receiving different output IDs;
  • different source patients receiving the same output ID.

It also runs the existing research-release-v0.1 privacy checks on each readable, listed candidate file. Original source files are not expected to pass those checks. Add --policy FILE for project-specific checks on candidate files.

Source patients are grouped by IssuerOfPatientID and PatientID. If the issuer is absent, equal source IDs are treated as the same patient. Use a single source ID namespace or populate issuers before comparing data from multiple sites. Output PatientID values must be unique across source patients in this comparison. Patient name consistency is not checked; output names still receive the ordinary pseudonym-format check.

Read the results​

Exit codes match scan: 0 means no findings, 1 means warnings, and 2 means errors, unreadable files, or invalid inputs. Empty manifests and duplicate pairings are rejected instead of producing a clean result. Parser warnings also stop a file from passing. Their original text is suppressed because it may contain identifiers; inspect affected files in a restricted environment.

Comparison reports omit original paths, identifiers, UIDs, and other raw metadata. A reference such as pair-000002/candidate points to the second data row of the manifest (excluding its header). Use the restricted manifest to locate that file. For files omitted from the manifest, unpaired-source-000001 or unpaired-candidate-000001 identifies the first omitted relative path in sorted order on that side. These references are local to this comparison, not stable identifiers across different manifests or inventories.

The JSON comparison section counts source files, candidate files, manifest pairs, readable pairs, and pairs checked for identity consistency. The usual scan counts refer to readable, listed candidate files. Dataset errors can therefore exist even when all scanned candidates pass their individual checks.

In HTML, Charts and audit details shows pairing coverage: both files readable, a missing file, or an unreadable file. Missing takes precedence when a pair also has an unreadable file; files outside the manifest are excluded from this chart. These are coverage counts, not privacy verdicts: a readable pair can still have findings.

Limits of this first implementation​

This comparison expects one output file for every source file. It does not yet support intentional exclusions, splitting or merging files, date-shift checks, UID-reference checks, pixel comparisons, or MultiQC output for comparisons. Equal file counts alone do not establish completeness: the manifest must cover both inventories and every listed file must be readable.

A consistent pseudonym is not proof of anonymity. Review the findings and the remaining privacy risks before sharing data.