Sharing a linelist externally — with collaborators, for CRAN
examples, in a research supplement — requires that directly identifying
information is removed. molting() automates this using a
cryptographic hash: each row gets a unique fingerprint derived from its
identifiers, the identifiers are stripped, and the fingerprint is stored
in a lookup table that only authorised personnel hold.
The name comes from the natural process of moulting: a bird sheds its distinctive, identifiable plumage and temporarily becomes more uniform. The old plumage is not destroyed — it re-grows from the same follicles. The lookup table is those follicles.
molting() uses regular-expression pattern matching to
detect PII columns. The default patterns cover:
name, surname,
firstname, lastnamedob, birthmrn, urn,
medicare, patient_id, subject_id,
_id$address, street,
phone, emailAge category variables (age2cat,
age5cat, age10cat, etc.) are automatically
preserved — they are not directly identifying.
patient_data <- data.frame(
patient_name = c("John Doe", "Jane Smith"),
dob = as.Date(c("1980-01-01", "1975-05-15")),
mrn = c("12345", "67890"),
age5cat = factor(c("18-64", "18-64")), # preserved automatically
diagnosis = c("Condition A", "Condition B"),
lab_value = c(120, 95)
)
result <- suppressMessages(molting(patient_data))
names(result$deidentified) # hash + retained columns
#> [1] "row_hash" "age5cat" "diagnosis" "lab_value"
names(result$lookup) # hash + removed columns
#> [1] "row_hash" "patient_name" "dob" "mrn"
molting() returns a named list with two elements when
return_lookup = TRUE (the default):
$deidentified — the de-identified data frame, with
row_hash as the first column$lookup — the lookup table, with row_hash
plus all removed identifier columnsstr(result, max.level = 1)
#> List of 2
#> $ deidentified: tibble [2 × 4] (S3: tbl_df/tbl/data.frame)
#> $ lookup : tibble [2 × 4] (S3: tbl_df/tbl/data.frame)
head(result$lookup)
#> # A tibble: 2 × 4
#> row_hash patient_name dob mrn
#> <chr> <chr> <chr> <chr>
#> 1 89573bbf928ef324ba95e8d04fd1701dfc5c15265ab21b3dbbb2… John Doe 1980… 12345
#> 2 7b2536bf2d008eb3555f9404d1762dcc529e1a7d5e6880c66341… Jane Smith 1975… 67890
Store $lookup securely and separately from
$deidentified. Consider encrypting the lookup file before
archiving. In a Queensland Health context, the lookup table should
remain within the Health Service network.
The default is SHA-256, which provides strong collision resistance for typical surveillance dataset sizes (tens of thousands of rows). For very large datasets where speed matters more than collision resistance, SHA-1 or MD5 are faster but should not be used where linkage integrity is critical.
# SHA-256 (default, recommended)
result_256 <- molting(patient_data, hash_method = "sha256")
# MD5 (shorter hash, faster, lower collision resistance)
result_md5 <- molting(patient_data, hash_method = "md5")
# Blake3 (fast and cryptographically strong — good for large datasets)
result_b3 <- molting(patient_data, hash_method = "blake3")
By default, molting() hashes all detected PII columns.
Supply id_cols to override this — useful when you want a
shorter, more stable hash based only on a true unique identifier.
# Hash only on MRN and DOB — more stable if name variations exist
result_ids <- suppressMessages(
molting(patient_data, id_cols = c("mrn", "dob"))
)
result_ids$lookup
#> # A tibble: 2 × 4
#> row_hash patient_name dob mrn
#> <chr> <chr> <chr> <chr>
#> 1 255ad8c4452e80121bdee96a35de7cea0c25475a5e9e6f73088c… John Doe 1980… 12345
#> 2 107510d61f22148ec63b636654c41efea411429ace78b9c371f0… Jane Smith 1975… 67890
Use additional_pii_cols for dataset-specific identifiers
that don’t match the default patterns.
patient_data2 <- patient_data
patient_data2$study_code <- c("SC-001","SC-002")
result2 <- suppressMessages(
molting(patient_data2, additional_pii_cols = "study_code")
)
names(result2$deidentified)
#> [1] "row_hash" "age5cat" "diagnosis" "lab_value"
If you genuinely do not need to relink (e.g. producing a public-use
file), set return_lookup = FALSE. This is irreversible —
there is no way to recover the original identifiers.
deidentified_only <- suppressMessages(
molting(patient_data, return_lookup = FALSE)
)
class(deidentified_only) # a data frame, not a list
#> [1] "tbl_df" "tbl" "data.frame"
If two rows produce the same hash (extremely rare with SHA-256 for
realistic dataset sizes but possible with very short hashes like CRC32),
molting() warns you. If a collision is detected, switch to
a stronger algorithm or add more columns to id_cols.
Use [homing()] to relink the de-identified data when authorised (see
vignette("homing")).
For aggregated outputs — monthly counts from roost() —
de-identification may not be necessary at all if the counts are not
small enough to be re-identifying. The ABS cell-suppression threshold of
5 is a useful rule of thumb.