---
title: "homing(): Relinking De-identified Data"
subtitle: "The authorised path back from anonymisation"
author: "Dr Nicolas Smoll, SCPHU, Sunshine Coast Hospital and Health Service"
date: "`r Sys.Date()`"
output:
  html_document:
    toc: true
    toc_depth: 3
    toc_float: true
    theme: flatly
  pdf_document:
    toc: true
    toc_depth: 3
    number_sections: true
    latex_engine: xelatex
vignette: >
  %\VignetteIndexEntry{homing(): Relinking De-identified Data}
  %\VignetteEngine{knitr::rmarkdown}
  %\VignetteEncoding{UTF-8}
---

```{r setup, include=FALSE}
knitr::opts_chunk$set(collapse = TRUE, comment = "#>", warning = FALSE, message = FALSE)
library(mudnester)
```

## When to use `homing()`

`homing()` is the complement to `molting()`. It is used when:

1. A de-identified dataset has been shared or stored, and
2. An authorised person subsequently needs to recover the original identifiers — for example, to contact individuals for follow-up, to correct a data error, or for a notifiable disease regulatory obligation.

The name is precise: homing pigeons navigate back to their loft regardless of where they were released, using an internal compass that only they carry. `homing()` navigates a de-identified dataset back to its identifiers using the lookup table — and only those who hold the lookup table can make that journey.

---

## The basic relink workflow

```{r workflow}
# Construct sample data
patient_data <- data.frame(
  patient_name = c("John Doe", "Jane Smith", "Alice Brown"),
  dob          = as.Date(c("1980-01-01", "1975-05-15", "1992-11-30")),
  mrn          = c("12345", "67890", "11111"),
  diagnosis    = c("Condition A", "Condition B", "Condition C"),
  severity     = c("mild", "moderate", "severe")
)

# Step 1: de-identify (typically done at data collection / storage time)
result <- suppressMessages(molting(patient_data))

# Step 2: share or archive result$deidentified
#         store result$lookup securely, separately

# Step 3: relink when authorised
relinked <- homing(
  deidentified_data = result$deidentified,
  lookup_table      = result$lookup
)

head(relinked)
```

The original identifiers (`patient_name`, `dob`, `mrn`) are joined back in via the `row_hash` column.

---

## Controlling the hash column name

If `molting()` was called with a custom `hash_col_name`, pass the same name to `homing()`.

```{r custom-hash}
result_custom <- suppressMessages(
  molting(patient_data, hash_col_name = "person_hash")
)

relinked_custom <- homing(
  result_custom$deidentified,
  result_custom$lookup,
  hash_col_name = "person_hash"
)

"patient_name" %in% names(relinked_custom)
```

---

## Removing the hash after relinking

If you want a clean re-identified dataset without the hash column:

```{r no-hash}
relinked_clean <- homing(
  result$deidentified,
  result$lookup,
  keep_hash = FALSE
)

names(relinked_clean)   # no row_hash column
```

---

## Partial matches and unmatched records

If the lookup table is incomplete (e.g. some records were excluded from the lookup for a legitimate reason, or the wrong lookup was supplied), `homing()` warns you about unmatched rows and returns them with `NA` in the identifier columns rather than silently dropping them.

```{r partial}
# Simulate a truncated lookup — only the first two rows
partial_lookup <- result$lookup[1:2, ]

relinked_partial <- homing(
  result$deidentified,
  partial_lookup
)

# Third row has NA identifiers
relinked_partial[, c("row_hash","patient_name","diagnosis")]
```

Always check the summary message for the matched count. A significantly lower matched count than expected usually means the wrong lookup was supplied.

---

## Security and governance checklist

Before using `homing()` in a production workflow, ensure:

- [ ] The relink is authorised by your organisation's privacy and governance framework
- [ ] The lookup table was stored separately from the de-identified data, with access logging
- [ ] The relinking event is documented in your data management plan
- [ ] The re-identified dataset is treated as identifiable data from the moment `homing()` returns
- [ ] The re-identified dataset is not stored in the same location as the de-identified data unless access controls are equivalent

In Queensland Health, re-identification for notifiable disease follow-up typically falls under Public Health Act 2005 obligations and does not require separate ethics approval, but document the basis for re-identification in your outbreak log.

---

## What comes next

After re-identification, the data is again fully identifiable. If you need to re-anonymise for a secondary analysis, run `molting()` again. See `vignette("molting")` for options.

If the purpose of the relink was to add clinical follow-up data, the updated dataset can be re-cleaned with `clean_the_nest()` before further analysis.
