---
title: "Masking variable names"
vignette: >
  %\VignetteIndexEntry{Masking variable names}
  %\VignetteEncoding{UTF-8}
  %\VignetteEngine{quarto::html}
knitr:
  opts_chunk:
    collapse: true
    comment: '#>'
editor_options: 
  chunk_output_type: console
---

In certain studies, variable names should be masked to prevent researcher bias. Examples can include exploratory factor analysis, network analysis, etc. The `vazul` package provides means to mask variable names in a dataset, ensuring that analyses can be conducted without preconceived notions about the variables.

```{r}
#| label: setup
#| message: false

library(vazul)
library(dplyr)
library(stats)

```

In this example, we will use the `williams` dataset from the `{vazul}` package.

```{r}
data("williams", package = "vazul")

head(williams)
glimpse(williams)

```

We will apply masking to the variables related to life history strategy, which are prefixed with `SexUnres`, `Impuls`, `Opport`, `InvEdu`, and `InvChild`. We'll mask each variable group separately with randomized letter prefixes (e.g., `C_01`, `A_01`, `E_01`, etc.). This way, variables within the same original scale keep a common prefix for the analysis, but analysts won't know which prefix corresponds to which original scale due to the randomization.


```{r}
set.seed(84)

# Sample 5 random letters for the 5 variable groups
random_prefixes <- paste0(sample(LETTERS, 5), "_")

masked_williams <-
    williams |> 
    mask_names(starts_with("SexUnres"), prefix = random_prefixes[1]) |>
    mask_names(starts_with("Impul"), prefix = random_prefixes[2]) |>
    mask_names(starts_with("Opport"), prefix = random_prefixes[3]) |>
    mask_names(starts_with("InvEdu"), prefix = random_prefixes[4]) |>
    mask_names(starts_with("InvChild"), prefix = random_prefixes[5])

# Show the randomized prefixes used (but not which corresponds to which)
sort(unique(sub("_.*", "_", grep("^[A-Z]_", names(masked_williams), value = TRUE))))
```

We can now perform an exploratory factor analysis (EFA) on the masked variables. Since the variable names are masked with randomized prefixes, we won't know which original variables correspond to which factor, thus preventing bias in interpreting the results.

```{r}

set.seed(123)
efa_blind <-
    masked_williams |> 
    select(matches("^[A-Z]_")) |>
    factanal(factors = 5, rotation = "varimax")
    
# Get the loadings of the EFA on the masked data
efa_blind |> 
    loadings() |> 
    print(cutoff = 0.3, sort = TRUE)

```

The loading table shows that factors are not necessarily loading to their original category. By using masking, researchers may be able to make decisions without being biased on the variable names.

Applying the same analysis on the original dataset reveals the names of the variables. Please note that the loadings may differ slightly due to the randomness in the factor analysis process.

## Preserving a fixed suffix with `keep_suffixes`

Some of the `williams` columns end in `_r`, marking items that are reverse-scored (e.g.
`SexUnres_4_r`, `Impuls_2_r`). That suffix is not itself sensitive information to hide — it is
a fixed analysis-relevant tag that a researcher needs to see in order to reverse-code the item
correctly before scoring, even while the rest of the name stays masked. By default, `mask_names()`
masks the suffix away along with everything else, so `SexUnres_4_r` becomes an opaque `C_02`
with no indication that it needs to be reverse-coded.

The `keep_suffixes` argument preserves one or more literal suffixes verbatim in the masked name,
so the reverse-coding information survives masking:

```{r}
set.seed(84)
masked_with_suffix <-
    williams |>
    mask_names(starts_with("SexUnres"), prefix = "C_", keep_suffixes = "_r")

masked_with_suffix |>
    select(matches("^C_")) |>
    names()
```

Notice that the two reverse-scored columns keep their `_r` suffix (e.g. `C_01_r`) while the
rest of the name is masked as usual, and that masked columns are also sorted alphabetically by
their masked name. If more than one supplied suffix could match the same column, the longest
match is kept and `mask_names()` issues a warning naming the affected column(s).

`keep_suffixes` covers the common case of a single, fixed suffix that should always be preserved
literally (not remapped to a randomized label). When *both* the prefix and the suffix carry
meaning that must be masked - but consistently, so the same value always maps to the same masked
label - the long-format approach shown next is the right tool instead.

## Masking names with meaningful prefixes *and* suffixes

Sometimes both the prefix and the suffix of a set of column names carry meaning that should
stay consistent across columns. For example, `exp_pre`, `exp_post`, `ctl_pre`, and `ctl_post`
encode both an experimental condition (`exp`/`ctl`) and a measurement occasion (`pre`/`post`).

Masking these names directly with repeated `mask_names()` calls, as above, does not preserve
that consistency: each call assigns new labels independently, so there is no guarantee that,
say, `pre` is masked to the same label in the `exp_` columns as in the `ctl_` columns. In this
situation we recommend reshaping the data to long format first, so that the condition and
occasion information become *values* in their own columns rather than fragments of the column
names. `mask_variables()` can then mask those columns directly, which guarantees that a given
value (e.g. `"exp"` or `"pre"`) is always mapped to the same masked label everywhere it occurs.
If a wide format is required for the confirmatory analysis, the masked data can be reshaped
back afterward.

```{r}
#| message: false
library(tidyr)

wide_demo <- data.frame(
  id = 1:4,
  exp_pre = c(10, 12, 9, 11),
  exp_post = c(15, 14, 13, 16),
  ctl_pre = c(8, 9, 10, 7),
  ctl_post = c(9, 10, 11, 8)
)
wide_demo
```

```{r}
# 1) Reshape to long format, splitting the column names into
#    a `condition` and a `time` variable
long_demo <- wide_demo |>
  pivot_longer(
    cols = -id,
    names_to = c("condition", "time"),
    names_sep = "_"
  )
long_demo
```

```{r}
# 2) Mask the condition and time variables. Each variable keeps its own,
#    internally consistent mapping (every "exp" becomes the same masked
#    label, every "pre" becomes the same masked label, and so on).
set.seed(2024)
long_masked <- long_demo |>
  mask_variables(condition, time)

long_masked |>
  count(condition, time)
```

```{r}
# 3) Reshape back to wide format if a wide layout is needed for analysis
long_masked |>
  unite(masked_name, condition, time) |>
  pivot_wider(names_from = masked_name, values_from = value)
```

Because masking is applied to `condition` and `time` as data values rather than as fragments
of the column names, the masked labels for `"exp"`/`"ctl"` and `"pre"`/`"post"` are consistent
across every column, both before and after reshaping. We recommend this long-format workflow
whenever meaningful prefixes and suffixes both need to be masked; `mask_names()` remains the
simpler choice when only a single, non-overlapping piece of the name needs masking, and
`mask_names(..., keep_suffixes = ...)` covers the case of a fixed literal suffix that should
be preserved as-is rather than masked (see above).

```{r}
set.seed(123)
efa_orig <-
    williams |> 
    select(SexUnres_1:InvChild_2_r) |>
    factanal(factors = 5, rotation = "varimax")

# Get the loadings of the EFA on the original data

efa_orig |> 
    loadings() |> 
    print(cutoff = 0.3, sort = TRUE)

```




