Package {searchlight}


Title: Audited Analysis of Police Stop and Search Records
Version: 0.1.0
Description: Acquire and audit public police stop and search archives for England and Wales, retaining provenance and explicit coverage information. Designed for exposure-based ethnic disparity analysis with separately reported sampling and assumption uncertainty. Contains public sector information licensed under the Open Government Licence v3.0.
License: MIT + file LICENSE
Encoding: UTF-8
Language: en-GB
Depends: R (≥ 4.1.0)
Imports: CARBayes (≥ 6.1.1), cli, commonmark, digest, dplyr (≥ 1.1.1), httr2, jsonlite, MASS, readr, rlang, sf, spdep, splines, suncalc, posterior, ggplot2, tidyr, tibble, withr, xml2
URL: https://github.com/BlackThrive/searchlight, https://blackthrive.github.io/searchlight/
BugReports: https://github.com/BlackThrive/searchlight/issues
RoxygenNote: 7.3.3
Suggests: testthat (≥ 3.0.0), httptest2, knitr, rmarkdown, lme4, lintr, spelling, vdiffr, zip
VignetteBuilder: knitr
Config/testthat/edition: 3
NeedsCompilation: no
Packaged: 2026-09-22 04:33:08 UTC; musta
Author: Mustapha Wasseja [aut, cre], Sarah Hamed [aut], Souci Frissa [aut], Black Thrive Global [cph, fnd]
Maintainer: Mustapha Wasseja <mustapha.wasseja.mohammed@gmail.com>
Repository: CRAN
Date/Publication: 2026-09-30 11:40:02 UTC

searchlight: audited police stop and search analysis

Description

Public archive records describe events, not unique people. Missing submissions must not be treated as zero events. Self-defined and officer-defined ethnicity are distinct measurements. See the installed NOTES directory for decisions and verified sources.

Author(s)

Maintainer: Mustapha Wasseja mustapha.wasseja.mohammed@gmail.com

Authors:

Other contributors:

See Also

Useful links:


Download and verify the necessary archive snapshots

Description

Selects the newest available snapshots that cover the requested months. Forces are recorded as extraction filters; the source only offers full ZIPs. Checks publisher MD5 before recording local SHA-256. Existing files must match both the publisher and any recorded SHA-256; corrupt files are never replaced silently. The manifest and ZIPs are written to dir.

Usage

sl_archive_download(months, forces = NULL, dir = sl_cache_dir(), index = NULL)

Arguments

months

Character vector of YYYY-MM months.

forces

Force identifiers, or NULL for all forces.

dir

Cache directory.

index

Optional previously retrieved archive index.

Value

Manifest tibble, or NULL when a resource is unavailable.

See Also

sl_archive_snapshot(), sl_list_versions()

Other ingest: sl_archive_index(), sl_archive_snapshot(), sl_assign_geography(), sl_boundaries(), sl_cache_clear(), sl_cache_dir(), sl_contract(), sl_exposure(), sl_list_versions(), sl_population(), sl_population_crosstab(), sl_read_records(), sl_sample(), sl_select_version()

Examples

index <- readRDS(system.file("extdata", "archive-index.rds",
  package = "searchlight"
))
index$archive_file[1]

List available bulk archive snapshots

Description

Parses advertised month ranges and publisher MD5 checksums. The complete rolling snapshot can contain several years. Results are cached only after this function is called. Source failures return NULL with a message.

Usage

sl_archive_index(
  dir = sl_cache_dir(),
  refresh = FALSE,
  url = "https://data.police.uk/data/archive/"
)

Arguments

dir

Cache directory.

refresh

Refresh the cached index.

url

Archive listing URL, normally the official source.

Value

A tibble with archive name, snapshot month, coverage range, URL and MD5, or NULL if unavailable.

See Also

sl_archive_download()

Other ingest: sl_archive_download(), sl_archive_snapshot(), sl_assign_geography(), sl_boundaries(), sl_cache_clear(), sl_cache_dir(), sl_contract(), sl_exposure(), sl_list_versions(), sl_population(), sl_population_crosstab(), sl_read_records(), sl_sample(), sl_select_version()

Examples

index <- readRDS(system.file("extdata", "archive-index.rds",
  package = "searchlight"
))
index[, c("archive_file", "month_start", "month_end")]

Extract immutable stop and search CSV snapshots

Description

Other archive contents are not extracted. Retains source ZIPs and raw CSV bytes. A manifest records archive and CSV SHA-256, row count and original member path. ZIP traversal paths are refused. Previously extracted CSVs are hash-checked.

Usage

sl_archive_snapshot(dir = sl_cache_dir(), months = NULL, forces = NULL)

Arguments

dir

Cache containing archive ZIPs and optionally a download manifest.

months

Optional YYYY-MM filter; defaults to the download request.

forces

Optional force filter; defaults to the download request.

Value

A tibble of CSV versions. Writes snapshots and their manifest in dir.

See Also

sl_select_version(), sl_read_records()

Other ingest: sl_archive_download(), sl_archive_index(), sl_assign_geography(), sl_boundaries(), sl_cache_clear(), sl_cache_dir(), sl_contract(), sl_exposure(), sl_list_versions(), sl_population(), sl_population_crosstab(), sl_read_records(), sl_sample(), sl_select_version()

Examples

sl_list_versions(system.file("extdata", "sample", package = "searchlight"))

Assign anonymised points to local geographies

Description

Uses British National Grid (EPSG:27700). Missing or invalid coordinates are retained with NA assignments. On shared polygon edges, ties are resolved by code and flagged ambiguous. Boundary distances describe anonymised snap points, not the unknown true locations; resulting error may be systematic. The boundary-sensitive share uses assigned events with measured distances. Missing-coordinate and unassigned shares use all supplied events.

Usage

sl_assign_geography(
  records,
  boundaries,
  types = names(boundaries),
  threshold = 50
)

Arguments

records

Contract-bearing records.

boundaries

A named list of sf layers, or one sf layer with metadata.

types

Names of the layers to assign; defaults to supplied layers.

threshold

Boundary-sensitive distance in metres.

Value

Records with geography codes, per-layer distances and flags, and an updated contract. No event rows are dropped or duplicated.

See Also

sl_boundaries(), sl_location_quality()

Other ingest: sl_archive_download(), sl_archive_index(), sl_archive_snapshot(), sl_boundaries(), sl_cache_clear(), sl_cache_dir(), sl_contract(), sl_exposure(), sl_list_versions(), sl_population(), sl_population_crosstab(), sl_read_records(), sl_sample(), sl_select_version()

Examples

b <- readRDS(system.file("extdata", "sample-boundaries.rds",
  package = "searchlight"
))
x <- sl_sample()[1:3, ]
b$msoa21 <- b$msoa21[b$msoa21$geography_code %in% x$msoa21, ]
sl_assign_geography(x, b, "msoa21")

Cross-check annual totals against the Home Office

Description

Years end on 31 March. Ratios are withheld unless twelve force-months are submitted or refreshed and the current records contain every selected row. The benchmark is not an equality guarantee: scope of powers, event versus person/vehicle reporting, publication cut-offs and revisions can differ.

Usage

sl_benchmark(records, ppap_table = NULL, tolerance = 0.1)

Arguments

records

Contract-bearing records.

ppap_table

Table with force_id, year_end and published_total. Default is the bundled Home Office table SS_20, year ending March 2025.

tolerance

Flag absolute proportional differences exceeding this value.

Value

A tibble with observed/published counts, coverage, ratio and flags.

See Also

sl_coverage()

Other audit: sl_coverage(), sl_coverage_compare(), sl_location_quality(), sl_quality(), sl_timestamp_quality()

Examples

ppap <- utils::read.csv(system.file("extdata", "ppap-totals.csv",
  package = "searchlight"
))
sl_benchmark(sl_sample(), ppap)

Fetch a versioned ONS boundary layer

Description

Uses a verified catalogue; latest means the latest verified vintage in the installed catalogue, not an unrecorded live change. Boundaries are BGC (20 m generalised, coast clipped). Region polygons cover England. Census hierarchy and current administrative lookups are attached with separate vintages.

Usage

sl_boundaries(
  type,
  vintage = "latest",
  dir = sl_cache_dir(),
  codes = NULL,
  lookups = TRUE
)

Arguments

type

One of lsoa21, msoa21, ward, lad, pfa, region.

vintage

A catalogue vintage (YYYY-MM), or latest.

dir

Explicit cache directory.

codes

Optional ONS codes to limit acquisition.

lookups

Fetch and attach the ONS lookup tables.

Value

An sf object with geography_code, geography_name, source metadata and lookups, or NULL with a message when the source is unavailable.

See Also

sl_assign_geography(), sl_population()

Other ingest: sl_archive_download(), sl_archive_index(), sl_archive_snapshot(), sl_assign_geography(), sl_cache_clear(), sl_cache_dir(), sl_contract(), sl_exposure(), sl_list_versions(), sl_population(), sl_population_crosstab(), sl_read_records(), sl_sample(), sl_select_version()

Examples

boundaries <- readRDS(system.file("extdata", "sample-boundaries.rds",
  package = "searchlight"
))
boundaries$lsoa21[1, "geography_code"]

Clear a marked searchlight cache

Description

Only removes searchlight-managed entries in a directory bearing the marker created by acquisition. Other user files are retained. This is destructive.

Usage

sl_cache_clear(dir = sl_cache_dir())

Arguments

dir

Cache directory.

Value

Invisibly, paths removed.

See Also

sl_cache_dir()

Other ingest: sl_archive_download(), sl_archive_index(), sl_archive_snapshot(), sl_assign_geography(), sl_boundaries(), sl_cache_dir(), sl_contract(), sl_exposure(), sl_list_versions(), sl_population(), sl_population_crosstab(), sl_read_records(), sl_sample(), sl_select_version()

Examples

d <- tempfile()
dir.create(d)
file.create(file.path(d, ".searchlight-cache"))
sl_cache_clear(d)
unlink(d, recursive = TRUE)

Locate the searchlight cache

Description

Returns the cache path without creating it. Acquisition functions explicitly create and write to this directory. Set options(searchlight.cache_dir = ...) to override the platform default, or pass an explicit dir.

Usage

sl_cache_dir()

Value

A character path.

See Also

sl_archive_index()

Other ingest: sl_archive_download(), sl_archive_index(), sl_archive_snapshot(), sl_assign_geography(), sl_boundaries(), sl_cache_clear(), sl_contract(), sl_exposure(), sl_list_versions(), sl_population(), sl_population_crosstab(), sl_read_records(), sl_sample(), sl_select_version()

Examples

sl_cache_dir()

Inspect an ingestion contract

Description

Contracts retain source provenance through subsetting and dplyr verbs. The coverage and timestamp tables describe the ingested sources; scope records the currently visible row count and force-months after filtering. Filtering does not turn a submitted file into a missing submission.

Usage

sl_contract(x)

Arguments

x

Records, aggregates or a contract.

Value

A named list of class sl_contract. summary() returns coverage.

See Also

sl_read_records(), sl_sample()

Other ingest: sl_archive_download(), sl_archive_index(), sl_archive_snapshot(), sl_assign_geography(), sl_boundaries(), sl_cache_clear(), sl_cache_dir(), sl_exposure(), sl_list_versions(), sl_population(), sl_population_crosstab(), sl_read_records(), sl_sample(), sl_select_version()

Examples

sl_contract(sl_sample())

Fit count models with population-time exposure

Description

Models recorded events, never unique people or a population-minus-stops cell. A population offset is multiplied by months_submitted/12 when available. Other offset columns must already represent exposure over the observed period. Unknown ethnicity and rows without positive exposure are reported as exclusions. Moran's I uses mean Pearson residuals per area and is exploratory after fitting. The exponentiated intercept is a baseline rate per exposure unit; other exponentiated coefficients are multiplicative rate effects. Confidence bounds are reported on this exponentiated scale.

Usage

sl_count_model(
  counts,
  formula,
  family = c("poisson", "negbin"),
  offset = "population",
  random = NULL,
  boundaries = NULL,
  conf_level = 0.95
)

Arguments

counts

Counts with a contract and exposure column, usually sl_rates.

formula

A formula with response n and desired fixed predictors.

family

Poisson or negative binomial.

offset

Name of the population or population-time exposure column.

random

Optional one-sided lme4 formula, for example ~ (1 | force_id).

boundaries

Optional sf polygons keyed by geography_code for residuals.

conf_level

Wald confidence level.

Value

A tidy sl_count_model coefficient table with model, diagnostics and excluded row/event counts in attributes, carrying the ingestion contract.

See Also

sl_rates(), sl_spatial_disparity()

Other inference: sl_hit_rates(), sl_simulate(), sl_spatial_disparity(), sl_spatial_map(), sl_veil_of_darkness()

Examples

r <- readRDS(system.file("extdata", "example-rates.rds",
  package = "searchlight"
))
r <- r[r$ethnicity %in% c("White", "Black"), ]
sl_count_model(r, n ~ ethnicity)
# Supply matching boundaries to request the residual Moran diagnostic.

Count events while preserving submission strata

Description

Force and month always remain in the result, even if omitted from by: these identify independent reporting and time-exposure cells. Unknown ethnicity remains a level. Missing submissions have NA counts; submitted combinations without events have zero. Filtering records does not narrow the source time grid: pass months and forces explicitly to select the analysis period.

Usage

sl_counts(
  records,
  by = c("pfa", "month", "ethnicity_5"),
  units = NULL,
  months = NULL,
  forces = NULL
)

Arguments

records

Contract-bearing records.

by

Additional grouping columns, including exactly one ethnicity field.

units

Optional force_id and geography mapping defining the full area universe, including areas without events. Otherwise uses contract units or observed areas. PFA defaults to the published force-code mapping.

months, forces

Explicit analysis scope; defaults to contract coverage.

Value

An sl_counts tibble with canonical geography_code, ethnicity, n, force_id and month columns, original grouping fields, and coverage status.

See Also

sl_rates(), sl_coverage()

Other rates: sl_rate_ratio(), sl_rates()

Examples

counts <- sl_counts(sl_sample())
head(counts)

Audit submission coverage on a full force-month grid

Description

Counts come from the immutable selected files, including after filtering records. A submitted zero-row CSV is zero; an absent CSV is NA. Refreshes describe differing version counts. Partial suspicion takes precedence over refreshes, and uses preceding calendar months (at most twelve), not twelve observed files. A changelog issue does not create a nonexistent submission.

Usage

sl_coverage(
  records,
  forces = NULL,
  months = NULL,
  changelog = NULL,
  partial_threshold = 0.2
)

Arguments

records

Contract-bearing records.

forces

Force IDs; default is the contract's requested forces.

months

Months; default is the contract's requested months.

changelog

Table with force_id, month, note and optional issue flag. Default uses the bundled, dated changelog extract.

partial_threshold

Fraction of trailing median, default 0.2.

Value

An sl_coverage tibble with contract; plot gives a coverage heatmap.

See Also

sl_quality(), sl_coverage_compare(), sl_benchmark()

Other audit: sl_benchmark(), sl_coverage_compare(), sl_location_quality(), sl_quality(), sl_timestamp_quality()

Examples

audit <- sl_coverage(sl_sample())
audit
# plot(audit) displays the submission heatmap.

Compare coverage patterns across two periods

Description

Periods are explicit month vectors sorted chronologically. Comparability means equal length, matching submitted positions, at least one submitted position, and no partial suspicion. Identical missing positions are listed but allow comparison of the same observed subset. This tests coverage only, not seasonality, reporting practice or population comparability.

Usage

sl_coverage_compare(coverage, period_a, period_b)

Arguments

coverage

Output of sl_coverage.

period_a, period_b

Month vectors for comparison.

Value

A tibble with comparable, submitted positions, and a list of gaps.

See Also

sl_coverage()

Other audit: sl_benchmark(), sl_coverage(), sl_location_quality(), sl_quality(), sl_timestamp_quality()

Examples

sl_coverage_compare(sl_coverage(sl_sample()), "2026-05", "2026-06")

Compare compatible population denominator scenarios

Description

Exposure tables must have identical geography/classification and cell keys. Each sampling interval is conditional on its exposure scenario; the range of point estimates across scenarios is reported separately.

Usage

sl_denominator_scenarios(
  counts,
  exposures,
  reference = "White",
  comparison = "Black"
)

Arguments

counts

Contract-bearing event counts.

exposures

Named list of sl_exposure tables.

reference, comparison

Groups to compare.

Value

An sl_sensitivity tibble with scenario ratios and assumption ranges.

See Also

sl_exposure(), sl_missing_ethnicity_bounds()

Other sensitivity: sl_missing_ethnicity_bounds(), sl_ranking_stability(), sl_standardise()

Examples

p <- readRDS(system.file("extdata", "sample-population.rds",
  package = "searchlight"
))$msoa21
c <- readRDS(system.file("extdata", "example-counts.rds",
  package = "searchlight"
))
head(sl_denominator_scenarios(c, list(resident = p)))

Construct an explicit exposure table

Description

Population is an exposure. Unknown ethnicity has no Census population denominator and must have NA population. Additional age_band and sex columns identify disjoint strata. Counts may exceed exposure.

Usage

sl_exposure(
  data,
  geography,
  classification = c("5", "19"),
  source = "user_supplied"
)

Arguments

data

A data frame with geography_code, ethnicity and population.

geography

Geography type and vintage, for example msoa21.

classification

Self-defined classification, 5 or 19.

source

Description of a user-supplied exposure.

Value

A tibble with class sl_exposure and an exposure contract.

See Also

sl_population(), sl_population_crosstab()

Other ingest: sl_archive_download(), sl_archive_index(), sl_archive_snapshot(), sl_assign_geography(), sl_boundaries(), sl_cache_clear(), sl_cache_dir(), sl_contract(), sl_list_versions(), sl_population(), sl_population_crosstab(), sl_read_records(), sl_sample(), sl_select_version()

Examples

sl_exposure(data.frame(
  geography_code = "area1", ethnicity = "White",
  population = 1000
), geography = "example")

Describe three separate outcomes conditional on being searched

Description

Computes each requested outcome separately with Wilson binomial intervals. Unknown ethnicity remains in descriptive tables and missing outcomes are not failures. Logistic comparisons use known ethnicity, force and object fixed effects; constant controls are explicitly omitted. These are associations among searches. Differing hit rates alone do not establish discrimination; selection, differing risk distributions and infra-marginality matter.

Usage

sl_hit_rates(
  records,
  by = c("pfa", "object_group"),
  outcome = c("any_action", "arrest", "outcome_linked_to_object"),
  reference = "White",
  comparison = "Black",
  conf_level = 0.95
)

Arguments

records

Search records with the three derived binary outcome measures.

by

Descriptive grouping columns; ethnicity_5 is always retained.

outcome

Separate outcomes to calculate (all three by default).

reference, comparison

Known self-defined ethnicity groups.

conf_level

Wilson and logistic Wald confidence level.

Value

An sl_hit_rates tibble. Models, coefficients and exclusions describe separate logistic models. The ingestion contract is retained.

See Also

sl_read_records(), sl_veil_of_darkness()

Other inference: sl_count_model(), sl_simulate(), sl_spatial_disparity(), sl_spatial_map(), sl_veil_of_darkness()

Examples

hits <- sl_hit_rates(sl_sample())
head(hits)
attr(hits, "coefficients")

List force-month archive versions

Description

List force-month archive versions

Usage

sl_list_versions(dir = sl_cache_dir())

Arguments

dir

A snapshot cache or bundled sample directory.

Value

A tibble with each force-month's archive, CSV hash and row count.

See Also

sl_archive_snapshot(), sl_select_version()

Other ingest: sl_archive_download(), sl_archive_index(), sl_archive_snapshot(), sl_assign_geography(), sl_boundaries(), sl_cache_clear(), sl_cache_dir(), sl_contract(), sl_exposure(), sl_population(), sl_population_crosstab(), sl_read_records(), sl_sample(), sl_select_version()

Examples

sl_list_versions(system.file("extdata", "sample", package = "searchlight"))

Diagnose anonymised location quality

Description

Snap-point concentration is the share at the most common reported coordinate within an LSOA, among records with usable coordinates in that LSOA. It is not the concentration at true incident locations. Missing LSOA assignments remain a separate group. Per-geography boundary diagnostics remain in the contract.

Usage

sl_location_quality(records)

Arguments

records

Contract-bearing records, optionally with assigned LSOAs.

Value

A force-month-LSOA tibble with missingness, concentration and boundary sensitivity, carrying the contract.

See Also

sl_assign_geography(), sl_timestamp_quality()

Other audit: sl_benchmark(), sl_coverage(), sl_coverage_compare(), sl_quality(), sl_timestamp_quality()

Examples

sl_location_quality(sl_sample()[1:100, ])

Bound disparity under allocations of Unknown ethnicity

Description

Extreme allocations send every Unknown event to one of the two groups. Proportional allocation uses all known groups, not only the comparison pair. force_object_mar uses known-group shares within force and object, pooled over areas and months; it refuses strata with no known ethnicity. These are assumptions, not confidence intervals. Force/object allocation also assumes pooled known composition applies to each recipient area/month; MAR alone does not establish that transportability. The tipping share solves the equation where the ratio equals one, with all remaining Unknown sent to the reference. A tipping value outside zero to one means no feasible tipping allocation.

Usage

sl_missing_ethnicity_bounds(
  counts,
  population,
  reference = "White",
  comparison = "Black",
  scenarios = c("all_to_reference", "all_to_comparison", "proportional",
    "force_object_mar")
)

Arguments

counts

Contract-bearing event counts.

population

Marginal ethnic-group exposure table.

reference, comparison

Groups to compare.

scenarios

Allocation assumptions to report.

Value

An sl_sensitivity tibble with scenario ratios, extreme bounds, tipping allocation and separately labelled baseline sampling intervals.

See Also

sl_rates(), sl_denominator_scenarios()

Other sensitivity: sl_denominator_scenarios(), sl_ranking_stability(), sl_standardise()

Examples

c <- readRDS(system.file("extdata", "example-counts.rds",
  package = "searchlight"
))
p <- readRDS(system.file("extdata", "sample-population.rds",
  package = "searchlight"
))$msoa21
head(sl_missing_ethnicity_bounds(c, p))

Census 2021 ethnicity populations

Description

Fetches TS021 (NM_2041_1), with exact codes verified against returned cells. Supply ONS codes or an sf boundary object. Current ward and LAD boundaries must not be silently matched to Census 2021 units: unchanged codes alone do not demonstrate unchanged boundaries. A new exposure source must be supplied for changed units. Census disclosure control may cause small inconsistencies between separately published tables. Unknown is retained with NA exposure.

Usage

sl_population(geography, classification = c("5", "19"), dir = sl_cache_dir())

Arguments

geography

ONS geography codes, or an sf layer with geography_code.

classification

Census classification, 5 or 19.

dir

Explicit cache directory, written only on this call.

Value

An sl_exposure tibble, or NULL with a message if unavailable.

See Also

sl_exposure(), sl_population_crosstab()

Other ingest: sl_archive_download(), sl_archive_index(), sl_archive_snapshot(), sl_assign_geography(), sl_boundaries(), sl_cache_clear(), sl_cache_dir(), sl_contract(), sl_exposure(), sl_list_versions(), sl_population_crosstab(), sl_read_records(), sl_sample(), sl_select_version()

Examples

p <- readRDS(system.file("extdata", "sample-population.rds",
  package = "searchlight"
))
head(p$msoa21)

Census ethnicity by age and sex

Description

Uses RM032 (NM_2132_1), available at Census 2021 small areas. Five published age bands are combined to under 25, 25-34 and 35+, the common partition with police age bands. Census sex and police-recorded gender are different measurements; any analysis using them must state this assumption. Unsupported geography returns NULL and a message. No age-sex cross-tabulation is imputed from marginal totals.

Usage

sl_population_crosstab(
  geography,
  dir = sl_cache_dir(),
  classification = c("5", "19")
)

Arguments

geography

ONS geography codes, or an sf layer with geography_code.

dir

Explicit cache directory, written only on this call.

classification

Census classification, 5 or 19.

Value

An sl_exposure tibble with age_band and sex, or NULL.

See Also

sl_population(), sl_exposure()

Other ingest: sl_archive_download(), sl_archive_index(), sl_archive_snapshot(), sl_assign_geography(), sl_boundaries(), sl_cache_clear(), sl_cache_dir(), sl_contract(), sl_exposure(), sl_list_versions(), sl_population(), sl_read_records(), sl_sample(), sl_select_version()

Examples

p <- readRDS(system.file("extdata", "sample-crosstab.rds",
  package = "searchlight"
))
head(p)

Attach submission, timestamp and location diagnostics

Description

Pass population to complete the population metadata in the records contract; population counts remain a separate exposure table. Diagnostics describe the current records; coverage continues to describe the original source files.

Usage

sl_quality(records, population = NULL, threshold = 0.05, changelog = NULL)

Arguments

records

Contract-bearing records.

population

Optional sl_exposure table.

threshold

Midnight-share threshold.

changelog

Optional changelog table passed to sl_coverage.

Value

Records with an updated contract and attached quality tables.

See Also

sl_coverage(), sl_timestamp_quality(), sl_location_quality()

Other audit: sl_benchmark(), sl_coverage(), sl_coverage_compare(), sl_location_quality(), sl_timestamp_quality()

Examples

records <- sl_quality(sl_sample()[1:100, ])
sl_contract(records)$timestamps

Assess area rank uncertainty under count models or posterior draws

Description

Bootstrap draws event counts from the fitted Poisson or negative-binomial model, then ranks event-rate ratios (largest is rank 1). Quasi-Poisson does not define a count distribution and is refused. Positive totals in both groups are required for plug-in bootstrap; sparse zero-count areas need a suitable posterior model. Ties receive random ranks under the saved seed. Optional scenario rows are analysed separately, never pooled into a sampling distribution. Pairwise stability means ordering probability above 0.95.

Usage

sl_ranking_stability(
  rate_ratios,
  method = c("bootstrap", "posterior"),
  n = 1000,
  seed = 1
)

Arguments

rate_ratios

An sl_rate_ratio table, or a contract-bearing table with geography_code and a posterior_draws matrix attribute (draws by area).

method

Parametric bootstrap or posterior ranks.

n

Number of draws; default 1000.

seed

Reproducible seed, restored on exit.

Value

An sl_sensitivity tibble of median/95% rank intervals, with attributes rank_probabilities (matrix) and pairwise (ordering probabilities).

See Also

sl_rate_ratio(), sl_denominator_scenarios()

Other sensitivity: sl_denominator_scenarios(), sl_missing_ethnicity_bounds(), sl_standardise()

Examples

# Posterior draws can also be supplied by an independently fitted model.
x <- tibble::tibble(geography_code = c("a", "b"))
attr(x, "contract") <- sl_contract(sl_sample())
attr(x, "posterior_draws") <- cbind(a = c(1, 2, 3), b = c(3, 2, 1))
sl_ranking_stability(x, "posterior", n = 3)

Estimate an event stop-rate ratio with exposure offsets

Description

Fits n ~ ethnicity with log(population-time) offset within each area and other grouping stratum, pooling submitted months. Poisson intervals are exact conditional intervals for two event counts; zero counts are supported without adding pseudocounts. Quasi-Poisson and negative binomial require replicated cells and use model-based log intervals. Dispersion is Pearson chi-square divided by residual degrees of freedom. This is not a person-level risk ratio.

Usage

sl_rate_ratio(
  rates,
  reference = "White",
  comparison = "Black",
  method = c("poisson", "quasipoisson", "negbin"),
  conf_level = 0.95
)

Arguments

rates

Output from sl_rates.

reference, comparison

Self-defined ethnicity groups.

method

Count model family.

conf_level

Sampling confidence level.

Value

An sl_rate_ratio tibble with sampling intervals, dispersion, counts, exposures and fitted model list column, carrying the ingestion contract.

See Also

sl_rates(), sl_missing_ethnicity_bounds()

Other rates: sl_counts(), sl_rates()

Examples

r <- readRDS(system.file("extdata", "example-rates.rds",
  package = "searchlight"
))
head(sl_rate_ratio(r))

Calculate rates using observed population-time exposure

Description

The primary rate is annualised events per 1,000 resident person-years: n / (population * submitted_months / 12) * per. The separate period_rate is n / population * per for that cell. Neither is a probability of a person being searched. NA submission counts are excluded from time exposure. Partial submissions remain flagged and describe reported events only.

Usage

sl_rates(counts, population, per = 1000)

Arguments

counts

Output from sl_counts.

population

A compatible sl_exposure table.

per

Rate scaling, default 1000.

Value

An sl_rates tibble with population, exposure, rate and period_rate.

See Also

sl_counts(), sl_rate_ratio(), sl_exposure()

Other rates: sl_counts(), sl_rate_ratio()

Examples

c <- readRDS(system.file("extdata", "example-counts.rds",
  package = "searchlight"
))
p <- readRDS(system.file("extdata", "sample-population.rds",
  package = "searchlight"
))$msoa21
head(sl_rates(c, p))

Read selected immutable archive records

Description

Validates the fixed CSV schema, checksum and row count, preserving every event (including identical rows) and raw ethnicity strings. Timestamps retain their supplied UTC offsets and are displayed in Europe/London. Absent or refused ethnicity is Unknown. Missing outcome measures remain NA, not negative outcomes. The filename defines the submission month. Offset conversion can cross a month boundary; date_month_crossing records this without moving the event into another source file. A timestamp must match the filename in either its supplied clock or the London clock.

Usage

sl_read_records(dir = sl_cache_dir(), versions = NULL)

Arguments

dir

Snapshot cache directory.

versions

One selected version per force-month; defaults to latest.

Value

An sl_records tibble with a validated ingestion contract.

See Also

sl_contract(), sl_select_version(), sl_sample()

Other ingest: sl_archive_download(), sl_archive_index(), sl_archive_snapshot(), sl_assign_geography(), sl_boundaries(), sl_cache_clear(), sl_cache_dir(), sl_contract(), sl_exposure(), sl_list_versions(), sl_population(), sl_population_crosstab(), sl_sample(), sl_select_version()

Examples

d <- system.file("extdata", "sample", package = "searchlight")
v <- sl_select_version(sl_list_versions(d))
records <- sl_read_records(d, v[1, ])
summary(sl_contract(records))

Write an offline report of audited records and analysis results

Description

Renders the installed Markdown/Rmd template to a self-contained HTML file. No Pandoc, network, model fitting or external web assets are required. Supplied estimates must retain the same source snapshots and ethnicity classification as records. Their analysis scope and diagnostics are shown. Coverage describes source files; filtering does not redefine coverage. Sampling intervals and assumption ranges have separate labelled sections. The report includes a summary, section navigation and expandable audit details. Tables retain their original column names in header tooltips.

Usage

sl_report(
  records,
  file,
  estimates = list(),
  title = "Stop and search: evidence and assumptions",
  assumptions = character(),
  limitations = character(),
  max_rows = 500,
  overwrite = FALSE
)

Arguments

records

Contract-bearing search records.

file

Explicit HTML output path. Its parent directory must exist.

estimates

Uniquely named list of contract-bearing result tables.

title

Report title, treated as text.

assumptions, limitations

Additional statements, treated as text.

max_rows

Maximum displayed rows per table; truncation is labelled.

overwrite

Whether to replace an existing output file.

Value

Invisibly, a one-row tibble with the output path and record count, carrying the ingestion contract. Writes only the requested output file.

See Also

sl_coverage(), sl_rate_ratio(), sl_missing_ethnicity_bounds()

Examples

rates <- readRDS(system.file("extdata", "example-rates.rds",
  package = "searchlight"
))
path <- tempfile(fileext = ".html")
sl_report(sl_sample(), path, estimates = list(ratios = sl_rate_ratio(rates)))
file.exists(path)

Load the bundled real archive sample

Description

Load the bundled real archive sample

Usage

sl_sample()

Value

An sl_records tibble for the documented sample forces and months.

See Also

sl_read_records(), sl_contract()

Other ingest: sl_archive_download(), sl_archive_index(), sl_archive_snapshot(), sl_assign_geography(), sl_boundaries(), sl_cache_clear(), sl_cache_dir(), sl_contract(), sl_exposure(), sl_list_versions(), sl_population(), sl_population_crosstab(), sl_read_records(), sl_select_version()

Examples

x <- sl_sample()
table(x$force_id, x$month)

Select one complete version per force-month

Description

Never deduplicates event rows. Ties are resolved by snapshot month and archive name. max_rows chooses the newest version among equal row counts. Differences greater than 1% of the smallest count are reported (zero to positive is infinite).

Usage

sl_select_version(
  versions,
  rule = c("latest", "earliest", "max_rows", "manual"),
  manual = NULL
)

Arguments

versions

Output from sl_list_versions().

rule

Version selection rule.

manual

For the manual rule, a tibble with force_id, month, archive_file, containing exactly one choice per force-month.

Value

Selected versions with list-column alternatives and a differences attribute. Each row carries the selection rule.

See Also

sl_read_records()

Other ingest: sl_archive_download(), sl_archive_index(), sl_archive_snapshot(), sl_assign_geography(), sl_boundaries(), sl_cache_clear(), sl_cache_dir(), sl_contract(), sl_exposure(), sl_list_versions(), sl_population(), sl_population_crosstab(), sl_read_records(), sl_sample()

Examples

d <- system.file("extdata", "sample", package = "searchlight")
v <- sl_list_versions(d)
sl_select_version(v)

Simulate event counts with independent disparity and total intensity

Description

The disparity surface varies east-west and expected total events north-south. Their lattice correlation is zero by construction. Counts are Poisson events, then unknown ethnicity is generated by binomial thinning. MCAR has a common probability, MAR varies by synthetic force, and MNAR varies by ethnicity. True latent rate ratios and complete realised event ratios are distinct.

Usage

sl_simulate(
  side = 6,
  surface = c("smooth", "discontinuous", "small"),
  population = c(White = 5000, Black = 1000),
  rate = 0.025,
  missingness = c("none", "mcar", "mar", "mnar"),
  missing_rate = 0.2,
  months = 12,
  seed = 1
)

Arguments

side

Lattice width and height, at least three.

surface

Smooth, discontinuous, or small comparison-population setting.

population

Resident reference/comparison populations per area.

rate

Mean total annual event rate per resident.

missingness

Missing-ethnicity mechanism.

missing_rate

Baseline missing probability.

months

Number of submitted months, between one and twelve.

seed

Reproducible seed, restored on exit.

Value

A list of counts, complete_counts, population, boundaries, truth and settings. Counts carry a clearly synthetic ingestion contract.

See Also

sl_spatial_disparity(), sl_missing_ethnicity_bounds()

Other inference: sl_count_model(), sl_hit_rates(), sl_spatial_disparity(), sl_spatial_map(), sl_veil_of_darkness()

Examples

s <- sl_simulate(side = 3, missingness = "mnar", missing_rate = 0.3)
s$truth
sl_missing_ethnicity_bounds(s$counts, s$population,
  scenarios = c("all_to_reference", "all_to_comparison")
)

Estimate a multivariate spatial stop-rate disparity surface

Description

Uses CARBayes MVS.CARleroux with two Poisson responses and a matrix of log population-time offsets. Each ethnicity has its own spatial effect phi_ig. The equivalent shared/contrast decomposition is u_i=(phi_ir+phi_ic)/2, v_ir=(phi_ir-phi_ic)/2 and v_ic=-v_ir. These components have the covariance induced by the multivariate prior; they are not separately independent priors. The full posterior ratio is exp(alpha_c-alpha_r+phi_ic-phi_ir). Smoothing stabilises under stated assumptions; it cannot repair exposure, missing ethnicity, missing submissions or anonymised location errors.

Usage

sl_spatial_disparity(
  counts,
  population,
  boundaries,
  reference = "White",
  comparison = "Black",
  backend = c("carbayes", "inla"),
  burnin = 2000,
  n.sample = 12000,
  thin = 10,
  chains = 2,
  k = 2,
  adjacency = NULL,
  boundary_threshold = 0.2,
  seed = 1,
  ...
)

Arguments

counts

Marginal event counts with a valid contract.

population

Compatible marginal ethnicity exposure table.

boundaries

sf polygons with geography_code.

reference, comparison

Known self-defined ethnicity groups.

backend

CARBayes. INLA is explicitly deferred beyond version 0.1.0.

burnin, n.sample, thin

MCMC controls per chain, including burn-in in n.sample.

chains

Number of independent chains, at least two, run on one core.

k

Additional ratio threshold for exceedance probability.

adjacency

Optional symmetric adjacency matrix named with geography codes. Otherwise rook adjacency is used; islands require an explicit decision.

boundary_threshold

Warn if the assigned boundary-sensitive share exceeds this proportion (default 0.2).

seed

Reproducible seed, restored on exit.

...

Prior and sampler arguments for MVS.CARleroux, such as rho or MALA.

Value

An sl_spatial tibble with 90/95 percent intervals, exceedance probabilities, crude ratios, rank-normalised R-hat and bulk/tail ESS. Joint posterior draws, fitted model, diagnostics and boundaries are attributes.

See Also

sl_simulate(), sl_spatial_map(), sl_ranking_stability()

Other inference: sl_count_model(), sl_hit_rates(), sl_simulate(), sl_spatial_map(), sl_veil_of_darkness()

Examples

# A precomputed synthetic fit avoids MCMC in installed examples.
path <- system.file("extdata", "example-spatial.rds",
  package = "searchlight"
)
if (nzchar(path)) {
  fit <- readRDS(path)
  head(fit)
  head(sl_spatial_map(fit))
}

Join spatial posterior summaries to their polygons

Description

Join spatial posterior summaries to their polygons

Usage

sl_spatial_map(x)

Arguments

x

An sl_spatial fit.

Value

An sf object with summaries and its ingestion contract.

See Also

sl_spatial_disparity()

Other inference: sl_count_model(), sl_hit_rates(), sl_simulate(), sl_spatial_disparity(), sl_veil_of_darkness()

Examples

path <- system.file("extdata", "example-spatial.rds",
  package = "searchlight"
)
if (nzchar(path)) head(sl_spatial_map(readRDS(path)))

Directly standardise age-sex stop-event rates

Description

Uses common age-sex weights across ethnicity groups. Census sex is compared with police-recorded gender only under an explicit measurement assumption. Unknown age, gender and ethnicity are retained in exclusion diagnostics. A positively weighted stratum with zero exposure makes the standardised rate undefined; it is never silently dropped. Sampling intervals combine exact Poisson stratum intervals with a Bonferroni correction and fixed weights. They are conservative, including when every observed stratum count is zero.

Usage

sl_standardise(
  counts,
  crosstab,
  standard = c("england_wales", "study_population"),
  per = 1000
)

Arguments

counts

Counts grouped by age_band and sex.

crosstab

Compatible Census or explicit ethnicity-age-sex exposure.

standard

England/Wales Census 2021 or pooled study population weights.

per

Rate scaling, default 1000 per person-year.

Value

An sl_sensitivity tibble with crude and standardised rates, sampling intervals, excluded-event counts, weights and the ingestion contract.

See Also

sl_population_crosstab(), sl_rates()

Other sensitivity: sl_denominator_scenarios(), sl_missing_ethnicity_bounds(), sl_ranking_stability()

Examples

p <- readRDS(system.file("extdata", "sample-crosstab.rds",
  package = "searchlight"
))
c <- readRDS(system.file("extdata", "example-demographic-counts.rds",
  package = "searchlight"
))
head(suppressWarnings(sl_standardise(c, p)))

Diagnose timestamp completeness and midnight heaping

Description

Reports both London-clock midnight and midnight in the source string. Some date-only summer records carry a UTC offset, so checking London midnight alone would miss them. Reliability requires both shares below threshold and no missing timestamps. This is a screening diagnostic, not proof that all supplied times are accurate. Empty and absent force-months are not reliable.

Usage

sl_timestamp_quality(records, threshold = 0.05)

Arguments

records

Contract-bearing records.

threshold

Maximum permitted midnight share, default 0.05.

Value

A tibble with force-month diagnostics and within-force share range.

See Also

sl_quality(), sl_location_quality()

Other audit: sl_benchmark(), sl_coverage(), sl_coverage_compare(), sl_location_quality(), sl_quality()

Examples

sl_timestamp_quality(sl_sample())

Diagnose ethnicity composition in daylight and darkness

Description

Excludes every force-month that fails or lacks the contract's time-quality screen, retaining an explicit exclusion table. Annual evening clock-time bounds use each distinct location and year. Civil darkness begins at dusk; ambiguous sunset-to-dusk observations are excluded. Times and clock changes use Europe/London, including British summer time. The logistic comparison controls clock time (natural spline), weekday, month and force where these vary. The DST option further restricts to weeks around both clock changes and adds transition and running-day controls; it is a local darkness comparison. The method was developed for vehicle stops. Transfer to pedestrian searches requires assumptions about visibility, activity, deployment and selection.

Usage

sl_veil_of_darkness(
  records,
  boundaries = NULL,
  twilight = c("civil", "sunset"),
  window = "intertwilight",
  design = c("intertwilight", "dst"),
  reference = "White",
  comparison = "Black",
  dst_weeks = 3,
  conf_level = 0.95
)

Arguments

records

Contract-bearing search records.

boundaries

Optional sf study region; restrict published points to it. Missing locations are never replaced by centroids.

twilight

Civil twilight end or sunset definition of darkness.

window

Currently only intertwilight, the annual evening overlap window.

design

Annual intertwilight or local DST-window comparison.

reference, comparison

Known self-defined ethnicity groups.

dst_weeks

Weeks on either side of each clock change for the DST design.

conf_level

Logistic Wald confidence level.

Value

An sl_veil tibble with an odds ratio, interval and fit status. Attributes steps, excluded_force_months, assumptions, data and model retain the full design audit. A gated or unidentifiable sample has an undefined ratio.

See Also

sl_timestamp_quality(), sl_hit_rates()

Other inference: sl_count_model(), sl_hit_rates(), sl_simulate(), sl_spatial_disparity(), sl_spatial_map()

Examples

# Date-only records are explicitly refused by the time-quality screen.
x <- sl_sample()[1:10, ]
day <- as.Date(x$date, tz = "Europe/London")
x$date <- as.POSIXct(paste(day, "00:00:00"), tz = "Europe/London")
x$date_raw <- format(x$date, "%Y-%m-%dT%H:%M:%S%z")
attr(x, "contract")$timestamps <- sl_timestamp_quality(x)
sl_veil_of_darkness(x)