| Title: | Audited Analysis of Police Stop and Search Records |
| Version: | 0.1.0 |
| Description: | Acquire and audit public police stop and search archives for England and Wales, retaining provenance and explicit coverage information. Designed for exposure-based ethnic disparity analysis with separately reported sampling and assumption uncertainty. Contains public sector information licensed under the Open Government Licence v3.0. |
| License: | MIT + file LICENSE |
| Encoding: | UTF-8 |
| Language: | en-GB |
| Depends: | R (≥ 4.1.0) |
| Imports: | CARBayes (≥ 6.1.1), cli, commonmark, digest, dplyr (≥ 1.1.1), httr2, jsonlite, MASS, readr, rlang, sf, spdep, splines, suncalc, posterior, ggplot2, tidyr, tibble, withr, xml2 |
| URL: | https://github.com/BlackThrive/searchlight, https://blackthrive.github.io/searchlight/ |
| BugReports: | https://github.com/BlackThrive/searchlight/issues |
| RoxygenNote: | 7.3.3 |
| Suggests: | testthat (≥ 3.0.0), httptest2, knitr, rmarkdown, lme4, lintr, spelling, vdiffr, zip |
| VignetteBuilder: | knitr |
| Config/testthat/edition: | 3 |
| NeedsCompilation: | no |
| Packaged: | 2026-09-22 04:33:08 UTC; musta |
| Author: | Mustapha Wasseja [aut, cre], Sarah Hamed [aut], Souci Frissa [aut], Black Thrive Global [cph, fnd] |
| Maintainer: | Mustapha Wasseja <mustapha.wasseja.mohammed@gmail.com> |
| Repository: | CRAN |
| Date/Publication: | 2026-09-30 11:40:02 UTC |
searchlight: audited police stop and search analysis
Description
Public archive records describe events, not unique people.
Missing submissions must not be treated as zero events.
Self-defined and officer-defined ethnicity
are distinct measurements. See the installed NOTES directory for
decisions
and verified sources.
Author(s)
Maintainer: Mustapha Wasseja mustapha.wasseja.mohammed@gmail.com
Authors:
Sarah Hamed Sarah.Hamed@blackthrive.org
Souci Frissa Souci.Frissa@blackthrive.org
Other contributors:
Black Thrive Global [copyright holder, funder]
See Also
Useful links:
Report bugs at https://github.com/BlackThrive/searchlight/issues
Download and verify the necessary archive snapshots
Description
Selects the newest available snapshots that cover the requested months.
Forces are recorded as extraction filters; the source only offers full
ZIPs.
Checks publisher MD5 before recording local SHA-256. Existing files
must match
both the publisher and any recorded SHA-256; corrupt files are never
replaced
silently. The manifest and ZIPs are written to dir.
Usage
sl_archive_download(months, forces = NULL, dir = sl_cache_dir(), index = NULL)
Arguments
months |
Character vector of YYYY-MM months. |
forces |
Force identifiers, or |
dir |
Cache directory. |
index |
Optional previously retrieved archive index. |
Value
Manifest tibble, or NULL when a resource is unavailable.
See Also
sl_archive_snapshot(), sl_list_versions()
Other ingest:
sl_archive_index(),
sl_archive_snapshot(),
sl_assign_geography(),
sl_boundaries(),
sl_cache_clear(),
sl_cache_dir(),
sl_contract(),
sl_exposure(),
sl_list_versions(),
sl_population(),
sl_population_crosstab(),
sl_read_records(),
sl_sample(),
sl_select_version()
Examples
index <- readRDS(system.file("extdata", "archive-index.rds",
package = "searchlight"
))
index$archive_file[1]
List available bulk archive snapshots
Description
Parses advertised month ranges and publisher MD5 checksums. The complete
rolling snapshot can contain several years. Results are cached only after
this function is called. Source failures return NULL with a message.
Usage
sl_archive_index(
dir = sl_cache_dir(),
refresh = FALSE,
url = "https://data.police.uk/data/archive/"
)
Arguments
dir |
Cache directory. |
refresh |
Refresh the cached index. |
url |
Archive listing URL, normally the official source. |
Value
A tibble with archive name, snapshot month, coverage range,
URL and MD5,
or NULL if unavailable.
See Also
Other ingest:
sl_archive_download(),
sl_archive_snapshot(),
sl_assign_geography(),
sl_boundaries(),
sl_cache_clear(),
sl_cache_dir(),
sl_contract(),
sl_exposure(),
sl_list_versions(),
sl_population(),
sl_population_crosstab(),
sl_read_records(),
sl_sample(),
sl_select_version()
Examples
index <- readRDS(system.file("extdata", "archive-index.rds",
package = "searchlight"
))
index[, c("archive_file", "month_start", "month_end")]
Extract immutable stop and search CSV snapshots
Description
Other archive contents are not extracted. Retains source ZIPs and raw CSV bytes. A manifest records archive and CSV SHA-256, row count and original member path. ZIP traversal paths are refused. Previously extracted CSVs are hash-checked.
Usage
sl_archive_snapshot(dir = sl_cache_dir(), months = NULL, forces = NULL)
Arguments
dir |
Cache containing archive ZIPs and optionally a download manifest. |
months |
Optional YYYY-MM filter; defaults to the download request. |
forces |
Optional force filter; defaults to the download request. |
Value
A tibble of CSV versions. Writes snapshots and their manifest
in dir.
See Also
sl_select_version(), sl_read_records()
Other ingest:
sl_archive_download(),
sl_archive_index(),
sl_assign_geography(),
sl_boundaries(),
sl_cache_clear(),
sl_cache_dir(),
sl_contract(),
sl_exposure(),
sl_list_versions(),
sl_population(),
sl_population_crosstab(),
sl_read_records(),
sl_sample(),
sl_select_version()
Examples
sl_list_versions(system.file("extdata", "sample", package = "searchlight"))
Assign anonymised points to local geographies
Description
Uses British National Grid (EPSG:27700). Missing or invalid coordinates are retained with NA assignments. On shared polygon edges, ties are resolved by code and flagged ambiguous. Boundary distances describe anonymised snap points, not the unknown true locations; resulting error may be systematic. The boundary-sensitive share uses assigned events with measured distances. Missing-coordinate and unassigned shares use all supplied events.
Usage
sl_assign_geography(
records,
boundaries,
types = names(boundaries),
threshold = 50
)
Arguments
records |
Contract-bearing records. |
boundaries |
A named list of sf layers, or one sf layer with metadata. |
types |
Names of the layers to assign; defaults to supplied layers. |
threshold |
Boundary-sensitive distance in metres. |
Value
Records with geography codes, per-layer distances and flags, and an updated contract. No event rows are dropped or duplicated.
See Also
sl_boundaries(), sl_location_quality()
Other ingest:
sl_archive_download(),
sl_archive_index(),
sl_archive_snapshot(),
sl_boundaries(),
sl_cache_clear(),
sl_cache_dir(),
sl_contract(),
sl_exposure(),
sl_list_versions(),
sl_population(),
sl_population_crosstab(),
sl_read_records(),
sl_sample(),
sl_select_version()
Examples
b <- readRDS(system.file("extdata", "sample-boundaries.rds",
package = "searchlight"
))
x <- sl_sample()[1:3, ]
b$msoa21 <- b$msoa21[b$msoa21$geography_code %in% x$msoa21, ]
sl_assign_geography(x, b, "msoa21")
Cross-check annual totals against the Home Office
Description
Years end on 31 March. Ratios are withheld unless twelve force-months are submitted or refreshed and the current records contain every selected row. The benchmark is not an equality guarantee: scope of powers, event versus person/vehicle reporting, publication cut-offs and revisions can differ.
Usage
sl_benchmark(records, ppap_table = NULL, tolerance = 0.1)
Arguments
records |
Contract-bearing records. |
ppap_table |
Table with force_id, year_end and published_total. Default is the bundled Home Office table SS_20, year ending March 2025. |
tolerance |
Flag absolute proportional differences exceeding this value. |
Value
A tibble with observed/published counts, coverage, ratio and flags.
See Also
Other audit:
sl_coverage(),
sl_coverage_compare(),
sl_location_quality(),
sl_quality(),
sl_timestamp_quality()
Examples
ppap <- utils::read.csv(system.file("extdata", "ppap-totals.csv",
package = "searchlight"
))
sl_benchmark(sl_sample(), ppap)
Fetch a versioned ONS boundary layer
Description
Uses a verified catalogue; latest means the latest verified vintage in the installed catalogue, not an unrecorded live change. Boundaries are BGC (20 m generalised, coast clipped). Region polygons cover England. Census hierarchy and current administrative lookups are attached with separate vintages.
Usage
sl_boundaries(
type,
vintage = "latest",
dir = sl_cache_dir(),
codes = NULL,
lookups = TRUE
)
Arguments
type |
One of lsoa21, msoa21, ward, lad, pfa, region. |
vintage |
A catalogue vintage (YYYY-MM), or latest. |
dir |
Explicit cache directory. |
codes |
Optional ONS codes to limit acquisition. |
lookups |
Fetch and attach the ONS lookup tables. |
Value
An sf object with geography_code, geography_name, source metadata
and lookups, or NULL with a message when the source is unavailable.
See Also
sl_assign_geography(), sl_population()
Other ingest:
sl_archive_download(),
sl_archive_index(),
sl_archive_snapshot(),
sl_assign_geography(),
sl_cache_clear(),
sl_cache_dir(),
sl_contract(),
sl_exposure(),
sl_list_versions(),
sl_population(),
sl_population_crosstab(),
sl_read_records(),
sl_sample(),
sl_select_version()
Examples
boundaries <- readRDS(system.file("extdata", "sample-boundaries.rds",
package = "searchlight"
))
boundaries$lsoa21[1, "geography_code"]
Clear a marked searchlight cache
Description
Only removes searchlight-managed entries in a directory bearing the marker created by acquisition. Other user files are retained. This is destructive.
Usage
sl_cache_clear(dir = sl_cache_dir())
Arguments
dir |
Cache directory. |
Value
Invisibly, paths removed.
See Also
Other ingest:
sl_archive_download(),
sl_archive_index(),
sl_archive_snapshot(),
sl_assign_geography(),
sl_boundaries(),
sl_cache_dir(),
sl_contract(),
sl_exposure(),
sl_list_versions(),
sl_population(),
sl_population_crosstab(),
sl_read_records(),
sl_sample(),
sl_select_version()
Examples
d <- tempfile()
dir.create(d)
file.create(file.path(d, ".searchlight-cache"))
sl_cache_clear(d)
unlink(d, recursive = TRUE)
Locate the searchlight cache
Description
Returns the cache path without creating it. Acquisition functions
explicitly
create and write to this directory. Set options(searchlight.cache_dir = ...)
to override the platform default, or pass an explicit dir.
Usage
sl_cache_dir()
Value
A character path.
See Also
Other ingest:
sl_archive_download(),
sl_archive_index(),
sl_archive_snapshot(),
sl_assign_geography(),
sl_boundaries(),
sl_cache_clear(),
sl_contract(),
sl_exposure(),
sl_list_versions(),
sl_population(),
sl_population_crosstab(),
sl_read_records(),
sl_sample(),
sl_select_version()
Examples
sl_cache_dir()
Inspect an ingestion contract
Description
Contracts retain source provenance through subsetting and dplyr verbs. The
coverage and timestamp tables describe the ingested sources; scope
records
the currently visible row count and force-months after filtering. Filtering
does not turn a submitted file into a missing submission.
Usage
sl_contract(x)
Arguments
x |
Records, aggregates or a contract. |
Value
A named list of class sl_contract. summary() returns coverage.
See Also
sl_read_records(), sl_sample()
Other ingest:
sl_archive_download(),
sl_archive_index(),
sl_archive_snapshot(),
sl_assign_geography(),
sl_boundaries(),
sl_cache_clear(),
sl_cache_dir(),
sl_exposure(),
sl_list_versions(),
sl_population(),
sl_population_crosstab(),
sl_read_records(),
sl_sample(),
sl_select_version()
Examples
sl_contract(sl_sample())
Fit count models with population-time exposure
Description
Models recorded events, never unique people or a population-minus-stops cell. A population offset is multiplied by months_submitted/12 when available. Other offset columns must already represent exposure over the observed period. Unknown ethnicity and rows without positive exposure are reported as exclusions. Moran's I uses mean Pearson residuals per area and is exploratory after fitting. The exponentiated intercept is a baseline rate per exposure unit; other exponentiated coefficients are multiplicative rate effects. Confidence bounds are reported on this exponentiated scale.
Usage
sl_count_model(
counts,
formula,
family = c("poisson", "negbin"),
offset = "population",
random = NULL,
boundaries = NULL,
conf_level = 0.95
)
Arguments
counts |
Counts with a contract and exposure column, usually sl_rates. |
formula |
A formula with response n and desired fixed predictors. |
family |
Poisson or negative binomial. |
offset |
Name of the population or population-time exposure column. |
random |
Optional one-sided lme4 formula, for example ~ (1 | force_id). |
boundaries |
Optional sf polygons keyed by geography_code for residuals. |
conf_level |
Wald confidence level. |
Value
A tidy sl_count_model coefficient table with model, diagnostics and excluded row/event counts in attributes, carrying the ingestion contract.
See Also
sl_rates(), sl_spatial_disparity()
Other inference:
sl_hit_rates(),
sl_simulate(),
sl_spatial_disparity(),
sl_spatial_map(),
sl_veil_of_darkness()
Examples
r <- readRDS(system.file("extdata", "example-rates.rds",
package = "searchlight"
))
r <- r[r$ethnicity %in% c("White", "Black"), ]
sl_count_model(r, n ~ ethnicity)
# Supply matching boundaries to request the residual Moran diagnostic.
Count events while preserving submission strata
Description
Force and month always remain in the result, even if omitted from by: these identify independent reporting and time-exposure cells. Unknown ethnicity remains a level. Missing submissions have NA counts; submitted combinations without events have zero. Filtering records does not narrow the source time grid: pass months and forces explicitly to select the analysis period.
Usage
sl_counts(
records,
by = c("pfa", "month", "ethnicity_5"),
units = NULL,
months = NULL,
forces = NULL
)
Arguments
records |
Contract-bearing records. |
by |
Additional grouping columns, including exactly one ethnicity field. |
units |
Optional force_id and geography mapping defining the full area universe, including areas without events. Otherwise uses contract units or observed areas. PFA defaults to the published force-code mapping. |
months, forces |
Explicit analysis scope; defaults to contract coverage. |
Value
An sl_counts tibble with canonical geography_code, ethnicity, n, force_id and month columns, original grouping fields, and coverage status.
See Also
Other rates:
sl_rate_ratio(),
sl_rates()
Examples
counts <- sl_counts(sl_sample())
head(counts)
Audit submission coverage on a full force-month grid
Description
Counts come from the immutable selected files, including after filtering records. A submitted zero-row CSV is zero; an absent CSV is NA. Refreshes describe differing version counts. Partial suspicion takes precedence over refreshes, and uses preceding calendar months (at most twelve), not twelve observed files. A changelog issue does not create a nonexistent submission.
Usage
sl_coverage(
records,
forces = NULL,
months = NULL,
changelog = NULL,
partial_threshold = 0.2
)
Arguments
records |
Contract-bearing records. |
forces |
Force IDs; default is the contract's requested forces. |
months |
Months; default is the contract's requested months. |
changelog |
Table with force_id, month, note and optional issue flag. Default uses the bundled, dated changelog extract. |
partial_threshold |
Fraction of trailing median, default 0.2. |
Value
An sl_coverage tibble with contract; plot gives a coverage heatmap.
See Also
sl_quality(), sl_coverage_compare(), sl_benchmark()
Other audit:
sl_benchmark(),
sl_coverage_compare(),
sl_location_quality(),
sl_quality(),
sl_timestamp_quality()
Examples
audit <- sl_coverage(sl_sample())
audit
# plot(audit) displays the submission heatmap.
Compare coverage patterns across two periods
Description
Periods are explicit month vectors sorted chronologically. Comparability means equal length, matching submitted positions, at least one submitted position, and no partial suspicion. Identical missing positions are listed but allow comparison of the same observed subset. This tests coverage only, not seasonality, reporting practice or population comparability.
Usage
sl_coverage_compare(coverage, period_a, period_b)
Arguments
coverage |
Output of sl_coverage. |
period_a, period_b |
Month vectors for comparison. |
Value
A tibble with comparable, submitted positions, and a list of gaps.
See Also
Other audit:
sl_benchmark(),
sl_coverage(),
sl_location_quality(),
sl_quality(),
sl_timestamp_quality()
Examples
sl_coverage_compare(sl_coverage(sl_sample()), "2026-05", "2026-06")
Compare compatible population denominator scenarios
Description
Exposure tables must have identical geography/classification and cell keys. Each sampling interval is conditional on its exposure scenario; the range of point estimates across scenarios is reported separately.
Usage
sl_denominator_scenarios(
counts,
exposures,
reference = "White",
comparison = "Black"
)
Arguments
counts |
Contract-bearing event counts. |
exposures |
Named list of sl_exposure tables. |
reference, comparison |
Groups to compare. |
Value
An sl_sensitivity tibble with scenario ratios and assumption ranges.
See Also
sl_exposure(), sl_missing_ethnicity_bounds()
Other sensitivity:
sl_missing_ethnicity_bounds(),
sl_ranking_stability(),
sl_standardise()
Examples
p <- readRDS(system.file("extdata", "sample-population.rds",
package = "searchlight"
))$msoa21
c <- readRDS(system.file("extdata", "example-counts.rds",
package = "searchlight"
))
head(sl_denominator_scenarios(c, list(resident = p)))
Construct an explicit exposure table
Description
Population is an exposure. Unknown ethnicity has no Census population denominator and must have NA population. Additional age_band and sex columns identify disjoint strata. Counts may exceed exposure.
Usage
sl_exposure(
data,
geography,
classification = c("5", "19"),
source = "user_supplied"
)
Arguments
data |
A data frame with geography_code, ethnicity and population. |
geography |
Geography type and vintage, for example msoa21. |
classification |
Self-defined classification, 5 or 19. |
source |
Description of a user-supplied exposure. |
Value
A tibble with class sl_exposure and an exposure contract.
See Also
sl_population(), sl_population_crosstab()
Other ingest:
sl_archive_download(),
sl_archive_index(),
sl_archive_snapshot(),
sl_assign_geography(),
sl_boundaries(),
sl_cache_clear(),
sl_cache_dir(),
sl_contract(),
sl_list_versions(),
sl_population(),
sl_population_crosstab(),
sl_read_records(),
sl_sample(),
sl_select_version()
Examples
sl_exposure(data.frame(
geography_code = "area1", ethnicity = "White",
population = 1000
), geography = "example")
Describe three separate outcomes conditional on being searched
Description
Computes each requested outcome separately with Wilson binomial intervals. Unknown ethnicity remains in descriptive tables and missing outcomes are not failures. Logistic comparisons use known ethnicity, force and object fixed effects; constant controls are explicitly omitted. These are associations among searches. Differing hit rates alone do not establish discrimination; selection, differing risk distributions and infra-marginality matter.
Usage
sl_hit_rates(
records,
by = c("pfa", "object_group"),
outcome = c("any_action", "arrest", "outcome_linked_to_object"),
reference = "White",
comparison = "Black",
conf_level = 0.95
)
Arguments
records |
Search records with the three derived binary outcome measures. |
by |
Descriptive grouping columns; ethnicity_5 is always retained. |
outcome |
Separate outcomes to calculate (all three by default). |
reference, comparison |
Known self-defined ethnicity groups. |
conf_level |
Wilson and logistic Wald confidence level. |
Value
An sl_hit_rates tibble. Models, coefficients and exclusions describe separate logistic models. The ingestion contract is retained.
See Also
sl_read_records(), sl_veil_of_darkness()
Other inference:
sl_count_model(),
sl_simulate(),
sl_spatial_disparity(),
sl_spatial_map(),
sl_veil_of_darkness()
Examples
hits <- sl_hit_rates(sl_sample())
head(hits)
attr(hits, "coefficients")
List force-month archive versions
Description
List force-month archive versions
Usage
sl_list_versions(dir = sl_cache_dir())
Arguments
dir |
A snapshot cache or bundled sample directory. |
Value
A tibble with each force-month's archive, CSV hash and row count.
See Also
sl_archive_snapshot(), sl_select_version()
Other ingest:
sl_archive_download(),
sl_archive_index(),
sl_archive_snapshot(),
sl_assign_geography(),
sl_boundaries(),
sl_cache_clear(),
sl_cache_dir(),
sl_contract(),
sl_exposure(),
sl_population(),
sl_population_crosstab(),
sl_read_records(),
sl_sample(),
sl_select_version()
Examples
sl_list_versions(system.file("extdata", "sample", package = "searchlight"))
Diagnose anonymised location quality
Description
Snap-point concentration is the share at the most common reported coordinate within an LSOA, among records with usable coordinates in that LSOA. It is not the concentration at true incident locations. Missing LSOA assignments remain a separate group. Per-geography boundary diagnostics remain in the contract.
Usage
sl_location_quality(records)
Arguments
records |
Contract-bearing records, optionally with assigned LSOAs. |
Value
A force-month-LSOA tibble with missingness, concentration and boundary sensitivity, carrying the contract.
See Also
sl_assign_geography(), sl_timestamp_quality()
Other audit:
sl_benchmark(),
sl_coverage(),
sl_coverage_compare(),
sl_quality(),
sl_timestamp_quality()
Examples
sl_location_quality(sl_sample()[1:100, ])
Bound disparity under allocations of Unknown ethnicity
Description
Extreme allocations send every Unknown event to one of the two groups. Proportional allocation uses all known groups, not only the comparison pair. force_object_mar uses known-group shares within force and object, pooled over areas and months; it refuses strata with no known ethnicity. These are assumptions, not confidence intervals. Force/object allocation also assumes pooled known composition applies to each recipient area/month; MAR alone does not establish that transportability. The tipping share solves the equation where the ratio equals one, with all remaining Unknown sent to the reference. A tipping value outside zero to one means no feasible tipping allocation.
Usage
sl_missing_ethnicity_bounds(
counts,
population,
reference = "White",
comparison = "Black",
scenarios = c("all_to_reference", "all_to_comparison", "proportional",
"force_object_mar")
)
Arguments
counts |
Contract-bearing event counts. |
population |
Marginal ethnic-group exposure table. |
reference, comparison |
Groups to compare. |
scenarios |
Allocation assumptions to report. |
Value
An sl_sensitivity tibble with scenario ratios, extreme bounds, tipping allocation and separately labelled baseline sampling intervals.
See Also
sl_rates(), sl_denominator_scenarios()
Other sensitivity:
sl_denominator_scenarios(),
sl_ranking_stability(),
sl_standardise()
Examples
c <- readRDS(system.file("extdata", "example-counts.rds",
package = "searchlight"
))
p <- readRDS(system.file("extdata", "sample-population.rds",
package = "searchlight"
))$msoa21
head(sl_missing_ethnicity_bounds(c, p))
Census 2021 ethnicity populations
Description
Fetches TS021 (NM_2041_1), with exact codes verified against returned cells. Supply ONS codes or an sf boundary object. Current ward and LAD boundaries must not be silently matched to Census 2021 units: unchanged codes alone do not demonstrate unchanged boundaries. A new exposure source must be supplied for changed units. Census disclosure control may cause small inconsistencies between separately published tables. Unknown is retained with NA exposure.
Usage
sl_population(geography, classification = c("5", "19"), dir = sl_cache_dir())
Arguments
geography |
ONS geography codes, or an sf layer with geography_code. |
classification |
Census classification, 5 or 19. |
dir |
Explicit cache directory, written only on this call. |
Value
An sl_exposure tibble, or NULL with a message if unavailable.
See Also
sl_exposure(), sl_population_crosstab()
Other ingest:
sl_archive_download(),
sl_archive_index(),
sl_archive_snapshot(),
sl_assign_geography(),
sl_boundaries(),
sl_cache_clear(),
sl_cache_dir(),
sl_contract(),
sl_exposure(),
sl_list_versions(),
sl_population_crosstab(),
sl_read_records(),
sl_sample(),
sl_select_version()
Examples
p <- readRDS(system.file("extdata", "sample-population.rds",
package = "searchlight"
))
head(p$msoa21)
Census ethnicity by age and sex
Description
Uses RM032 (NM_2132_1), available at Census 2021 small areas. Five published age bands are combined to under 25, 25-34 and 35+, the common partition with police age bands. Census sex and police-recorded gender are different measurements; any analysis using them must state this assumption. Unsupported geography returns NULL and a message. No age-sex cross-tabulation is imputed from marginal totals.
Usage
sl_population_crosstab(
geography,
dir = sl_cache_dir(),
classification = c("5", "19")
)
Arguments
geography |
ONS geography codes, or an sf layer with geography_code. |
dir |
Explicit cache directory, written only on this call. |
classification |
Census classification, 5 or 19. |
Value
An sl_exposure tibble with age_band and sex, or NULL.
See Also
sl_population(), sl_exposure()
Other ingest:
sl_archive_download(),
sl_archive_index(),
sl_archive_snapshot(),
sl_assign_geography(),
sl_boundaries(),
sl_cache_clear(),
sl_cache_dir(),
sl_contract(),
sl_exposure(),
sl_list_versions(),
sl_population(),
sl_read_records(),
sl_sample(),
sl_select_version()
Examples
p <- readRDS(system.file("extdata", "sample-crosstab.rds",
package = "searchlight"
))
head(p)
Attach submission, timestamp and location diagnostics
Description
Pass population to complete the population metadata in the records contract; population counts remain a separate exposure table. Diagnostics describe the current records; coverage continues to describe the original source files.
Usage
sl_quality(records, population = NULL, threshold = 0.05, changelog = NULL)
Arguments
records |
Contract-bearing records. |
population |
Optional sl_exposure table. |
threshold |
Midnight-share threshold. |
changelog |
Optional changelog table passed to sl_coverage. |
Value
Records with an updated contract and attached quality tables.
See Also
sl_coverage(), sl_timestamp_quality(), sl_location_quality()
Other audit:
sl_benchmark(),
sl_coverage(),
sl_coverage_compare(),
sl_location_quality(),
sl_timestamp_quality()
Examples
records <- sl_quality(sl_sample()[1:100, ])
sl_contract(records)$timestamps
Assess area rank uncertainty under count models or posterior draws
Description
Bootstrap draws event counts from the fitted Poisson or negative-binomial model, then ranks event-rate ratios (largest is rank 1). Quasi-Poisson does not define a count distribution and is refused. Positive totals in both groups are required for plug-in bootstrap; sparse zero-count areas need a suitable posterior model. Ties receive random ranks under the saved seed. Optional scenario rows are analysed separately, never pooled into a sampling distribution. Pairwise stability means ordering probability above 0.95.
Usage
sl_ranking_stability(
rate_ratios,
method = c("bootstrap", "posterior"),
n = 1000,
seed = 1
)
Arguments
rate_ratios |
An sl_rate_ratio table, or a contract-bearing table with geography_code and a posterior_draws matrix attribute (draws by area). |
method |
Parametric bootstrap or posterior ranks. |
n |
Number of draws; default 1000. |
seed |
Reproducible seed, restored on exit. |
Value
An sl_sensitivity tibble of median/95% rank intervals, with attributes rank_probabilities (matrix) and pairwise (ordering probabilities).
See Also
sl_rate_ratio(), sl_denominator_scenarios()
Other sensitivity:
sl_denominator_scenarios(),
sl_missing_ethnicity_bounds(),
sl_standardise()
Examples
# Posterior draws can also be supplied by an independently fitted model.
x <- tibble::tibble(geography_code = c("a", "b"))
attr(x, "contract") <- sl_contract(sl_sample())
attr(x, "posterior_draws") <- cbind(a = c(1, 2, 3), b = c(3, 2, 1))
sl_ranking_stability(x, "posterior", n = 3)
Estimate an event stop-rate ratio with exposure offsets
Description
Fits n ~ ethnicity with log(population-time) offset within each area and other grouping stratum, pooling submitted months. Poisson intervals are exact conditional intervals for two event counts; zero counts are supported without adding pseudocounts. Quasi-Poisson and negative binomial require replicated cells and use model-based log intervals. Dispersion is Pearson chi-square divided by residual degrees of freedom. This is not a person-level risk ratio.
Usage
sl_rate_ratio(
rates,
reference = "White",
comparison = "Black",
method = c("poisson", "quasipoisson", "negbin"),
conf_level = 0.95
)
Arguments
rates |
Output from sl_rates. |
reference, comparison |
Self-defined ethnicity groups. |
method |
Count model family. |
conf_level |
Sampling confidence level. |
Value
An sl_rate_ratio tibble with sampling intervals, dispersion, counts, exposures and fitted model list column, carrying the ingestion contract.
See Also
sl_rates(), sl_missing_ethnicity_bounds()
Other rates:
sl_counts(),
sl_rates()
Examples
r <- readRDS(system.file("extdata", "example-rates.rds",
package = "searchlight"
))
head(sl_rate_ratio(r))
Calculate rates using observed population-time exposure
Description
The primary rate is annualised events per 1,000 resident person-years: n / (population * submitted_months / 12) * per. The separate period_rate is n / population * per for that cell. Neither is a probability of a person being searched. NA submission counts are excluded from time exposure. Partial submissions remain flagged and describe reported events only.
Usage
sl_rates(counts, population, per = 1000)
Arguments
counts |
Output from sl_counts. |
population |
A compatible sl_exposure table. |
per |
Rate scaling, default 1000. |
Value
An sl_rates tibble with population, exposure, rate and period_rate.
See Also
sl_counts(), sl_rate_ratio(), sl_exposure()
Other rates:
sl_counts(),
sl_rate_ratio()
Examples
c <- readRDS(system.file("extdata", "example-counts.rds",
package = "searchlight"
))
p <- readRDS(system.file("extdata", "sample-population.rds",
package = "searchlight"
))$msoa21
head(sl_rates(c, p))
Read selected immutable archive records
Description
Validates the fixed CSV schema, checksum and row count, preserving every event (including identical rows) and raw ethnicity strings. Timestamps retain their supplied UTC offsets and are displayed in Europe/London. Absent or refused ethnicity is Unknown. Missing outcome measures remain NA, not negative outcomes. The filename defines the submission month. Offset conversion can cross a month boundary; date_month_crossing records this without moving the event into another source file. A timestamp must match the filename in either its supplied clock or the London clock.
Usage
sl_read_records(dir = sl_cache_dir(), versions = NULL)
Arguments
dir |
Snapshot cache directory. |
versions |
One selected version per force-month; defaults to latest. |
Value
An sl_records tibble with a validated ingestion contract.
See Also
sl_contract(), sl_select_version(), sl_sample()
Other ingest:
sl_archive_download(),
sl_archive_index(),
sl_archive_snapshot(),
sl_assign_geography(),
sl_boundaries(),
sl_cache_clear(),
sl_cache_dir(),
sl_contract(),
sl_exposure(),
sl_list_versions(),
sl_population(),
sl_population_crosstab(),
sl_sample(),
sl_select_version()
Examples
d <- system.file("extdata", "sample", package = "searchlight")
v <- sl_select_version(sl_list_versions(d))
records <- sl_read_records(d, v[1, ])
summary(sl_contract(records))
Write an offline report of audited records and analysis results
Description
Renders the installed Markdown/Rmd template to a self-contained HTML file. No Pandoc, network, model fitting or external web assets are required. Supplied estimates must retain the same source snapshots and ethnicity classification as records. Their analysis scope and diagnostics are shown. Coverage describes source files; filtering does not redefine coverage. Sampling intervals and assumption ranges have separate labelled sections. The report includes a summary, section navigation and expandable audit details. Tables retain their original column names in header tooltips.
Usage
sl_report(
records,
file,
estimates = list(),
title = "Stop and search: evidence and assumptions",
assumptions = character(),
limitations = character(),
max_rows = 500,
overwrite = FALSE
)
Arguments
records |
Contract-bearing search records. |
file |
Explicit HTML output path. Its parent directory must exist. |
estimates |
Uniquely named list of contract-bearing result tables. |
title |
Report title, treated as text. |
assumptions, limitations |
Additional statements, treated as text. |
max_rows |
Maximum displayed rows per table; truncation is labelled. |
overwrite |
Whether to replace an existing output file. |
Value
Invisibly, a one-row tibble with the output path and record count, carrying the ingestion contract. Writes only the requested output file.
See Also
sl_coverage(), sl_rate_ratio(), sl_missing_ethnicity_bounds()
Examples
rates <- readRDS(system.file("extdata", "example-rates.rds",
package = "searchlight"
))
path <- tempfile(fileext = ".html")
sl_report(sl_sample(), path, estimates = list(ratios = sl_rate_ratio(rates)))
file.exists(path)
Load the bundled real archive sample
Description
Load the bundled real archive sample
Usage
sl_sample()
Value
An sl_records tibble for the documented sample forces and months.
See Also
sl_read_records(), sl_contract()
Other ingest:
sl_archive_download(),
sl_archive_index(),
sl_archive_snapshot(),
sl_assign_geography(),
sl_boundaries(),
sl_cache_clear(),
sl_cache_dir(),
sl_contract(),
sl_exposure(),
sl_list_versions(),
sl_population(),
sl_population_crosstab(),
sl_read_records(),
sl_select_version()
Examples
x <- sl_sample()
table(x$force_id, x$month)
Select one complete version per force-month
Description
Never deduplicates event rows. Ties are resolved by snapshot month and
archive
name. max_rows chooses the newest version among equal row counts.
Differences
greater than 1% of the smallest count are reported (zero to positive
is infinite).
Usage
sl_select_version(
versions,
rule = c("latest", "earliest", "max_rows", "manual"),
manual = NULL
)
Arguments
versions |
Output from |
rule |
Version selection rule. |
manual |
For the manual rule, a tibble with |
Value
Selected versions with list-column alternatives and a differences
attribute. Each row carries the selection rule.
See Also
Other ingest:
sl_archive_download(),
sl_archive_index(),
sl_archive_snapshot(),
sl_assign_geography(),
sl_boundaries(),
sl_cache_clear(),
sl_cache_dir(),
sl_contract(),
sl_exposure(),
sl_list_versions(),
sl_population(),
sl_population_crosstab(),
sl_read_records(),
sl_sample()
Examples
d <- system.file("extdata", "sample", package = "searchlight")
v <- sl_list_versions(d)
sl_select_version(v)
Simulate event counts with independent disparity and total intensity
Description
The disparity surface varies east-west and expected total events north-south. Their lattice correlation is zero by construction. Counts are Poisson events, then unknown ethnicity is generated by binomial thinning. MCAR has a common probability, MAR varies by synthetic force, and MNAR varies by ethnicity. True latent rate ratios and complete realised event ratios are distinct.
Usage
sl_simulate(
side = 6,
surface = c("smooth", "discontinuous", "small"),
population = c(White = 5000, Black = 1000),
rate = 0.025,
missingness = c("none", "mcar", "mar", "mnar"),
missing_rate = 0.2,
months = 12,
seed = 1
)
Arguments
side |
Lattice width and height, at least three. |
surface |
Smooth, discontinuous, or small comparison-population setting. |
population |
Resident reference/comparison populations per area. |
rate |
Mean total annual event rate per resident. |
missingness |
Missing-ethnicity mechanism. |
missing_rate |
Baseline missing probability. |
months |
Number of submitted months, between one and twelve. |
seed |
Reproducible seed, restored on exit. |
Value
A list of counts, complete_counts, population, boundaries, truth and settings. Counts carry a clearly synthetic ingestion contract.
See Also
sl_spatial_disparity(), sl_missing_ethnicity_bounds()
Other inference:
sl_count_model(),
sl_hit_rates(),
sl_spatial_disparity(),
sl_spatial_map(),
sl_veil_of_darkness()
Examples
s <- sl_simulate(side = 3, missingness = "mnar", missing_rate = 0.3)
s$truth
sl_missing_ethnicity_bounds(s$counts, s$population,
scenarios = c("all_to_reference", "all_to_comparison")
)
Estimate a multivariate spatial stop-rate disparity surface
Description
Uses CARBayes MVS.CARleroux with two Poisson responses and a matrix of log population-time offsets. Each ethnicity has its own spatial effect phi_ig. The equivalent shared/contrast decomposition is u_i=(phi_ir+phi_ic)/2, v_ir=(phi_ir-phi_ic)/2 and v_ic=-v_ir. These components have the covariance induced by the multivariate prior; they are not separately independent priors. The full posterior ratio is exp(alpha_c-alpha_r+phi_ic-phi_ir). Smoothing stabilises under stated assumptions; it cannot repair exposure, missing ethnicity, missing submissions or anonymised location errors.
Usage
sl_spatial_disparity(
counts,
population,
boundaries,
reference = "White",
comparison = "Black",
backend = c("carbayes", "inla"),
burnin = 2000,
n.sample = 12000,
thin = 10,
chains = 2,
k = 2,
adjacency = NULL,
boundary_threshold = 0.2,
seed = 1,
...
)
Arguments
counts |
Marginal event counts with a valid contract. |
population |
Compatible marginal ethnicity exposure table. |
boundaries |
sf polygons with geography_code. |
reference, comparison |
Known self-defined ethnicity groups. |
backend |
CARBayes. INLA is explicitly deferred beyond version 0.1.0. |
burnin, n.sample, thin |
MCMC controls per chain, including burn-in in n.sample. |
chains |
Number of independent chains, at least two, run on one core. |
k |
Additional ratio threshold for exceedance probability. |
adjacency |
Optional symmetric adjacency matrix named with geography codes. Otherwise rook adjacency is used; islands require an explicit decision. |
boundary_threshold |
Warn if the assigned boundary-sensitive share exceeds this proportion (default 0.2). |
seed |
Reproducible seed, restored on exit. |
... |
Prior and sampler arguments for MVS.CARleroux, such as rho or MALA. |
Value
An sl_spatial tibble with 90/95 percent intervals, exceedance probabilities, crude ratios, rank-normalised R-hat and bulk/tail ESS. Joint posterior draws, fitted model, diagnostics and boundaries are attributes.
See Also
sl_simulate(), sl_spatial_map(), sl_ranking_stability()
Other inference:
sl_count_model(),
sl_hit_rates(),
sl_simulate(),
sl_spatial_map(),
sl_veil_of_darkness()
Examples
# A precomputed synthetic fit avoids MCMC in installed examples.
path <- system.file("extdata", "example-spatial.rds",
package = "searchlight"
)
if (nzchar(path)) {
fit <- readRDS(path)
head(fit)
head(sl_spatial_map(fit))
}
Join spatial posterior summaries to their polygons
Description
Join spatial posterior summaries to their polygons
Usage
sl_spatial_map(x)
Arguments
x |
An sl_spatial fit. |
Value
An sf object with summaries and its ingestion contract.
See Also
Other inference:
sl_count_model(),
sl_hit_rates(),
sl_simulate(),
sl_spatial_disparity(),
sl_veil_of_darkness()
Examples
path <- system.file("extdata", "example-spatial.rds",
package = "searchlight"
)
if (nzchar(path)) head(sl_spatial_map(readRDS(path)))
Directly standardise age-sex stop-event rates
Description
Uses common age-sex weights across ethnicity groups. Census sex is compared with police-recorded gender only under an explicit measurement assumption. Unknown age, gender and ethnicity are retained in exclusion diagnostics. A positively weighted stratum with zero exposure makes the standardised rate undefined; it is never silently dropped. Sampling intervals combine exact Poisson stratum intervals with a Bonferroni correction and fixed weights. They are conservative, including when every observed stratum count is zero.
Usage
sl_standardise(
counts,
crosstab,
standard = c("england_wales", "study_population"),
per = 1000
)
Arguments
counts |
Counts grouped by age_band and sex. |
crosstab |
Compatible Census or explicit ethnicity-age-sex exposure. |
standard |
England/Wales Census 2021 or pooled study population weights. |
per |
Rate scaling, default 1000 per person-year. |
Value
An sl_sensitivity tibble with crude and standardised rates, sampling intervals, excluded-event counts, weights and the ingestion contract.
See Also
sl_population_crosstab(), sl_rates()
Other sensitivity:
sl_denominator_scenarios(),
sl_missing_ethnicity_bounds(),
sl_ranking_stability()
Examples
p <- readRDS(system.file("extdata", "sample-crosstab.rds",
package = "searchlight"
))
c <- readRDS(system.file("extdata", "example-demographic-counts.rds",
package = "searchlight"
))
head(suppressWarnings(sl_standardise(c, p)))
Diagnose timestamp completeness and midnight heaping
Description
Reports both London-clock midnight and midnight in the source string. Some date-only summer records carry a UTC offset, so checking London midnight alone would miss them. Reliability requires both shares below threshold and no missing timestamps. This is a screening diagnostic, not proof that all supplied times are accurate. Empty and absent force-months are not reliable.
Usage
sl_timestamp_quality(records, threshold = 0.05)
Arguments
records |
Contract-bearing records. |
threshold |
Maximum permitted midnight share, default 0.05. |
Value
A tibble with force-month diagnostics and within-force share range.
See Also
sl_quality(), sl_location_quality()
Other audit:
sl_benchmark(),
sl_coverage(),
sl_coverage_compare(),
sl_location_quality(),
sl_quality()
Examples
sl_timestamp_quality(sl_sample())
Diagnose ethnicity composition in daylight and darkness
Description
Excludes every force-month that fails or lacks the contract's time-quality screen, retaining an explicit exclusion table. Annual evening clock-time bounds use each distinct location and year. Civil darkness begins at dusk; ambiguous sunset-to-dusk observations are excluded. Times and clock changes use Europe/London, including British summer time. The logistic comparison controls clock time (natural spline), weekday, month and force where these vary. The DST option further restricts to weeks around both clock changes and adds transition and running-day controls; it is a local darkness comparison. The method was developed for vehicle stops. Transfer to pedestrian searches requires assumptions about visibility, activity, deployment and selection.
Usage
sl_veil_of_darkness(
records,
boundaries = NULL,
twilight = c("civil", "sunset"),
window = "intertwilight",
design = c("intertwilight", "dst"),
reference = "White",
comparison = "Black",
dst_weeks = 3,
conf_level = 0.95
)
Arguments
records |
Contract-bearing search records. |
boundaries |
Optional sf study region; restrict published points to it. Missing locations are never replaced by centroids. |
twilight |
Civil twilight end or sunset definition of darkness. |
window |
Currently only intertwilight, the annual evening overlap window. |
design |
Annual intertwilight or local DST-window comparison. |
reference, comparison |
Known self-defined ethnicity groups. |
dst_weeks |
Weeks on either side of each clock change for the DST design. |
conf_level |
Logistic Wald confidence level. |
Value
An sl_veil tibble with an odds ratio, interval and fit status. Attributes steps, excluded_force_months, assumptions, data and model retain the full design audit. A gated or unidentifiable sample has an undefined ratio.
See Also
sl_timestamp_quality(), sl_hit_rates()
Other inference:
sl_count_model(),
sl_hit_rates(),
sl_simulate(),
sl_spatial_disparity(),
sl_spatial_map()
Examples
# Date-only records are explicitly refused by the time-quality screen.
x <- sl_sample()[1:10, ]
day <- as.Date(x$date, tz = "Europe/London")
x$date <- as.POSIXct(paste(day, "00:00:00"), tz = "Europe/London")
x$date_raw <- format(x$date, "%Y-%m-%dT%H:%M:%S%z")
attr(x, "contract")$timestamps <- sl_timestamp_quality(x)
sl_veil_of_darkness(x)