---
title: "Getting Started with deriva"
output: rmarkdown::html_vignette
vignette: >
  %\VignetteIndexEntry{Getting Started with deriva}
  %\VignetteEngine{knitr::rmarkdown}
  %\VignetteEncoding{UTF-8}
---

```{r setup, include = FALSE}
knitr::opts_chunk$set(
  collapse = TRUE,
  comment = "#>"
)
old_options <- options(width = 70)
```

```{r load}
library(deriva)
```

## What is concept drift?

A machine learning model trained on historical data operates under the implicit assumption that the data-generating process remains stable over time. When this assumption breaks down — because user behaviour shifts, sensor calibration drifts, or the world simply changes — the model's predictions degrade without any obvious error being raised. This phenomenon is called **concept drift**.

Monitoring for drift requires a *drift detector*: an algorithm that reads a stream of per-observation signals (typically prediction errors) and raises a flag when the signal's distribution has changed significantly. The `deriva` package provides a tidy interface to a catalogue of 22 such detectors, designed to compose naturally with the tidymodels ecosystem.

## The running example: a credit-approval model

This vignette works through `credit_monitoring`, a dataset shipped with the package: the per-observation error stream of a credit-approval classifier in production.

```{r data}
head(credit_monitoring)
```

For its first 500 observations the model runs at its expected 5% error rate. After observation 500, a shift in the credit market raises the error rate to 30% — `drift_true` records this ground truth, which a real deployment would not have access to (`deriva`'s job is to infer it from `error` alone).

## Quick start

The one-shot shortcut `detect_drift()` runs a detector over an existing column and returns the data annotated with `.warning` and `.drift` flags.

```{r quickstart}
result <- detect_drift(credit_monitoring, .col = error, method = "ddm")

# Where was drift flagged?
subset(result, .drift)
```

The detector correctly identifies the distributional change shortly after the known drift point (observation 500), with no false drift detections in the 500 stable observations before it.

## The deriva interface

`deriva` follows the same three-verb pattern as tidymodels: **specify → fit → advance**.

### 1. Specify a detector

`drift_detector()` creates an inert specification — no computation happens here.

```{r spec}
spec <- drift_detector("ddm", min_instances = 30)
spec
```

Pass method hyperparameters as named arguments. Unknown parameters, and values outside a parameter's valid range, raise an informative error.

### 2. Fit on a baseline

`fit()` runs the detector over the **baseline period** — the stable window against which future observations are compared. Here, that is the first 500 observations, before the market shift.

```{r fit}
baseline <- credit_monitoring[credit_monitoring$t <= 500, ]

fitted <- fit(spec, baseline, signal = error)
fitted
```

The fitted object is **immutable**: it stores the internal engine state after processing the baseline, ready to receive new data.

### 3. Advance over new batches

`advance()` feeds a new batch to the detector and returns a **new** fitted object with the state updated and the annotated batch appended to the history. The original object is not modified.

```{r advance}
batch1 <- credit_monitoring[credit_monitoring$t >= 501 & credit_monitoring$t <= 700, ]
batch2 <- credit_monitoring[credit_monitoring$t >= 701, ]

fitted2 <- advance(fitted,  batch1)
fitted3 <- advance(fitted2, batch2)
fitted3
```

Batches can be any size — including a single observation for true streaming use. The market shift (observation 501) falls inside `batch1`; `deriva` flags it there, and `batch2` confirms the model has settled into its new, worse error rate with no further alarms.

### 4. Inspect results

**`augment()`** returns the retained history as a tibble (the last `keep` rows of the fitted object — see `?drift_detector`).

```{r augment}
history <- augment(fitted3)
tail(history[, c("t", "error", ".phase", ".warning", ".drift")], 10)
```

**`tidy()`** extracts the detected drift points.

```{r tidy}
tidy(fitted3)
```

**`glance()`** gives a one-row summary.

```{r glance}
glance(fitted3)
```

**`autoplot()`** plots the running mean of the signal with warning (orange) and drift (red) markers, and a dotted line separating baseline from stream.

```{r autoplot, eval = requireNamespace("ggplot2", quietly = TRUE), fig.width = 6.5, fig.height = 3.5}
library(ggplot2)
autoplot(fitted3)
```

## Bridging from tidymodels

In a real workflow, the signal column comes from model predictions, not a shipped dataset. `add_prediction_error()` converts the output of tidymodels' `augment()` (which contains truth and estimate columns) into a `.error` column that drift detectors can consume.

```{r bridge}
# Simulate tidymodels augment() output for a classifier
predictions <- data.frame(
  time     = 1:8,
  truth    = factor(c("yes","no","yes","yes","no","yes","no","yes")),
  .pred_class = factor(c("yes","no","yes","no" ,"no","no" ,"no","yes"))
)

add_prediction_error(predictions, truth = truth)
```

For regression problems, `.error` is the absolute prediction error; for classification it is a 0/1 mismatch indicator. Chaining straight into `fit()` needs care: `fit()`'s first argument is the *spec*, not the data, so build the annotated data first and pass the spec and data to `fit()` explicitly:

```r
monitoring_data <- model |>
  augment(new_data = production_data) |>
  add_prediction_error(truth = y)

fit(drift_detector("page_hinkley"), monitoring_data, signal = .error)
```

## Distribution-based detectors

Some detectors monitor the distribution of a numeric stream directly, without requiring labelled errors — see `vignette("distribution-detectors")` for a worked example with `"kswin"`.

## Available methods

`deriva` ships with 22 drift detectors across two signal types.

| Signal type | Methods |
|---|---|
| `"error"` (0/1 errors) | `ddm`, `eddm`, `hddm_a`, `hddm_w`, `ewma`, `rddm`, `stepd`, `fhddm`, `fhddms`, `mddm_a`, `mddm_e`, `mddm_g`, `wstd`, `ftdd`, `fpdd`, `fsdd` |
| `"distribution"` (numeric stream) | `kswin`, `adwin`, `page_hinkley`, `cusum`, `seed`, `seqdrift2` |

Use `drift_detector("<method>")` to inspect default hyperparameters for any method.

## Summary

The core `deriva` workflow is:

```r
drift_detector("ddm") |>          # specify
  fit(baseline, signal = error) |> # learn reference level
  advance(new_batch)               # update state, persist flags
```

Supplementary verbs — `augment()`, `tidy()`, `glance()`, `autoplot()` — follow the tidymodels convention and make it straightforward to inspect, summarise, and plot detection results at any point in the stream.

```{r cleanup, include = FALSE}
options(old_options)
```
