---
title: "Spatial Cross-Validation Methods"
author: "Mamadou SOW"
date: "`r Sys.Date()`"
output: rmarkdown::html_vignette
vignette: >
  %\VignetteIndexEntry{Spatial Cross-Validation Methods}
  %\VignetteEngine{knitr::rmarkdown}
  %\VignetteEncoding{UTF-8}
---

```{r setup, include=FALSE}
knitr::opts_chunk$set(echo = TRUE, warning = FALSE, message = FALSE)
```

## Overview

This vignette demonstrates the different spatial cross-validation methods available in `spatialcvR` and when to use each approach.

## Available Methods

`spatialcvR` implements four main spatial cross-validation methods:

1. **Spatial Block CV**: Divide space into rectangular blocks
2. **Buffered CV**: Exclude training observations within a buffer radius
3. **Spatial Clustering CV**: Group observations spatially using clustering
4. **Random Split**: Baseline random CV for comparison

## Setup

```{r}
library(spatialcvR)

# Load sample data
data(sample_spatial_data)
head(sample_spatial_data)
```

## Spatial Block Cross-Validation

### Concept

Spatial block CV divides the study area into a grid of rectangular blocks. Observations are assigned to folds based on which block they fall into. This ensures spatial separation between training and test sets.

### When to Use

- Data with relatively uniform spatial distribution
- When you want to ensure geographic coverage
- When computational efficiency is important
- As a default spatial CV method

### Basic Usage

```{r}
# Create spatial block folds
folds_block <- spatial_folds(
  data = sample_spatial_data,
  x = "longitude",
  y = "latitude",
  k = 5,
  method = "block",
  seed = 123
)

print(folds_block)
```

### Controlling Block Size

You can control the spatial resolution using either block size or number of blocks:

```{r}
# Specify block size
folds_block_size <- spatial_block_folds(
  data = sample_spatial_data,
  x = "longitude",
  y = "latitude",
  k = 5,
  block_size = c(200, 200),  # 200x200 unit blocks
  seed = 123
)

# Specify number of blocks
folds_n_blocks <- spatial_block_folds(
  data = sample_spatial_data,
  x = "longitude",
  y = "latitude",
  k = 5,
  n_blocks = c(5, 5),  # 5x5 grid
  seed = 123
)
```

### Assignment Strategies

Blocks can be assigned to folds using different strategies:

```{r}
# Systematic assignment (default)
folds_systematic <- spatial_block_folds(
  data = sample_spatial_data,
  x = "longitude",
  y = "latitude",
  k = 5,
  assignment = "systematic",
  seed = 123
)

# Random assignment
folds_random_assign <- spatial_block_folds(
  data = sample_spatial_data,
  x = "longitude",
  y = "latitude",
  k = 5,
  assignment = "random",
  seed = 123
)
```

## Buffered Cross-Validation

### Concept

Buffered CV ensures that for each test observation, no training observation falls within a specified buffer radius. This provides strict control over the minimum spatial separation.

### When to Use

- When you need precise control over train/test distance
- For point-based data with irregular sampling
- When buffer distance has ecological/physical meaning
- For conservative performance estimates

### Basic Usage

```{r}
# Create buffered folds
folds_buffer <- spatial_buffer_folds(
  data = sample_spatial_data,
  x = "longitude",
  y = "latitude",
  k = 5,
  buffer_radius = 100,  # 100 unit buffer
  seed = 123
)

print(folds_buffer)
```

### Important Notes

- Buffered CV requires a **projected CRS** for accurate distance calculations
- Using geographic coordinates (longitude/latitude) will produce approximate distances
- Larger buffer radii reduce the size of training sets
- This method is computationally more intensive than block CV

## Spatial Clustering Cross-Validation

### Concept

Spatial clustering CV groups spatially proximate observations using clustering algorithms (k-means), then assigns clusters to folds. This is useful for data with complex spatial structure.

### When to Use

- Data with irregular or clustered spatial patterns
- When natural spatial groupings exist
- For heterogeneous spatial distributions
- When block boundaries would be arbitrary

### Basic Usage

```{r}
# Create clustering folds
folds_cluster <- spatial_cluster_folds(
  data = sample_spatial_data,
  x = "longitude",
  y = "latitude",
  k = 5,
  n_clusters = 10,  # Number of spatial clusters
  seed = 123
)

print(folds_cluster)
```

### Understanding Cluster Assignment

```{r}
# Examine cluster centers
folds_cluster$parameters$cluster_centers
```

## Random Spatial Split (Baseline)

### Concept

Random spatial split performs standard random k-fold cross-validation without spatial constraints. This serves as a baseline to demonstrate the impact of spatial dependence.

### When to Use

- As a baseline for comparison
- When spatial dependence is minimal
- For initial exploratory analysis
- To demonstrate the value of spatial CV

### Basic Usage

```{r}
# Create random folds
folds_random <- spatial_split(
  data = sample_spatial_data,
  x = "longitude",
  y = "latitude",
  k = 5,
  seed = 123
)

print(folds_random)
```

## Choosing the Right Method

### Decision Flowchart

1. **Start with spatial block CV** - Good default choice
2. **Need precise distance control?** → Use buffered CV
3. **Complex spatial patterns?** → Use clustering CV
4. **Compare with random CV** → Always include baseline

### Method Comparison

| Method | Pros | Cons | Best For |
|--------|------|------|----------|
| Block CV | Fast, intuitive, geographic coverage | May split natural clusters | Most cases |
| Buffered CV | Precise distance control | Computationally intensive | Point data, strict requirements |
| Clustering CV | Handles complex patterns | Sensitive to cluster parameters | Heterogeneous data |
| Random CV | Fast, baseline | No spatial control | Comparison, minimal spatial dependence |

## Visualizing Folds

```{r}
# Plot individual fold
plot_spatial_folds(folds_block, sample_spatial_data, "longitude", "latitude", 
                   fold = 1, main = "Block CV - Fold 1")

# Plot all folds
plot_spatial_folds(folds_block, sample_spatial_data, "longitude", "latitude", 
                   fold = "all", main = "Block CV - All Folds")
```

## Practical Tips

### Start Simple

Begin with spatial block CV using default parameters:

```{r}
# Default spatial block CV
folds_default <- spatial_folds(
  data = sample_spatial_data,
  x = "longitude",
  y = "latitude",
  k = 5
)
```

### Adjust Based on Data Characteristics

- **Dense data**: Use larger blocks or more clusters
- **Sparse data**: Use smaller blocks or fewer clusters
- **Strong spatial patterns**: Consider clustering CV
- **Precise distance requirements**: Use buffered CV

### Always Compare with Random CV

```{r}
# Compare spatial vs random
folds_spatial <- spatial_folds(sample_spatial_data, "longitude", "latitude", 
                               k = 5, method = "block", seed = 123)
folds_random <- spatial_folds(sample_spatial_data, "longitude", "latitude", 
                             k = 5, method = "random", seed = 123)

# Analyze spatial leakage for both
leakage_spatial <- detect_spatial_leakage(sample_spatial_data, folds_spatial, 
                                           "longitude", "latitude")
leakage_random <- detect_spatial_leakage(sample_spatial_data, folds_random, 
                                          "longitude", "latitude")

print(leakage_spatial)
print(leakage_random)
```

## Common Issues and Solutions

### Issue: Too few observations per fold

**Solution**: Reduce k or use fewer blocks/clusters

```{r}
# Reduce number of folds
folds_k3 <- spatial_folds(sample_spatial_data, "longitude", "latitude", k = 3)
```

### Issue: Uneven fold sizes

**Solution**: This is normal for spatial methods; consider using stratified approaches if class imbalance is severe

### Issue: Geographic gaps in training data

**Solution**: Increase block size or number of clusters to improve coverage

## Next Steps

- Learn about [spatial leakage detection](spatial-leakage.html)
- Explore [model evaluation and comparison](model-evaluation.html)

## Key Takeaways

1. **Spatial block CV** is a good default method for most cases
2. **Buffered CV** provides precise distance control when needed
3. **Clustering CV** handles complex spatial patterns
4. **Always compare** with random CV to assess spatial dependence impact
5. **Visualize folds** to understand spatial separation