---
title: "Understanding Relationships: Correlation Analysis"
output: rmarkdown::html_vignette
vignette: >
  %\VignetteIndexEntry{Understanding Relationships: Correlation Analysis}
  %\VignetteEngine{knitr::rmarkdown}
  %\VignetteEncoding{UTF-8}
---

```{r, include = FALSE}
knitr::opts_chunk$set(
  collapse = TRUE,
  comment = "#>",
  message = FALSE,
  warning = FALSE
)
```

```{r setup}
library(mariposa)
library(dplyr)
data(survey_data)
```

## Overview

Correlation measures how two variables move together. mariposa provides three methods for different situations:

| Method | Function | Best for |
|--------|----------|----------|
| Pearson's *r* | `pearson_cor()` | Linear relationships between continuous variables |
| Spearman's $\rho$ | `spearman_rho()` | Monotonic relationships, ordinal data, or data with outliers |
| Kendall's $\tau$ | `kendall_tau()` | Ordinal data, small samples, or many tied values |

All three support survey weights, multiple variables (correlation matrices), and grouped analysis.

## Pearson Correlation

### Basic Usage

```{r}
survey_data %>%
  pearson_cor(age, income)
```

Interpretation of *r*:

- $r = 1$: Perfect positive correlation
- $r = 0$: No linear relationship
- $r = -1$: Perfect negative correlation

### With Survey Weights

```{r}
survey_data %>%
  pearson_cor(age, income, weights = sampling_weight)
```

### Correlation Matrix

Pass multiple variables to get all pairwise correlations:

```{r}
survey_data %>%
  pearson_cor(trust_government, trust_media, trust_science,
              weights = sampling_weight)
```

### Detailed Output

```{r}
result <- survey_data %>%
  pearson_cor(age, income, weights = sampling_weight)

summary(result)
```

### Grouped Analysis

```{r}
survey_data %>%
  group_by(region) %>%
  pearson_cor(age, income, weights = sampling_weight)
```

## Spearman Correlation

Use Spearman's $\rho$ when the relationship is monotonic but not necessarily linear, or when working with ordinal data or data with outliers:

```{r}
survey_data %>%
  spearman_rho(political_orientation, environmental_concern,
               weights = sampling_weight)
```

### Multiple Variables

```{r}
survey_data %>%
  spearman_rho(political_orientation, environmental_concern,
               life_satisfaction, trust_government,
               weights = sampling_weight)
```

## Kendall's Tau

Use Kendall's $\tau$ for ordinal data, small samples ($n < 30$), or data with many tied values. It is more robust than Spearman but typically produces smaller absolute values:

```{r}
survey_data %>%
  kendall_tau(political_orientation, life_satisfaction,
              weights = sampling_weight)
```

## Partial Correlation

`partial_cor()` (SPSS `PARTIAL CORR`) removes the influence of control
variables from a correlation. Comparing the partial against the
zero-order correlation shows how much of the original association the
controls explain:

```{r}
# Does the satisfaction-income correlation survive controlling for age?
survey_data %>%
  partial_cor(life_satisfaction, income, controls = age,
              weights = sampling_weight)
```

Multiple analysis variables and multiple controls work together; with
three or more variables, `summary()` adds the partial-correlation
matrix:

```{r}
result <- survey_data %>%
  partial_cor(trust_government, trust_media, trust_science,
              controls = c(age, political_orientation))
summary(result)
```

## Interpreting Correlations

### Strength Guidelines

| Absolute value of *r* | Interpretation |
|------------------------|----------------|
| 0.0 -- 0.3 | Weak |
| 0.3 -- 0.7 | Moderate |
| 0.7 -- 1.0 | Strong |

These are general guidelines. In survey research, correlations of 0.2--0.4 between different constructs are common and often substantively meaningful.

### Choosing the Right Method

1. **Both variables continuous, relationship is linear** → Pearson
2. **Monotonic but non-linear relationship, or outliers present** → Spearman
3. **Ordinal data, small sample, or many ties** → Kendall

### Comparing Methods

If Pearson and Spearman give very different results, the relationship may be non-linear:

```{r}
pearson_result <- survey_data %>%
  pearson_cor(life_satisfaction, income, weights = sampling_weight)

spearman_result <- survey_data %>%
  spearman_rho(life_satisfaction, income, weights = sampling_weight)

kendall_result <- survey_data %>%
  kendall_tau(life_satisfaction, income, weights = sampling_weight)

comparison <- data.frame(
  Method = c("Pearson", "Spearman", "Kendall"),
  Correlation = c(pearson_result$correlations$correlation[1],
                  spearman_result$correlations$correlation[1],
                  kendall_result$correlations$correlation[1]),
  P_Value = c(pearson_result$correlations$p_value[1],
              spearman_result$correlations$p_value[1],
              kendall_result$correlations$p_value[1])
)
print(comparison)
```

### Correlation Does Not Imply Causation

A significant correlation does not mean one variable causes the other:

- A **third variable** may drive both (e.g., warm weather → both ice cream sales and drowning)
- The **direction of causality** is unknown
- There may be **no causal link** at all

## Complete Example

```{r}
# 1. Correlation matrix for key variables
cor_result <- survey_data %>%
  pearson_cor(age, income, life_satisfaction,
              political_orientation, environmental_concern,
              weights = sampling_weight)
cor_result

# 2. Identify the strongest correlations
significant <- cor_result$correlations %>%
  filter(p_value < 0.05) %>%
  arrange(desc(abs(correlation)))
print(significant)

# 3. Check if relationships vary by region
survey_data %>%
  group_by(region) %>%
  pearson_cor(age, income, weights = sampling_weight)
```

## Reporting Results (APA Style)

Include the correlation coefficient, confidence interval, p-value, and sample size:

> "There was a moderate positive correlation between age and income (*r* = .34, 95% CI [.29, .39], *p* < .001, *N*~eff~ = 2,341), indicating that older respondents tended to report higher incomes."

## Practical Tips

1. **Always check the scatterplot.** Correlations can miss non-linear relationships, and a single outlier can inflate or deflate the coefficient.

2. **Use Spearman when in doubt.** It makes fewer assumptions than Pearson and works with both continuous and ordinal data.

3. **Report the magnitude, not just significance.** With large samples, even $r = .05$ can be significant but is practically meaningless.

4. **Use correlation matrices to prioritize.** Before regression, check which variables are most strongly related to your outcome.

## Summary

1. **Pearson** measures linear relationships between continuous variables
2. **Spearman** measures monotonic relationships and works with ordinal data
3. **Kendall** is the most robust choice for ordinal data and small samples
4. All three support **weights**, **correlation matrices**, and **group_by()**
5. Correlation does **not** imply causation

## Next Steps

- Build predictive models --- see `vignette("regression-analysis")`
- Compare groups --- see `vignette("hypothesis-testing")`
- Construct reliable scales --- see `vignette("scale-analysis")`
