---
title: "Comparing Groups and Testing Hypotheses"
output: rmarkdown::html_vignette
vignette: >
  %\VignetteIndexEntry{Comparing Groups and Testing Hypotheses}
  %\VignetteEngine{knitr::rmarkdown}
  %\VignetteEncoding{UTF-8}
---

```{r, include = FALSE}
knitr::opts_chunk$set(
  collapse = TRUE,
  comment = "#>",
  message = FALSE,
  warning = FALSE
)
```

```{r setup}
library(mariposa)
library(dplyr)
data(survey_data)
```

## Overview

Statistical tests determine whether observed differences between groups are real or due to random chance. This guide covers all hypothesis tests in mariposa, organized by the type of data and research question.

### Choosing the Right Test

| Your data | 2 groups | 3+ groups | Paired |
|-----------|----------|-----------|--------|
| **Continuous, normal** | `t_test()` | `oneway_anova()` | *paired t-test planned* |
| **Continuous, non-normal** | `mann_whitney()` | `kruskal_wallis()` | `wilcoxon_test()` |
| **Categorical** | `chi_square()` | `chi_square()` | `mcnemar_test()` |
| **Small sample, categorical** | `fisher_test()` | --- | --- |
| **Multiple factors** | --- | `factorial_anova()` | `friedman_test()` |
| **With covariate** | --- | `ancova()` | --- |
| **Proportion vs. expected** | `binomial_test()` | `chisq_gof()` | --- |

### Checking Normality First

The parametric tests in the first row assume (approximately) normally
distributed variables. `normality_test()` produces the SPSS EXAMINE
"Tests of Normality" table — Kolmogorov-Smirnov with Lilliefors
correction and Shapiro-Wilk:

```{r}
survey_data %>%
  group_by(gender) %>%
  normality_test(age, income)
```

Run it grouped (as above) when you are about to compare groups: the
assumption that matters is normality *within* each group. With large
samples these tests flag even trivial deviations, so combine them with
the skewness and kurtosis from `describe()` before switching to a
non-parametric alternative.

## t-Tests

### Independent Samples

Compare two groups on a continuous variable:

```{r}
survey_data %>%
  t_test(life_satisfaction, group = gender, weights = sampling_weight)
```

The output includes both Student's t-test (equal variances assumed) and Welch's t-test (not assumed). When in doubt, use Welch --- it is more robust.

For the detailed output with group descriptives, Levene's test, and confidence intervals:

```{r}
survey_data %>%
  t_test(life_satisfaction, group = gender, weights = sampling_weight) %>%
  summary()
```

### Multiple Variables at Once

```{r}
survey_data %>%
  t_test(trust_government, trust_media, trust_science,
         group = gender, weights = sampling_weight)
```

### One-Sample t-Test

Test whether a mean differs from a specific value:

```{r}
# Is average life satisfaction different from the scale midpoint (3)?
survey_data %>%
  t_test(life_satisfaction, mu = 3, weights = sampling_weight)
```

### Grouped Analysis

Run separate tests per subgroup:

```{r}
survey_data %>%
  group_by(region) %>%
  t_test(income, group = gender, weights = sampling_weight)
```

## One-Way ANOVA

Compare means across three or more groups:

```{r}
result <- survey_data %>%
  oneway_anova(life_satisfaction, group = education, weights = sampling_weight)
result
```

The effect size $\eta^2$ (eta-squared) indicates how much variance is explained by group membership:

- **Small**: $\eta^2 \approx 0.01$
- **Medium**: $\eta^2 \approx 0.06$
- **Large**: $\eta^2 \approx 0.14$

### Post-Hoc Tests

A significant ANOVA tells you that groups differ, but not *which* groups. Use post-hoc tests:

```{r}
# Tukey HSD: balanced comparison of all pairs
tukey_test(result)
```

```{r}
# Scheffe: more conservative (fewer false positives)
scheffe_test(result)
```

### Assumption Check

ANOVA assumes equal variances. Test with Levene's test:

```{r}
levene_test(result)
```

If Levene's test is significant ($p < .05$), variances are unequal. Use the Welch correction included in the ANOVA output.

## Factorial ANOVA

Test the effects of two or more factors and their interactions:

```{r}
survey_data %>%
  factorial_anova(dv = income, between = c(gender, education),
                  weights = sampling_weight)
```

The output uses Type III sums of squares and reports partial $\eta^2$ for each effect. Weighted analysis uses WLS estimation, matching SPSS UNIANOVA.

For the full output with descriptive statistics per cell:

```{r}
survey_data %>%
  factorial_anova(dv = life_satisfaction, between = c(gender, region),
                  weights = sampling_weight) %>%
  summary()
```

## ANCOVA

Compare groups while controlling for a covariate:

```{r}
survey_data %>%
  ancova(dv = income, between = education, covariate = age,
         weights = sampling_weight)
```

The output includes the covariate effect, the adjusted factor effect, and estimated marginal means (group means adjusted for the covariate).

## Non-Parametric Tests

Use these when data is not normally distributed, ordinal, or based on small samples.

### Mann-Whitney U Test

The non-parametric alternative to the independent t-test:

```{r}
survey_data %>%
  mann_whitney(political_orientation, group = region,
               weights = sampling_weight)
```

### Kruskal-Wallis H Test

The non-parametric alternative to one-way ANOVA (3+ groups):

```{r}
kw_result <- survey_data %>%
  kruskal_wallis(life_satisfaction, group = education)

kw_result
```

When significant, use Dunn's post-hoc test with Bonferroni correction:

```{r}
dunn_test(kw_result)
```

### Wilcoxon Signed-Rank Test

The non-parametric alternative to the paired t-test:

```{r}
data(longitudinal_data_wide)

longitudinal_data_wide %>%
  wilcoxon_test(score_T1, score_T2)
```

### Friedman Test

The non-parametric alternative to repeated-measures ANOVA (3+ measurements):

```{r}
friedman_result <- longitudinal_data_wide %>%
  friedman_test(score_T1, score_T2, score_T3)

friedman_result
```

When significant, use pairwise Wilcoxon post-hoc tests:

```{r}
pairwise_wilcoxon(friedman_result)
```

### Binomial Test

Test whether an observed proportion differs from an expected value:

```{r}
survey_data %>%
  binomial_test(gender)
```

## Categorical Tests

### Chi-Square Test of Independence

Test whether two categorical variables are related:

```{r}
survey_data %>%
  chi_square(education, employment, weights = sampling_weight)
```

A significant result means the variables are not independent --- knowing one tells you something about the other.

### Effect Sizes for Categorical Data

The helpers `phi()`, `cramers_v()`, and `goodman_gamma()` run the
chi-square analysis internally and return just the requested effect size
as a number (per group for grouped data). For the full test output, call
`chi_square()` directly.

```{r}
# Phi coefficient (2x2 tables)
survey_data %>%
  phi(gender, employment, weights = sampling_weight)
```

```{r}
# Cramer's V (larger tables)
survey_data %>%
  cramers_v(education, employment, weights = sampling_weight)
```

### Fisher's Exact Test

Use when expected cell frequencies are below 5:

```{r}
small_sample <- survey_data %>% slice_sample(n = 30)

small_sample %>%
  fisher_test(gender, region)
```

### Chi-Square Goodness-of-Fit

Test whether observed frequencies match expected proportions:

```{r}
# Equal proportions (default)
survey_data %>%
  chisq_gof(education)
```

```{r}
# Custom expected proportions
survey_data %>%
  chisq_gof(education, expected = c(0.30, 0.25, 0.25, 0.20))
```

### McNemar's Test

Compare paired proportions (e.g., before/after):

```{r, eval = FALSE}
test_data <- survey_data %>%
  mutate(
    trust_gov_high = ifelse(trust_government > 3, 1, 0),
    trust_media_high = ifelse(trust_media > 3, 1, 0)
  )

test_data %>%
  mcnemar_test(var1 = trust_gov_high, var2 = trust_media_high)
```

## Interpreting Results

### p-Values

- $p < .05$: The difference is statistically significant
- $p \geq .05$: No significant difference detected

"Not significant" does **not** mean "no difference" --- it means we cannot rule out chance given the sample size.

### Effect Sizes

With large samples, even tiny differences can be significant. Always check effect sizes:

| Test | Effect size | Small | Medium | Large |
|------|-------------|-------|--------|-------|
| t-test | Cohen's *d* | 0.20 | 0.50 | 0.80 |
| ANOVA | $\eta^2$ | 0.01 | 0.06 | 0.14 |
| Chi-square | Cramer's *V* | 0.10 | 0.30 | 0.50 |
| Correlation | *r* | 0.10 | 0.30 | 0.50 |

### Multiple Comparisons

Running many tests inflates false positive rates. Post-hoc tests (`tukey_test()`, `dunn_test()`, `pairwise_wilcoxon()`) handle this automatically with corrections.

## Complete Example

A typical hypothesis testing workflow:

```{r}
# 1. Describe the groups
survey_data %>%
  group_by(education) %>%
  describe(life_satisfaction, weights = sampling_weight)

# 2. Test for overall differences
anova_result <- survey_data %>%
  oneway_anova(life_satisfaction, group = education,
               weights = sampling_weight)
anova_result

# 3. Check assumptions
levene_test(anova_result)

# 4. Post-hoc: which groups differ?
tukey_test(anova_result)
```

## Practical Tips

1. **Check assumptions first.** Use `describe(show = "all")` to inspect skewness. For non-normal data, use non-parametric tests.

2. **Match the test to the data.** Normal continuous data: t-test / ANOVA. Non-normal or ordinal: Mann-Whitney / Kruskal-Wallis. Categorical: chi-square / Fisher.

3. **Always follow up significant omnibus tests.** Use `tukey_test()` for ANOVA, `dunn_test()` for Kruskal-Wallis, `pairwise_wilcoxon()` for Friedman.

4. **Report effect sizes alongside p-values.** A significant result with a negligible effect size may not be practically meaningful.

5. **Use weights when available.** They ensure results represent the population, not just the sample.

## Summary

### Parametric Tests
- **`t_test()`** compares means between two groups
- **`oneway_anova()`** extends to three or more groups, with `tukey_test()` / `scheffe_test()` post-hoc
- **`factorial_anova()`** tests multiple factors and interactions
- **`ancova()`** controls for a covariate

### Non-Parametric Tests
- **`mann_whitney()`**, **`kruskal_wallis()`** (with `dunn_test()`), **`wilcoxon_test()`**, **`friedman_test()`** (with `pairwise_wilcoxon()`), **`binomial_test()`**

### Categorical Tests
- **`chi_square()`**, **`fisher_test()`**, **`chisq_gof()`**, **`mcnemar_test()`**

### Effect Sizes
- **`phi()`**, **`cramers_v()`**, **`goodman_gamma()`**

## Next Steps

- Measure relationships between continuous variables --- see `vignette("correlation-analysis")`
- Build predictive models --- see `vignette("regression-analysis")`
- Construct reliable scales --- see `vignette("scale-analysis")`
