---
title: "Scale Analysis"
output: rmarkdown::html_vignette
vignette: >
  %\VignetteIndexEntry{Scale Analysis}
  %\VignetteEngine{knitr::rmarkdown}
  %\VignetteEncoding{UTF-8}
---

```{r setup, include = FALSE}
knitr::opts_chunk$set(
  collapse = TRUE,
  comment = "#>",
  message = FALSE,
  warning = FALSE
)
```

```{r}
library(mariposa)
library(dplyr)
data(survey_data)
```

## Overview

Scale analysis combines multiple survey items into a single score that reliably measures a concept. The typical workflow is:

1. **Check reliability** with `reliability()` --- do the items measure the same thing?
2. **Explore structure** with `efa()` --- how do items group into dimensions?
3. **Create scores** with `row_means()` --- compute mean indices
4. **Standardize** with `pomps()` --- transform to a comparable 0--100 scale

This guide uses the trust items from `survey_data` (trust in government, media, and science).

| Function | Purpose |
|----------|---------|
| `reliability()` | Internal consistency (Cronbach's Alpha, McDonald's Omega) |
| `efa()` | Discover underlying dimensions (factor analysis) |
| `row_means()` | Row-wise mean indices |
| `row_sums()` | Row-wise sums |
| `row_count()` | Count specific values per row |
| `pomps()` | Percent of Maximum Possible Scores (0--100) |

## Reliability Analysis

### Basic Usage

```{r}
reliability(survey_data, trust_government, trust_media, trust_science)
```

### Detailed Output

```{r}
rel <- reliability(survey_data, trust_government, trust_media, trust_science)
summary(rel)
```

The detailed output shows item-total correlations, alpha-if-item-deleted, and inter-item correlations --- matching SPSS RELIABILITY.

### Interpreting Cronbach's Alpha

- Alpha > 0.90: Excellent
- 0.80 -- 0.90: Good
- 0.70 -- 0.80: Acceptable
- 0.60 -- 0.70: Questionable
- Below 0.60: Reconsider your items

**Item-Total Correlation** shows how well each item fits the scale. Values above 0.40 indicate good fit; below 0.20 suggests the item does not belong.

**Alpha if Item Deleted** shows what happens without each item. If alpha increases when you remove an item, that item weakens the scale.

### McDonald's Omega

Alongside alpha, `reliability()` reports **McDonald's Omega**, a factor-model-based reliability coefficient. Alpha assumes all items measure the construct equally well; omega fits a one-factor model and lets each item carry its own loading, which usually makes it the more accurate estimate when item loadings differ. The same thresholds as for alpha are commonly applied. Omega needs at least 3 items (with 2 items it is reported as `NA`), and **Omega if Item Deleted** appears in the item-total table for scales of 4 or more items.

Note: omega is currently an R-only statistic (no SPSS reference run yet); see `vignette("spss-compatibility")` for its validation status.

### With Survey Weights

```{r}
reliability(survey_data, trust_government, trust_media, trust_science,
            weights = sampling_weight)
```

### Using tidyselect

```{r}
reliability(survey_data, starts_with("trust"))
```

### Grouped Analysis

Check whether reliability holds across subgroups:

```{r}
survey_data %>%
  group_by(region) %>%
  reliability(trust_government, trust_media, trust_science)
```

A scale that works well overall might be unreliable in specific subgroups. Always check when your sample spans diverse populations.

## Exploratory Factor Analysis

### Basic Usage

When you have many items, `efa()` reveals how they group into underlying dimensions:

```{r}
efa(survey_data,
    political_orientation, environmental_concern, life_satisfaction,
    trust_government, trust_media, trust_science)
```

### Detailed Output

```{r}
efa_result <- efa(survey_data,
    political_orientation, environmental_concern, life_satisfaction,
    trust_government, trust_media, trust_science)

summary(efa_result)
```

### Understanding the Output

**KMO (Kaiser-Meyer-Olkin)** measures sampling adequacy for factor analysis:

- Above 0.80: Good to excellent
- 0.60 -- 0.80: Acceptable
- Below 0.60: Factor analysis may not be appropriate

**Bartlett's Test** should be significant ($p < .05$), confirming that meaningful correlations exist.

**Eigenvalues** show variance explained per component. By default, components with eigenvalue > 1 are retained (Kaiser criterion).

**Factor Loadings** show item-component associations:

- Above 0.70: Strong
- 0.40 -- 0.70: Moderate
- Below 0.40: Suppressed by default

### Rotation Methods

**Varimax** (default) --- assumes uncorrelated factors:

```{r}
efa(survey_data,
    political_orientation, environmental_concern, life_satisfaction,
    trust_government, trust_media, trust_science,
    rotation = "varimax")
```

**Promax** --- allows correlated factors, produces Pattern and Structure matrices:

```{r}
efa(survey_data,
    political_orientation, environmental_concern, life_satisfaction,
    trust_government, trust_media, trust_science,
    rotation = "promax")
```

**Oblimin** --- another oblique rotation, common in psychology:

```{r, eval = FALSE}
# Requires GPArotation package
efa(survey_data,
    political_orientation, environmental_concern, life_satisfaction,
    trust_government, trust_media, trust_science,
    rotation = "oblimin")
```

### Extraction Methods

By default, `efa()` uses PCA (Principal Component Analysis). For a true factor analysis model, use Maximum Likelihood:

```{r}
efa(survey_data,
    political_orientation, environmental_concern, life_satisfaction,
    trust_government, trust_media, trust_science,
    extraction = "ml")
```

ML extraction provides a goodness-of-fit test and uses SMC (squared multiple correlations) as initial communalities. Combine any extraction with any rotation:

```{r}
efa(survey_data,
    political_orientation, environmental_concern, life_satisfaction,
    trust_government, trust_media, trust_science,
    extraction = "ml", rotation = "promax")
```

### Fixing the Number of Factors

```{r}
efa(survey_data,
    political_orientation, environmental_concern, life_satisfaction,
    trust_government, trust_media, trust_science,
    n_factors = 2)
```

### With Survey Weights

```{r}
efa(survey_data,
    political_orientation, environmental_concern, life_satisfaction,
    trust_government, trust_media, trust_science,
    weights = sampling_weight)
```

## Creating Scale Scores

After confirming reliability, create scores using the row operation functions. For details on `row_means()`, `row_sums()`, `row_count()`, and `pomps()`, see `vignette("data-transformation")`.

### Quick Scale Construction

```{r}
# Create mean index
survey_data <- survey_data %>%
  mutate(m_trust = row_means(., trust_government, trust_media, trust_science,
                             min_valid = 2))

# Transform to 0-100 scale
survey_data <- survey_data %>%
  mutate(trust_pomps = pomps(m_trust, scale_min = 1, scale_max = 5))

# Check the result
survey_data %>%
  describe(m_trust, trust_pomps)
```

### Using the Scale in Analysis

```{r}
# Group comparison
survey_data %>%
  t_test(m_trust, group = gender, weights = sampling_weight)
```

```{r}
# As a predictor in regression
survey_data %>%
  linear_regression(life_satisfaction ~ m_trust + age + income,
                    weights = sampling_weight)
```

## Complete Example

```{r}
# 1. Check reliability
rel <- reliability(survey_data, trust_government, trust_media, trust_science)
rel
summary(rel)

# 2. Explore factor structure
efa_result <- efa(survey_data, trust_government, trust_media, trust_science)
efa_result

# 3. Create mean index (Alpha was acceptable)
survey_data <- survey_data %>%
  mutate(m_trust = row_means(., trust_government, trust_media, trust_science,
                             min_valid = 2))

# 4. Transform to POMPS
survey_data <- survey_data %>%
  mutate(trust_pomps = pomps(m_trust, scale_min = 1, scale_max = 5))

# 5. Use in further analysis
survey_data %>%
  group_by(education) %>%
  describe(m_trust, trust_pomps, weights = sampling_weight)
```

## Practical Tips

1. **Always check reliability first.** A mean index from unreliable items produces meaningless results. Aim for Alpha > .70.

2. **Use `min_valid` wisely.** For 3--5 item scales, `min_valid = 2` is a reasonable compromise. For longer scales, require at least half the items.

3. **Specify theoretical `scale_min` / `scale_max` in `pomps()`.** Using observed values makes scores sample-dependent and non-comparable.

4. **Run EFA when you have 6+ items.** Factor analysis can reveal unexpected dimensions before you create indices.

5. **Check grouped reliability.** A scale that works well overall may be unreliable in specific subgroups (e.g., different regions or education levels).

## Summary

1. `reliability()` checks whether items form a consistent scale (Cronbach's Alpha)
2. `efa()` discovers underlying dimensions (PCA or ML extraction, Varimax/Oblimin/Promax rotation)
3. `row_means()` creates mean indices; `row_sums()` and `row_count()` provide alternatives
4. `pomps()` transforms scores to a comparable 0--100 scale
5. Always validate reliability **before** creating scale scores

## Next Steps

- Learn about row operations and transformations --- see `vignette("data-transformation")`
- Use scales in regression models --- see `vignette("regression-analysis")`
- Compare scale scores across groups --- see `vignette("hypothesis-testing")`
- Apply survey weights --- see `vignette("survey-weights")`
