Using OPCC

OPCC links Ontario postal codes to Statistics Canada 2021 census geographies through open, redistributable evidence.

Two access paths:

You want to… Go to
Look up postal codes or join census geographies to your own data Part 1 – uses pre-built artifacts downloaded automatically by the package; no source checkout needed
Verify artifacts or reproduce the DA roll-up from the DB artifact Part 2 – uses package functions only; no source checkout needed
Rebuild all artifacts from raw Statistics Canada sources Rebuild all artifacts – uses package functions for download and build; requires sf, dplyr, readr

Most code below is shown but not run when this vignette is built. Those chunks download versioned release artifacts, read files from paths you supply, or need packages listed under Suggests. Chunks that run with the installed package alone and no network are evaluated.

Part 1: Use OPCC with your data

Install

Requires R >= 4.1.

install.packages("OPCC")

Key concepts

OPCC links postal codes to two Statistics Canada census geographies:

Each postal code can map to multiple geographies. OPCC retains all defensible links with an allocation_weight (the share of observed address evidence for that link). The best_link flag marks the single strongest link per postal code. Unmatched postal codes are reported, never fabricated.

Normalize postal codes

All OPCC functions normalize internally, but you can normalize explicitly:

library(OPCC)

normalize_postal_code(c("m5v3a8", "K1A-0A6", "k1a0a6"))
#> [1] "M5V 3A8" "K1A 0A6" "K1A 0A6"

Look up a single postal code

The first call downloads and caches the artifact (~15 MB). Subsequent calls use the local cache.

# Best dissemination area (DA) link
pc_to_geo("M5V 3A8", level = "DA", all_links = FALSE)

# Every defensible DA link with allocation weights
pc_to_geo("M5V 3A8", level = "DA")

# Best dissemination block (DB) link
pc_to_geo("M5V 3A8", level = "DB", all_links = FALSE)

# Every DB link
pc_to_geo("M5V 3A8", level = "DB")

Unmatched postal codes are not silently dropped. They appear in the unmatched attribute of the result:

result <- pc_to_geo(c("M5V 3A8", "ZZZ 9Z9"), level = "DA")
attr(result, "unmatched")
#> [1] "ZZZ 9Z9"

Join OPCC geographies to your own data

The most common workflow: you have a data frame with a postal-code column and you want census geography identifiers. get_correspondence() and get_da_correspondence() download a release artifact, so the joins below are not evaluated here.

library(OPCC)

my_data <- data.frame(
  id = 1:5,
  postal_code = c("M5V 3A8", "K1A 0A6", "N6A 1B1", "P7A 1A1", "ZZZ 9Z9"),
  value = c(10, 20, 30, 40, 50),
  stringsAsFactors = FALSE
)

# --- DA join (one best DA per postal code) ---

da <- get_da_correspondence(vintage = "2026-07-20")
da_best <- da[da$best_link, c("postal_code", "DAUID", "allocation_weight")]

merged_da <- merge(
  my_data,
  da_best,
  by = "postal_code",
  all.x = TRUE
)
# Rows with unmatched postal codes keep NA for DAUID and allocation_weight.

# --- DB join (one best DB per postal code) ---

db <- get_correspondence(vintage = "2026-07-19-geonames-amendment")
db_best <- db[db$best_link, c("postal_code", "DBUID", "DAUID", "allocation_weight")]

merged_db <- merge(
  my_data,
  db_best,
  by = "postal_code",
  all.x = TRUE
)

# --- Keep all many-to-many links ---
# Skip the best_link filter to retain every candidate geography and its
# allocation weight. Your join will produce multiple rows per postal code.

merged_all <- merge(my_data, da, by = "postal_code", all.x = TRUE)

Columns returned by the join:

Column Meaning
postal_code Normalized A1A 1A1 postal code
DBUID 2021 Dissemination Block identifier (DB artifact)
DAUID 2021 Dissemination Area identifier (both artifacts)
allocation_weight Share of observed address evidence for this link; sums to 1 per postal code
best_link TRUE for the single strongest link per postal code
n_contributing_dbs Number of DBs contributing to a DA link (DA artifact)
contributing_dbuids Pipe-delimited DBUID list behind a DA link (DA artifact)
source_vintages Source evidence vintage(s) behind the link

With dplyr, which is a Suggests package and so is also not evaluated here:

library(dplyr)

da <- get_da_correspondence(vintage = "2026-07-20")

my_data |>
  mutate(postal_code = normalize_postal_code(postal_code)) |>
  left_join(
    da |> filter(best_link) |> select(postal_code, DAUID, allocation_weight),
    by = "postal_code"
  )

Download and cache artifacts

get_correspondence() and get_da_correspondence() download the versioned artifact, verify its SHA-256 checksum, cache it locally, and return a data frame. Subsequent calls use the cache.

By default, artifacts are cached in a session temporary directory, so OPCC never writes to your home filespace unless you ask it to. To keep artifacts across sessions, pass an explicit cache_dir, set the OPCC.cache_dir option or the OPCC_CACHE_DIR environment variable, or accept the one-time prompt that an interactive session offers. See ?clear_opcc_cache for the full resolution order.

Listing vintages reads an index shipped with the package, so it runs here:

list_vintages(level = "DB")
#> [1] "2026-06-26"                    "2026-07-19-geonames-amendment"
list_vintages(level = "DA")
#> [1] "2026-06-26" "2026-07-20"

Downloading a release does need the network, so the rest is shown but not run:

# Download using the default session cache
db <- get_correspondence(vintage = "2026-07-19-geonames-amendment")
da <- get_da_correspondence(vintage = "2026-07-20")

# Or use an explicit cache directory
cache <- file.path(tempdir(), "opcc-cache")
db <- get_correspondence(
  vintage = "2026-07-19-geonames-amendment",
  cache_dir = cache
)

# Pass a pre-loaded artifact to skip re-reading
pc_to_geo("M5V 3A8", level = "DA", correspondence = da)
pc_to_geo("M5V 3A8", level = "DB", correspondence = db)

Verify a release

# Download, checksum, and validate schema and invariants
validate_release(vintage = "2026-07-19-geonames-amendment", level = "DB",
                 cache_dir = cache)
validate_release(vintage = "2026-07-20", level = "DA", cache_dir = cache)

# Inspect provenance metadata
release_manifest(vintage = "2026-07-20", level = "DA", cache_dir = cache)

Offline and air-gapped use

After caching a verified release once, use offline = TRUE to require the local cache and block network access:

validate_release(vintage = "2026-07-20", level = "DA",
                 cache_dir = cache, offline = TRUE)
da <- get_da_correspondence(vintage = "2026-07-20",
                            cache_dir = cache, offline = TRUE)

You can also supply a local artifact file directly to pc_to_point():

pts <- pc_to_point("K0A 0A1",
                   point_file = "/path/to/opcc_m1_geonames_points.csv.gz")

Source-qualified point lookups

pc_to_point() returns supplementary GeoNames point observations. These are source-labeled and separate from NAR address evidence. It downloads its own point artifact, so the chunk is not evaluated:

pc_to_point("K0A 0A1")
pc_to_point("K0A 0A1", source = "geonames")

Import your own open evidence locally

Use this only for redistributable, open data. Local layers stay separate from canonical OPCC releases. The API rejects Canada Post, PCCF, and PCCF+ inputs. The paths below are placeholders for your own files, so the chunk is not evaluated.

my_source <- utils::read.csv("/path/to/municipal-postal-data.csv",
                             stringsAsFactors = FALSE)

Describing the source is pure metadata, so this part runs. A real checksum comes from your own file; a literal stands in for it here:

adapter <- new_source_adapter(
  source_id = "municipal_registry",
  licence = "Open Government Licence - Municipality",
  lineage = "Municipal open address registry",
  retrieval_date = "2026-07-20",
  schema_map = list(
    postal_code = "postal",
    latitude = "lat",
    longitude = "lon"
  ),
  checksum = strrep("0", 64L)
)
#> This source layer remains local and separate from canonical OPCC releases. If redistribution is permitted, submit its contribution bundle as an OPCC issue or pull request.

adapter
#> $source_id
#> [1] "municipal_registry"
#> 
#> $licence
#> [1] "Open Government Licence - Municipality"
#> 
#> $lineage
#> [1] "Municipal open address registry"
#> 
#> $retrieval_date
#> [1] "2026-07-20"
#> 
#> $schema_map
#> $schema_map$postal_code
#> [1] "postal"
#> 
#> $schema_map$latitude
#> [1] "lat"
#> 
#> $schema_map$longitude
#> [1] "lon"
#> 
#> 
#> $endpoint
#> NULL
#> 
#> $checksum
#> [1] "0000000000000000000000000000000000000000000000000000000000000000"
#> 
#> $location_type
#> [1] "unknown"
#> 
#> $coordinate_method
#> [1] "unknown"
#> 
#> $authority_level
#> [1] "unknown"
#> 
#> $coverage_type
#> [1] "unknown"
#> 
#> $update_frequency
#> [1] "unknown"
#> 
#> attr(,"class")
#> [1] "opcc_source_adapter"

Building the layer, profiling it, and packaging a contribution bundle all need your own data, so they are shown but not run:

layer <- build_source_layer(my_source, adapter, on_invalid = "quarantine")
profile_source_layer(layer)

# `output_dir` must be an explicit path you choose; nothing is written
# to the working directory by default.
bundle <- contribution_bundle(layer,
                              output_dir = file.path(tempdir(), "contributions"),
                              fixture_rows = 100L)
contribution_issue_url(bundle)

Part 2: Reproduce and verify artifacts

Every published artifact can be downloaded, verified, and reproduced using package functions alone. No source checkout is needed.

Verify published artifacts

Download each artifact, check its SHA-256 against the release index, and validate schema and correspondence invariants (unique keys, weight sums, one best link per postal code):

validate_release(vintage = "2026-07-19-geonames-amendment", level = "DB")
validate_release(vintage = "2026-07-20", level = "DA")

Inspect the full provenance metadata (source checksums, code version, build timestamp, row counts):

release_manifest(vintage = "2026-07-19-geonames-amendment", level = "DB")
release_manifest(vintage = "2026-07-20", level = "DA")

Reproduce the DA artifact from the DB artifact

The DA correspondence (M5) is a deterministic attribute roll-up of the DB correspondence (M2). aggregate_da_correspondence() performs the same computation as the M5 build script: it sums DB allocation weights within each DA, preserves contributing-DB lineage, and selects the best DA link.

The published DA artifacts were built from the NAR-only M2 baseline (2026-06-26), so use that vintage to reproduce them exactly. Both lookups download a release artifact, so the chunk is not evaluated:

db <- get_correspondence(vintage = "2026-06-26")
da_reproduced <- aggregate_da_correspondence(db)

# Compare with the published DA artifact
da_published <- get_da_correspondence(vintage = "2026-07-20")
stopifnot(identical(nrow(da_reproduced), nrow(da_published)))
stopifnot(identical(
  sort(unique(da_reproduced$postal_code)),
  sort(unique(da_published$postal_code))
))

The GeoNames amendment vintage (2026-07-19-geonames-amendment) adds 17,334 supplementary postal codes. Rolling it up with aggregate_da_correspondence() produces a valid superset of the published DA, but it will not match the published artifact row-for-row.

Reproduce offline

After caching artifacts once, all verification and reproduction works without network access:

cache <- file.path(tempdir(), "opcc-cache")

validate_release(vintage = "2026-06-26", level = "DB",
                 cache_dir = cache, offline = TRUE)
validate_release(vintage = "2026-07-20", level = "DA",
                 cache_dir = cache, offline = TRUE)

db <- get_correspondence(vintage = "2026-06-26",
                         cache_dir = cache, offline = TRUE)
da <- aggregate_da_correspondence(db)

Rebuild all artifacts from raw sources

The full pipeline from raw Statistics Canada inputs to verified artifacts uses package functions only. No source checkout is needed. Each step is a separate function call so you can inspect intermediate results.

The build requires sf (system libraries GDAL, GEOS, PROJ), dplyr, and readr. On Ubuntu: sudo apt-get install libgdal-dev libgeos-dev libproj-dev. On macOS: brew install gdal geos proj. On Windows: install Rtools and use the sf binary from CRAN.

None of the build steps below are evaluated when this vignette is built: they download the raw public inputs and need sf, dplyr, and readr. Each step also consumes the output of the step before it. Raw source files are large, so all download helpers require an explicit directory that you manage.

packages <- c("digest", "dplyr", "jsonlite", "readr", "sf")
missing <- packages[!vapply(packages, requireNamespace, logical(1),
                            quietly = TRUE)]
if (length(missing)) install.packages(missing)

Step 1: Download all public inputs. Each download is cached; re-running skips files already present.

build_cache <- file.path(tempdir(), "opcc-build")
dir.create(build_cache, recursive = TRUE, showWarnings = FALSE)

nar_dir <- download_nar(cache_dir = build_cache)
geonames_txt <- download_geonames(cache_dir = build_cache)
bounds <- download_census_boundaries(cache_dir = build_cache)
gaf_csv <- download_gaf(cache_dir = build_cache)

Step 2: Build postal code centroids (M1). Reads NAR address points and GeoNames, computes the mean coordinate per postal code. NAR centroids take priority; GeoNames fills gaps.

output_dir is required: OPCC never picks a destination for large build products on your behalf.

centroids_csv <- build_centroids(nar_dir, geonames_txt,
                                 output_dir = file.path(build_cache, "centroids"))

Step 3: Assign centroids to DBs and join GAF (M1). Validates Ontario boundary membership, assigns each centroid to one 2021 Dissemination Block via point-in-polygon, and joins the Geographic Attribute File for DAUID and higher geographies.

rollup_csv <- build_db_assignment(
  centroids_csv, bounds$province, bounds$db, gaf_csv,
  output_dir = file.path(build_cache, "rollup")
)

Step 4: Build the DB correspondence (M2). Reads raw NAR address points, assigns them to DBs, aggregates evidence per postal-code/DB pair, and appends GeoNames supplementary links.

m2_csv <- build_m2(nar_dir, bounds$db, gaf_csv, rollup_csv,
                   output_dir = file.path(build_cache, "m2"))

Step 5: Reproduce the DA roll-up (M5). This is a deterministic attribute roll-up of the M2 artifact. No spatial join or new input is needed.

db <- utils::read.csv(m2_csv, stringsAsFactors = FALSE)
da <- aggregate_da_correspondence(db)

Step 6: Verify.

stopifnot(!anyDuplicated(db[c("postal_code", "DBUID")]))
weights <- tapply(db$allocation_weight, db$postal_code, sum)
stopifnot(all(abs(weights - 1) < 1e-8))
stopifnot(all(tapply(db$best_link, db$postal_code, sum) == 1L))

Exact byte-level rebuild. To reproduce the published artifact bytes exactly, check out the producer revision recorded in the published manifest rather than the tip of main. See docs/m2-reproduction.md and docs/m5-reproduction.md for the pinned revisions, expected SHA-256 values, and full verification sequences.