ZIP code is a convenient join key, but it is rarely the geography you actually want to analyze at. You usually want to roll observations up to county, state, metro area, or some custom region (sales territory, school district, study site). This post shows how to do that in R using the zipcodeR package.

There are two patterns, and they cover almost everything:

  1. Use built-in geographies. zipcodeR ships with a database that already maps every US ZIP to its state, county, county FIPS code, primary city, and more. A single join enriches your data.
  2. Use a custom crosswalk. When the region you care about is custom (sales territories, study clusters, custom service areas), maintain a separate zipcode → region lookup and join it on.

Setup

install.packages("zipcodeR")
library(zipcodeR)
library(dplyr)

For this post I will use a small example dataset of patient counts by ZIP:

patients <- tibble::tribble(
  ~zipcode, ~count,
  "08901",      45,
  "08902",      23,
  "07001",      67,
  "08540",      18,
  "10001",      12,
  "10002",       9,
)

1. Look up region attributes for a single ZIP

The workhorse function is reverse_zipcode():

reverse_zipcode("08901") %>%
  select(zipcode, major_city, county, state, county_fips, population)
#> # A tibble: 1 × 6
#>   zipcode major_city    county      state county_fips population
#>   <chr>   <chr>         <chr>       <chr> <chr>            <int>
#> 1 08901   New Brunswick Middlesex   NJ    34023            44537

The full result has 24 columns including timezone, area codes, lat/lng centroid, and ZIP type. For region rollups the columns you need are state, county, and county_fips. The county FIPS code is the canonical 5-digit identifier used by every Census, BLS, and BEA dataset — always prefer it over the county name for joins.

2. Enrich a data frame with state and county

To add region columns to your dataset, look up every ZIP at once and join the result back onto your data:

zip_info <- reverse_zipcode(patients$zipcode) %>%
  select(zipcode, state, county, county_fips, msa = bounds_north)  # rename as needed

patients_enriched <- patients %>%
  left_join(zip_info, by = "zipcode")

patients_enriched
#> # A tibble: 6 × 6
#>   zipcode count state county      county_fips     msa
#>   <chr>   <int> <chr> <chr>       <chr>         <dbl>
#> 1 08901      45 NJ    Middlesex   34023          40.6
#> 2 08902      23 NJ    Middlesex   34023          40.6
#> ...

reverse_zipcode() is vectorized — pass it the whole vector of ZIPs at once rather than calling it inside mutate() per row. The single batch call is dramatically faster.

3. Roll observations up to region

Once every row has a region, the rollup is standard dplyr:

patients_enriched %>%
  group_by(state, county) %>%
  summarize(total_patients = sum(count), .groups = "drop") %>%
  arrange(desc(total_patients))
#> # A tibble: 4 × 3
#>   state county      total_patients
#>   <chr> <chr>                <int>
#> 1 NJ    Middlesex               68
#> 2 NJ    Union                   67
#> 3 NJ    Mercer                  18
#> 4 NY    New York                21

For state-level rollups, group only by state. For metro-level, you’ll need an MSA crosswalk (see below) — zipcodeR includes county FIPS, and the Census’s CBSA-to-county mapping handles the county→metro step.

4. Apply a custom region crosswalk

When your region scheme isn’t a built-in geography, the pattern is the same but you bring your own lookup table. Suppose you have sales territories defined as:

territories <- tibble::tribble(
  ~zipcode, ~territory,
  "08901",  "NJ-Central",
  "08902",  "NJ-Central",
  "07001",  "NJ-North",
  "08540",  "NJ-Central",
  "10001",  "NY-Manhattan",
  "10002",  "NY-Manhattan",
)

Join it onto your data and roll up the same way:

patients %>%
  inner_join(territories, by = "zipcode") %>%
  group_by(territory) %>%
  summarize(total_patients = sum(count), .groups = "drop")

For larger custom region schemes (e.g. all 5,000+ NJ ZIPs assigned to school districts), build the crosswalk once, save it as a CSV or parquet, and load it into every analysis script.

A common gotcha: ZIP-to-region many-to-many

Some ZIP codes span more than one county or state. (The boundary between Camden County, NJ and Philadelphia, PA is the most famous example, but it happens elsewhere.) reverse_zipcode() returns the primary county/state for each ZIP — usually the one containing the post office — but for analytical work that crosses jurisdictional boundaries you may need a one-to-many crosswalk.

For that scenario, the canonical source is the HUD USPS ZIP Code Crosswalk Files, which provide quarterly snapshots of every ZIP’s distribution across counties, tracts, and CBSAs with proportions. Load that as a separate crosswalk and join with the proportional weights as needed.

zipcodeR is free and open source under the GPL-3 license. The CRAN package, full documentation, and source code are linked from the zipcodeR project page.

About Gavin Rozzi

Gavin Rozzi

Gavin Rozzi

I lead an in-house civic technology team in New Jersey state government. I write here about data science, software built for government, and how policy gets turned into working systems.