ZIP code is a convenient join key, but it is rarely the geography you actually want to analyze at. You usually want to roll observations up to county, state, metro area, or some custom region (sales territory, school district, study site). This post shows how to do that in R using the zipcodeR package.
There are two patterns, and they cover almost everything:
- Use built-in geographies. zipcodeR ships with a database that already maps every US ZIP to its state, county, county FIPS code, primary city, and more. A single join enriches your data.
- Use a custom crosswalk. When the region you care about is custom (sales territories, study clusters, custom service areas), maintain a separate
zipcode → regionlookup and join it on.
Setup
install.packages("zipcodeR")
library(zipcodeR)
library(dplyr)
For this post I will use a small example dataset of patient counts by ZIP:
patients <- tibble::tribble(
~zipcode, ~count,
"08901", 45,
"08902", 23,
"07001", 67,
"08540", 18,
"10001", 12,
"10002", 9,
)
1. Look up region attributes for a single ZIP
The workhorse function is reverse_zipcode():
reverse_zipcode("08901") %>%
select(zipcode, major_city, county, state, county_fips, population)
#> # A tibble: 1 × 6
#> zipcode major_city county state county_fips population
#> <chr> <chr> <chr> <chr> <chr> <int>
#> 1 08901 New Brunswick Middlesex NJ 34023 44537
The full result has 24 columns including timezone, area codes, lat/lng centroid, and ZIP type. For region rollups the columns you need are state, county, and county_fips. The county FIPS code is the canonical 5-digit identifier used by every Census, BLS, and BEA dataset — always prefer it over the county name for joins.
2. Enrich a data frame with state and county
To add region columns to your dataset, look up every ZIP at once and join the result back onto your data:
zip_info <- reverse_zipcode(patients$zipcode) %>%
select(zipcode, state, county, county_fips, msa = bounds_north) # rename as needed
patients_enriched <- patients %>%
left_join(zip_info, by = "zipcode")
patients_enriched
#> # A tibble: 6 × 6
#> zipcode count state county county_fips msa
#> <chr> <int> <chr> <chr> <chr> <dbl>
#> 1 08901 45 NJ Middlesex 34023 40.6
#> 2 08902 23 NJ Middlesex 34023 40.6
#> ...
reverse_zipcode() is vectorized — pass it the whole vector of ZIPs at once rather than calling it inside mutate() per row. The single batch call is dramatically faster.
3. Roll observations up to region
Once every row has a region, the rollup is standard dplyr:
patients_enriched %>%
group_by(state, county) %>%
summarize(total_patients = sum(count), .groups = "drop") %>%
arrange(desc(total_patients))
#> # A tibble: 4 × 3
#> state county total_patients
#> <chr> <chr> <int>
#> 1 NJ Middlesex 68
#> 2 NJ Union 67
#> 3 NJ Mercer 18
#> 4 NY New York 21
For state-level rollups, group only by state. For metro-level, you’ll need an MSA crosswalk (see below) — zipcodeR includes county FIPS, and the Census’s CBSA-to-county mapping handles the county→metro step.
4. Apply a custom region crosswalk
When your region scheme isn’t a built-in geography, the pattern is the same but you bring your own lookup table. Suppose you have sales territories defined as:
territories <- tibble::tribble(
~zipcode, ~territory,
"08901", "NJ-Central",
"08902", "NJ-Central",
"07001", "NJ-North",
"08540", "NJ-Central",
"10001", "NY-Manhattan",
"10002", "NY-Manhattan",
)
Join it onto your data and roll up the same way:
patients %>%
inner_join(territories, by = "zipcode") %>%
group_by(territory) %>%
summarize(total_patients = sum(count), .groups = "drop")
For larger custom region schemes (e.g. all 5,000+ NJ ZIPs assigned to school districts), build the crosswalk once, save it as a CSV or parquet, and load it into every analysis script.
A common gotcha: ZIP-to-region many-to-many
Some ZIP codes span more than one county or state. (The boundary between Camden County, NJ and Philadelphia, PA is the most famous example, but it happens elsewhere.) reverse_zipcode() returns the primary county/state for each ZIP — usually the one containing the post office — but for analytical work that crosses jurisdictional boundaries you may need a one-to-many crosswalk.
For that scenario, the canonical source is the HUD USPS ZIP Code Crosswalk Files, which provide quarterly snapshots of every ZIP’s distribution across counties, tracts, and CBSAs with proportions. Load that as a separate crosswalk and join with the proportional weights as needed.
Related work
- zipcodeR project page — full feature list, FAQ, and CRAN/GitHub links.
- How to calculate distance between ZIP codes in R — proximity calculations and radius searches.
- How to plot ZIP codes on a map in R with zipcodeR — turn your enriched data into choropleth maps.
- zipcodeR Software Impacts paper — the peer-reviewed publication.
zipcodeR is free and open source under the GPL-3 license. The CRAN package, full documentation, and source code are linked from the zipcodeR project page.