If you work with US data at the ZIP code level — patients, customers, properties, voters, anything — sooner or later you will need to compute distance between two ZIPs. Maybe you want to flag patients who live within 25 miles of a clinic, or build a dataset of stores within a delivery radius. The zipcodeR R package handles this with a single function and no external API calls.
This post covers the three patterns I use most often:
- Pairwise distance between two specific ZIPs.
- Batch distance calculations against a vector of ZIPs.
- Finding every ZIP within a radius of a target ZIP.
Setup
Install zipcodeR from CRAN:
install.packages("zipcodeR")
library(zipcodeR)
zipcodeR ships with a built-in SQLite database of every US ZIP code, including its centroid latitude and longitude. That means distance calculations work offline and at full speed — there is no API quota and no network call. Verify the install with a quick lookup:
reverse_zipcode("08901")
#> # A tibble: 1 × 24
#> zipcode zipcode_type major_city county state lat lng ...
#> <chr> <chr> <chr> <chr> <chr> <dbl> <dbl>
#> 1 08901 Standard New Brunsw Middlese NJ 40.49 -74.4
1. Distance between two ZIPs
The zip_distance() function returns the great-circle distance between two ZIP centroids:
zip_distance("08901", "10001")
#> [1] 28.42 # miles
By default the unit is miles. Pass units = "km" for kilometers. The calculation uses the haversine formula on the centroid coordinates, which is the right approximation for any distance under a few hundred miles — accurate to within a fraction of a percent for normal use.
A few edge cases to know about:
- Invalid or non-existent ZIPs return
NArather than throwing an error. This is usually what you want, especially in batch mode. - PO Box ZIPs (
zipcode_type == "PO Box") have centroid coordinates assigned to the post office handling that box. Distances are still meaningful but reflect the building, not a residential address. - Military ZIPs (
zipcode_type == "Military") and unique ZIPs (zipcode_type == "Unique", e.g. large corporate campuses) work the same way.
2. Batch pairwise distances
zip_distance() is vectorized. If you pass two vectors of equal length it returns a vector of pairwise distances:
origins <- c("08901", "10001", "07001", "08540")
destinations <- c("10001", "10001", "10001", "10001")
zip_distance(origins, destinations)
#> [1] 28.42 0.00 11.69 49.83
This pattern is the right approach when you have a data frame of (origin, destination) pairs — for example, patients and clinics, customers and stores. Add the result as a new column:
library(dplyr)
patient_clinic_pairs %>%
mutate(distance_miles = zip_distance(patient_zip, clinic_zip))
For an all-pairs (cross join) distance matrix, use outer() or tidyr::expand_grid() first:
library(tidyr)
zip_pairs <- expand_grid(
origin = unique(patients$zipcode),
dest = unique(clinics$zipcode)
)
zip_pairs <- zip_pairs %>%
mutate(distance_miles = zip_distance(origin, dest))
This is fast — even 100K row distance lookups complete in seconds because everything happens in-memory after the initial database query.
3. Finding every ZIP within a radius
For “show me all ZIPs within 25 miles of this one” the function is search_radius():
nearby <- search_radius("08901", radius = 25)
head(nearby)
#> # A tibble: 6 × 2
#> zipcode distance
#> <chr> <dbl>
#> 1 08901 0
#> 2 08902 1.4
#> 3 08816 2.1
#> 4 08882 2.8
#> 5 08854 3.3
#> 6 08810 3.7
The result is a tibble with two columns: the ZIP code and its distance from the input ZIP. The radius is in miles by default; pass units = "km" for kilometers. The input ZIP itself is included in the result with distance == 0 — drop it with filter(distance > 0) if you only want neighbors.
Putting it together: filter a dataset by proximity
A common end-to-end task is “give me every observation within N miles of a target ZIP.” search_radius() plus a join handles this in three lines:
library(dplyr)
target_zip <- "08901"
radius_miles <- 25
nearby_zips <- search_radius(target_zip, radius = radius_miles)$zipcode
my_data %>%
filter(zipcode %in% nearby_zips)
If you need the actual distance for each row in the filtered result (e.g. for distance-weighted analysis), do the join the other way:
my_data %>%
inner_join(search_radius(target_zip, radius = radius_miles), by = "zipcode") %>%
arrange(distance)
Where the data comes from
The ZIP centroid coordinates in zipcodeR are sourced from the U.S. Census Bureau’s ZIP Code Tabulation Areas (ZCTAs), which are the Census Bureau’s approximation of USPS ZIP delivery areas. ZCTAs and USPS ZIPs are not the same thing — USPS ZIPs are operational route designations, while ZCTAs are areal polygons that aggregate Census blocks. For statistical analysis (most academic and policy work) ZCTAs are what you actually want. For mailing-list applications you may need a separate USPS data source.
Related work
- zipcodeR project page — full feature list, FAQ, and CRAN/GitHub links.
- How to plot ZIP codes on a map in R with zipcodeR — the natural follow-on to this post.
- How to assign ZIP codes to geographic regions in R — when you need to roll ZIPs up to county, state, or MSA.
- zipcodeR Software Impacts paper — the peer-reviewed publication.
zipcodeR is free and open source under the GPL-3 license. If you use it in research, please cite the Software Impacts paper.