If you work with US data at the ZIP code level — patients, customers, properties, voters, anything — sooner or later you will need to compute distance between two ZIPs. Maybe you want to flag patients who live within 25 miles of a clinic, or build a dataset of stores within a delivery radius. The zipcodeR R package handles this with a single function and no external API calls.

This post covers the three patterns I use most often:

  1. Pairwise distance between two specific ZIPs.
  2. Batch distance calculations against a vector of ZIPs.
  3. Finding every ZIP within a radius of a target ZIP.

Setup

Install zipcodeR from CRAN:

install.packages("zipcodeR")
library(zipcodeR)

zipcodeR ships with a built-in SQLite database of every US ZIP code, including its centroid latitude and longitude. That means distance calculations work offline and at full speed — there is no API quota and no network call. Verify the install with a quick lookup:

reverse_zipcode("08901")
#> # A tibble: 1 × 24
#>   zipcode zipcode_type major_city county   state lat   lng   ...
#>   <chr>   <chr>        <chr>      <chr>    <chr> <dbl> <dbl>
#> 1 08901   Standard     New Brunsw Middlese NJ    40.49 -74.4

1. Distance between two ZIPs

The zip_distance() function returns the great-circle distance between two ZIP centroids:

zip_distance("08901", "10001")
#> [1] 28.42  # miles

By default the unit is miles. Pass units = "km" for kilometers. The calculation uses the haversine formula on the centroid coordinates, which is the right approximation for any distance under a few hundred miles — accurate to within a fraction of a percent for normal use.

A few edge cases to know about:

  • Invalid or non-existent ZIPs return NA rather than throwing an error. This is usually what you want, especially in batch mode.
  • PO Box ZIPs (zipcode_type == "PO Box") have centroid coordinates assigned to the post office handling that box. Distances are still meaningful but reflect the building, not a residential address.
  • Military ZIPs (zipcode_type == "Military") and unique ZIPs (zipcode_type == "Unique", e.g. large corporate campuses) work the same way.

2. Batch pairwise distances

zip_distance() is vectorized. If you pass two vectors of equal length it returns a vector of pairwise distances:

origins <- c("08901", "10001", "07001", "08540")
destinations <- c("10001", "10001", "10001", "10001")

zip_distance(origins, destinations)
#> [1] 28.42  0.00 11.69 49.83

This pattern is the right approach when you have a data frame of (origin, destination) pairs — for example, patients and clinics, customers and stores. Add the result as a new column:

library(dplyr)

patient_clinic_pairs %>%
  mutate(distance_miles = zip_distance(patient_zip, clinic_zip))

For an all-pairs (cross join) distance matrix, use outer() or tidyr::expand_grid() first:

library(tidyr)

zip_pairs <- expand_grid(
  origin = unique(patients$zipcode),
  dest = unique(clinics$zipcode)
)

zip_pairs <- zip_pairs %>%
  mutate(distance_miles = zip_distance(origin, dest))

This is fast — even 100K row distance lookups complete in seconds because everything happens in-memory after the initial database query.

3. Finding every ZIP within a radius

For “show me all ZIPs within 25 miles of this one” the function is search_radius():

nearby <- search_radius("08901", radius = 25)
head(nearby)
#> # A tibble: 6 × 2
#>   zipcode distance
#>   <chr>      <dbl>
#> 1 08901       0   
#> 2 08902       1.4
#> 3 08816       2.1
#> 4 08882       2.8
#> 5 08854       3.3
#> 6 08810       3.7

The result is a tibble with two columns: the ZIP code and its distance from the input ZIP. The radius is in miles by default; pass units = "km" for kilometers. The input ZIP itself is included in the result with distance == 0 — drop it with filter(distance > 0) if you only want neighbors.

Putting it together: filter a dataset by proximity

A common end-to-end task is “give me every observation within N miles of a target ZIP.” search_radius() plus a join handles this in three lines:

library(dplyr)

target_zip <- "08901"
radius_miles <- 25

nearby_zips <- search_radius(target_zip, radius = radius_miles)$zipcode

my_data %>%
  filter(zipcode %in% nearby_zips)

If you need the actual distance for each row in the filtered result (e.g. for distance-weighted analysis), do the join the other way:

my_data %>%
  inner_join(search_radius(target_zip, radius = radius_miles), by = "zipcode") %>%
  arrange(distance)

Where the data comes from

The ZIP centroid coordinates in zipcodeR are sourced from the U.S. Census Bureau’s ZIP Code Tabulation Areas (ZCTAs), which are the Census Bureau’s approximation of USPS ZIP delivery areas. ZCTAs and USPS ZIPs are not the same thing — USPS ZIPs are operational route designations, while ZCTAs are areal polygons that aggregate Census blocks. For statistical analysis (most academic and policy work) ZCTAs are what you actually want. For mailing-list applications you may need a separate USPS data source.

zipcodeR is free and open source under the GPL-3 license. If you use it in research, please cite the Software Impacts paper.

About Gavin Rozzi

Gavin Rozzi

Gavin Rozzi

I lead an in-house civic technology team in New Jersey state government. I write here about data science, software built for government, and how policy gets turned into working systems.