Stress-Testing Our Marathon Heat Number

We refit RunScore's heat cost eleven ways, dropped each of 92 races in turn, and found a real bug along the way. Here's what survived.

racecast·
WeatherHeatRunScore

When we rebuilt RunScore on 6.3 million marathon finishes, the number that carries the most weight in the whole model is one small coefficient: 0.35% slower per degree of WBGT above 50°F. Every heat cost on the site, every "how many minutes will this cost me" answer, traces back to it.

A single fitted number is easy to be wrong about and hard to know you're wrong about. So before we shipped it, we tried to break it: refit it eleven different ways, dropped each of the 92 races one at a time to see if any single race was propping it up, and rebuilt the whole pipeline after we found a genuine bug in our own results data. Here's what survived, what didn't, and the one chart we built and then chose not to publish.

Key takeaways
The headline number0.334% slower per °F of WBGT above 50°F, from race and year controls on 1,686 race-years. The shipped value is 0.35%.
Fit eleven waysDifferent controls, different subsets, different weighting. All eleven land between 0.31% and 0.37% per °F. Weighting each race day by its finishers, a cut outside the eleven, gives 0.40%.
Drop-one-race testRefit 92 times, leaving one race out each time. The slope never moves outside 0.321% to 0.340%.
A real bug, found and fixedSome results files stored gun time where chip time belonged. Fixing it moved 626 race-years by more than a minute, 140 by more than five.
The worst single fixNew York City 2025's median finish time moved 82 minutes faster once chip time replaced gun time.

Fitting the same number eleven ways

The headline heat estimate comes from one regression: log(finish time) against WBGT, with a control for each race and each year, clustered standard errors by race. If that's the only lens we look through, we're trusting a single set of choices, so we ran ten more versions, each changing one thing about the setup.

We tried the heat term alone, with no wind or rain terms sharing the equation. We swapped the per-year controls for a single straight-line trend. We gave 40 races their own baseline for each researched course era, in case a rerouted course was quietly doing the year control's job. We dropped 2020 through 2022, in case the pandemic's unusual fields were leaning on the estimate. We weighted races by how precise their median was, so a 30,000-runner field counts for more than a 200-runner one. We restricted to fields of at least 200, then 1,000. We split the original 13 races from the 79 added later. We isolated the 11 races outside the US, to check whether the slope is an American artifact. And we fit London alone, 25 years of one deep, fast field.

Eleven different ways of asking the same question, and every one that had enough races behind it to trust lands within six hundredths of a percentage point of the others. That's the kind of agreement that lets you ship a number.

The only real outlier is London on its own: 0.55%, with no usable confidence interval because a fit on one race has nowhere to draw uncertainty from. It's a reminder of what these ranges actually mean. A single race, however deep its field, isn't a sample size. It's an anecdote with a p-value attached.

The most demanding test isn't in that list of eleven. It's leave-one-race-out: refit the whole model 92 times, each time holding out one race entirely, and see how far the slope moves. Across all 92 refits, it stays inside 0.321% to 0.340%, a band under two hundredths of a point wide. No single race, not Boston, not Erie, not the smallest 200-runner field in the dataset, is carrying the result. If it were, dropping that race would show up as a spike, and nothing spikes.

The bug that almost gave us a real-looking finding

Partway through the rebuild we found something in our own results files that had nothing to do with statistics: some marathons' result exports stored gun time, the clock started at the first starting gun, in the field that should have held chip time, each runner's actual time from their own start mat to the finish. In a big wave start, where the last corral can cross the start line fifteen or twenty minutes after the gun, that difference compounds across thousands of runners.

It wasn't a small correction. Re-aggregating every race-year with the fix, taking the smaller of gun time and chip time wherever both were available, moved 626 race-years by more than a minute and 140 by more than five. The worst single case was New York City's 2025 marathon, whose median finish time moved 82 minutes faster once the fix was in, an extreme wave spread meeting a results file that had quietly been reporting gun time as if it were net time. Boston's 2021 COVID-era rolling start moved 64 minutes for a related reason, though that one turned out to be a genuine wave-start effect worth keeping, not a data error.

Before the fix, the data seemed to show something interesting: an extra penalty for hot, dry days on top of what WBGT alone predicted, as though dry heat were quietly worse than the humidity-adjusted number gave it credit for. It was a plausible story. It didn't survive. The bad files were concentrated in dry western races, San Francisco and Los Angeles chief among them, both with 17 years of results in the dataset. Once their finish times were corrected, the apparent dry-heat penalty shrank from a real-looking 1.6 percentage points to a statistically unremarkable 0.3 to 0.4, easily explained by noise. What had looked like a discovery about how heat works was actually a discovery about how our own results files were built. We'd rather find that out before publishing a chart about it than after.

What 13 races said, and what 92 said

Before the full rebuild, an early pass looked at just 13 marathons and found a heat cost of 0.43% per degree, noticeably steeper than what we ship today. It's worth walking through exactly how that number came down, because it happened in two separate steps that are easy to conflate into one.

The first step was rebuilding those same 13 races on the new pipeline. That changed several things at once, including 31 more race-years, a control for each year in place of a straight-line trend, and net times in place of gun times. Together they dropped the estimate to 0.34%, most of the way to where it would eventually land. The gun-versus-chip fix above was only part of that. Before the fix, the rebuilt 13 races fit 0.366%, so the bug accounts for about 0.02 of the 0.08-point drop, roughly a quarter. The second step was breadth: adding 79 more races, about 1,409 more race-years, moved the estimate only slightly further, to 0.334%, but it tightened the confidence interval considerably, from plus or minus about eight hundredths of a point down to about four hundredths. The original 13 races and the 79 added later now agree almost exactly, 0.344% versus 0.332%. The bug was part of the correction, not most of it. Most of the remaining value of the bigger sample was precision, not a different answer.

The chart we built and then didn't publish

One hypothesis felt intuitive enough that we wanted it to be true: heat above what's normal for a given race and season should cost more than the same WBGT reading on a day nobody was surprised by, because runners training through a typically mild spring haven't acclimated to a heat spike the way runners training through a typically hot summer have. It's a reasonable idea, and it has some support in the fitted model: a second heat term, measuring degrees above each race's own long-run normal, comes back at 0.113% per degree (standard error 0.057), on top of the 0.25% baseline rate for normal-for-the-race heat. That's roughly one and a half times the per-degree cost when the heat is a surprise.

But a fitted term with a standard error that size is a hint, not a chart. Before we built a visual around it, we looked directly at the two race-days it should explain best. Boston 2012 ran about 21°F above the race's normal WBGT, deep into unusual-heat territory, and the model missed that day's real slowdown by 5.5 percentage points, one of the dozen largest underestimates in the dataset and the largest on any day above 65°F WBGT. NYC 2022 ran almost as far above its own normal, about 20°F, and the model missed it by six tenths of a point, one of its more accurate hot-day estimates. Two races with nearly identical "how much hotter than usual" numbers, and one produced one of the model's biggest misses while the other was nearly exact.

Whatever explains Boston 2012, and something clearly does, it isn't simply distance above the race's seasonal normal. Plot the model's heat-only errors on hot days against each race's degrees above normal, and the line is flat. We built that chart. It didn't show a story, so we didn't ship it, and the acclimatization term stays out of the model rather than riding along on a scatter plot that would have implied more than the data actually says.

Why small samples mislead

Every thread here points at the same failure mode. The early sample was two majors, Boston and Chicago, plus eleven other races such as Carmel, Glass City and Lost Dutchman, eight of the 13 run in spring, 246 race-years in all. It also carried more than its share of the bad results files: 31 of the 141 race-years our sweep flagged for gun time, mostly Chicago before 2010, Boston before 2006, Oklahoma City and Pittsburgh. With that few race-years, a handful of bad files and a few pipeline choices could move the heat estimate by nearly a fifth, and nothing inside the 13 races would show it. The bug wasn't only theirs. San Francisco and Los Angeles, added later, had 17 bad years each, and the clearest case, New York's 82-minute gap, came from the bigger sample.

The honest version of "how much does heat cost a marathoner" isn't a single number we're confident in. It's a number, 0.334%, that moves by no more than about six hundredths of a point across the eleven ways we asked the question above, and doesn't depend on any one race being in the dataset. At a WBGT of 68°F, roughly what NYC 2022 averaged, the low and high ends of that range work out to about two and a half minutes apart for a four-hour marathoner. That's what the uncertainty in this number actually looks like in practice: real, worth stating honestly, and not the kind of gap that would change how anyone should plan a race.

We ship 0.35% per degree, a touch above our own 0.334% headline fit and inside the 0.31% to 0.37% range the eleven versions above span. Two cuts outside that list come out higher: weighting each race day by its number of finishers gives 0.40%, and the nine races whose median field is 5,000 runners or more fit about 0.42%. So the biggest marathons may cost a little more per degree than 0.35% says. It's the number in RunScore's heat cost today, and it's the one we spent the most time trying to break.

More guides

Boston Marathon
Check the current forecast, RunScore, and heat-adjusted pace for Boston
Stress-Testing Our Marathon Heat Number | racecast.io