Why We Rebuilt RunScore From 6.3 Million Marathon Finishes
The old RunScore learned from training runs. The new one is calibrated on 6.3 million marathon finishes. Why we switched, how it works, and where it's off.
The Erie Marathon started at 7 AM on September 13 with temperatures already in the low 70s. The median finisher ran 4:01:26, about 19 minutes slower than in Erie's recent cool years. In the days before the race, our Erie page rated that forecast "Good." Our own pace model, on the same page, said the weather would cost a 10:00-per-mile runner about 16 minutes.
Both numbers came from us. Only one was close.
Erie wasn't a one-off. Run the old score on the 2012 Boston Marathon, when temperatures reached the high 80s and the BAA let 2,160 runners defer to the next year, and it says Good. Same for 2022, the warmest NYC Marathon on record. So we rebuilt RunScore. Race-day scores now come straight out of the pace-cost model behind our pace adjustments, calibrated against 6.3 million marathon finish times. Here's how we did it, what the data said, and where the new score is still off.
Why training runs were the wrong data
The old RunScore was built from more than 350,000 logged runs, along with runners' ratings of how the weather felt. That's a sensible way to answer "is this nice running weather?" It's the wrong data for the question a race page is asking, which is how much slower this weather will make you on race day.
Training runs are mostly easy efforts at a time of day the runner picked, and a rating of how the weather felt measures comfort. A marathon is three to five hours of hard running from a start time someone else chose. Heat that's merely unpleasant on an easy hour keeps adding up over a race. If the question is race-day slowdown, the data that answers it is race-day finish times.
There was a second problem. The site had two opinions about every forecast. The score sat next to a pace model built from published research: WBGT heat stress (Ely et al., 2007), aerodynamic drag for wind, and Minetti's grade costs for hills. The two disagreed in ways that weren't noise. The score under-weighted humidity and heat, over-penalized light drizzle, and ignored the sun entirely. It also had a few plain bugs, among them a rain-chance term that was computed and then left out of the total, and a wind curve that started scoring stronger wind as better past about 28 mph.
One model, shown two ways
Instead of re-tuning the old formula, we dropped it for race day. The new RunScore starts from pace cost: the percentage slowdown our model expects from heat, wind, and rain over the hours runners are on course. Hills aren't in it. RunScore is weather only, and terrain is priced separately in the pace adjustments.
That cost maps onto the 0–10 scale like this:
| Weather cost | Score | Label |
|---|---|---|
| Under 1.2% | 8 to 10 | Great |
| 1.2 to 3.2% | 6.4 to 8 | Good |
| 3.2 to 5% | 4.8 to 6.4 | Fair |
| 5 to 8% | 3 to 4.8 | Tough |
| Over 8% | Below 3 | Brutal |
Any cost above 12% gets the lowest score. Applied to every race day in our sample, 54% land Great, 25% Good, 13% Fair, 7% Tough, and 0.3% Brutal. That last group is five race days in 27 years, the hottest in the sample, among them Med City 2012 and Grandma's 2006. Races organizers stopped for heat, like Chicago 2007, would land there too, but they're left out of the fit because their results don't reflect a full race for most of the field.
Heat and rain are the same calculation in the score and in our pace adjustments, so the two can't disagree about them. Wind is the one deliberate difference, for reasons covered below.
Calibrating it on 6.3 million finishes
A pace-cost model is only as good as its constants. Published studies give a range for how much a degree of heat costs a marathoner. We wanted our own number, fit to the exact heat estimate the site computes.
We pulled finisher results for 92 marathons from 2000 through 2026 (81 in the US, 11 elsewhere) and matched each edition to the hourly weather for its race morning, from our stored weather history and, for older years, Open-Meteo's historical archive. After exclusions, that left 1,686 race editions and about 6.3 million finish times.
The core idea is comparing each race to itself. Boston's hills are the same in a cool year and a hot one, so year-over-year differences at the same race are mostly weather. On top of that:
- Common trends. Finish times shift for reasons unrelated to weather: carbon-plated shoes, growing fields, the unusual fields of the COVID years. A control for each calendar year absorbs those across all races at once.
- Course changes. We researched route changes for 40 races and gave each route era its own baseline. It didn't move the heat estimate, which was reassuring.
- Broken years. Races cut short, stopped mid-race, run on a mismeasured course, or held virtually were excluded. Delayed starts were shifted to the real start time.
- Who was out when. When comparing faster and slower runners, each tenth of the field was scored against the weather of the hours it was actually on course, so a 3:00 runner isn't charged for the noon heat a 5:00 runner ran through.
We caught a real data bug partway through. Some results files had gun time, measured from the first starting gun, in the field that should hold chip time. In a big wave start that can add over an hour to the later waves. Fixing it moved 140 race editions by more than five minutes. The heat estimate barely changed (0.324 to 0.334% per degree), but two findings that had looked real didn't survive: an extra penalty for dry heat disappeared, and the extra cost of unseasonable heat shrank by about half. Both had leaned on the races with bad files, which were mostly dry western ones like San Francisco and Los Angeles. Their inflated times had looked like a dry-heat penalty.
An earlier pass on just 13 big races had pointed to a steeper heat cost (0.43% per degree) and a much harsher cold-rain penalty (about 8% per inch). Neither held up. Rebuilding the pipeline on the same 13 races already pulled both down, and adding 79 more races tightened the heat number and pushed cold rain toward zero. A small sample of famous races is an easy place to fool yourself.
What the data showed
Heat is the big one
Every degree of WBGT above 50°F costs the middle of the pack about 0.33%. (WBGT is the humidity-and-sun-weighted heat index race medical teams use, explained in our WBGT guide.) It usually reads a few degrees below the thermometer. 50°F WBGT is about a 59°F morning when it's dry and overcast, and closer to 50°F in full sun with high humidity, so a 60°F start is usually already past the line. Across the different ways we fit the full sample, the estimate stayed between 0.32 and 0.37%. Dropping any single race and refitting never pushed it outside 0.32 to 0.34. Here is every race day in the sample, with the line the score uses.
Faster runners lose less per degree: about 0.26% for the fastest tenth of the field, rising to about 0.33% by the midpack. That's after accounting for timing. Fast runners do finish before the day warms, but each group here is scored against its own hours on course. And on mornings when the temperature barely moved from gun to finish, the midpack still lost more per degree than the front. Slower runners really are more sensitive to heat.
The score uses 0.26% per degree for runners at 6:00 per mile or faster, rising to 0.35% at 9:10 per mile and holding there. On hot days, the runners hit hardest are the likeliest to skip the race or drop out, so they disappear from the finish times and the back of the pack looks less affected than it really is. That's the likely reason the two slowest groups in the chart dip, and it's why the score holds the cost flat for slower runners instead of letting it fall, a little above the 0.33% fit. At a WBGT of 68°F, about what NYC 2022 averaged, the score's heat cost comes to roughly 12 minutes for a 3:30 marathoner and 17 for a 4:30. Try your own numbers:
Rain is the weakest signal
Cold rain, below about 55°F, does seem to cost something, but the estimate is noisy: roughly 1 to 3% per inch, and we can't statistically rule out zero. Warm rain showed no cost at all, likely because the cooling offsets the soaked shoes and clothes. The score charges 3% per inch at 50°F and below, tapering to 1% per inch at 60°F and above. Both sit on the cautious side of what the data supports.
We also tested whether a forecast's chance of rain matters on its own. Once you know how much rain actually fell, it doesn't.
Headwinds barely show up
This was the surprise. Boston 2018 ran in the mid-30s with torrential rain and gusts to 32 mph, mostly in runners' faces. Full drag physics says that headwind should cost more than 20%. The median finisher ran about 3% slower than in Boston's cool years. Boston 2015, dry with a steady 17 mph headwind, prices at 16% by the physics. That field ran slightly faster than a typical cool Boston.
Across all 92 races, the days our physics flagged as the worst headwinds showed no measurable extra slowdown at the median. Tailwinds did show up, at roughly 40% of what the physics predicts. Finish times alone can't tell us why. Packs draft, city streets shelter runners, and wind measured at a weather station overstates what reaches the course. Probably all three.
So the score uses 15% of the wind physics. Anything from 10 to 25% fits the data, and 15% keeps big wind days visible without letting them dominate. The pace adjustments on our race pages keep the full physics, with a drafting toggle for running in a pack, because if you end up alone in that wind, it applies to you. Even at 15%, the score still marks down days like Boston 2015 and 2018 more than their results justify.
Testing it on days we know
Here's the old score against the new one on race days with known outcomes, next to how the median finisher actually did compared with that race's own cool-weather years.
| Race | Weather | Old score | New score | Median finisher vs. cool years |
|---|---|---|---|---|
| Boston 2012 | 73–85°F | 7.2 Good | 3.8 Tough | 11.7% slower |
| NYC 2022 | 71–74°F, humid | 7.6 Good | 4.0 Tough | 7.1% slower |
| Erie 2026 | 70–74°F | 7.4 Good | 3.2 Tough | 8.7% slower |
| London 2026 | 53–63°F, warming | 8.4 Great | 7.8 Good | 1.8% slower |
| Boston 2026 | 45–53°F, tailwind | 8.4 Great | 10 Great | 2.8% faster |
Scores are on the 0–10 scale. Weather is the hourly station data each score was computed from, which can differ from start-line reports. Erie's scores use race-week forecast weather, and its result comes from the 968 finishers in the posted results.
The warm days are where the old score failed, and the new one gets NYC, Erie, and London about right. London shows how much timing matters: Sabastian Sawe and Yomif Kejelcha became the first two people to break two hours in a marathon race there in April, finishing before the morning warmed up, while the midpack paid about 2%. Boston 2026, the fastest Boston ever run, now gets full credit for its tailwind.
Boston 2012 is the miss. We price it at about 6.6%, and the field ran almost 12% slower. Part of the gap is field composition, since 2,160 registered runners chose to defer rather than start. The rest we can't pin down. It was one of the biggest heat spikes in the sample, about 21°F of WBGT above a typical Boston, and April heat arrives before runners have acclimatized. Our data does show a modest, noisy extra cost for heat above a race's normal. But NYC 2022 was nearly as far above its own normal, and the score got NYC about right. So the early heat can't be the whole story, and we'd rather leave Boston 2012 as an honest miss than bend the model to fit it.
What the score can't tell you
Measured by what the field actually lost, Boston 2018 was a Good-grade day. The median finisher ran about 3% slower. That same morning, about 2,500 runners were treated in medical tents, most for cold-related problems. Both are true. A pace-cost score measures how much slower the typical finisher ran. It says nothing about the runner who became hypothermic and never finished.
We won't stretch the number to express danger it isn't built to measure. That would make it wrong on every other kind of day. The plan is a separate warning, shown alongside the score, for the combination behind 2018: cold, soaking rain, and strong wind together. It isn't live yet.
Where it's still wrong
- It's probably too kind to the hottest days, especially early-season heat spikes like Boston 2012.
- It's probably too harsh on big headwind days, even with wind damped to 15%.
- The data is observational. On hot days the runners hurt most are the likeliest to stay home or drop out, which makes finish times look kinder than the day was.
- It leans American. 81 of the 92 races are in the US. The 11 international races fit the same heat cost, but that's a thin check.
- It's marathon-only for now. Half marathons borrow the marathon heat cost, which is probably a bit harsh for a shorter race. A half-marathon fit is on the list.
- It's a field average. It doesn't know your heat tolerance, your training, or whether you'll be tucked into a pack.
More guides
- WBGT Explained — the heat-stress index behind the score's heat term
- How Wind Actually Affects Your Marathon — the drag physics our pace adjustments still use in full
- Boston 2026: The Tailwind, the Cold, and the Fastest Boston Marathon Ever — what a cold morning and a westerly did at Boston this year
- How Weather Affects Your Marathon Pace — the pace adjustments this score is built from
- Marathon Pacing Guide — adjusting effort and splits when conditions get tough
- Find your race — the live RunScore and weather cost for your next start line