What assumptions are baked into your race prediction model?

After launching my power law calculator earlier this week I got a couple emails from readers who noticed some counter-intuitive behavior.

Suppose you have an athlete who has run 800m in 2:25 and 1600m in 5:10. The power law calculator predicts a 3200m performance of 11:03—pretty reasonable. 

Then suppose your athlete improves their 800m time to 2:20. Enter that into the calculator instead and you get a 3200m time of 11:26—which is slower than their predicted 3200m time before! This seems wrong: shouldn’t the runner now be able to run faster over 3200m as well?  

Now, this wouldn’t be the first time a reader spotted a subtle issue with one of my calculators, but in this case, the calculator is working correctly, under the assumptions baked into the model. Here’s the intuition: 

Suppose you had two different athletes: Fast-Twitch Frannie, who runs 1600m in 5:10 and 800m in 2:20, and Slow-Twitch Sally, who can also run 5:10 for 1600m, but can only run 800m in 2:25. If these two athletes race over 3200m, who will win? 

From this perspective, it should be clear that we’d expect Slow-Twitch Sally to win, because she tends to “convert up” to longer distances better—hence my illustrative nickname.

All predictive models have assumptions

At an intuitive level, the power law model is working off a few key assumptions: an athlete has some “overall” fitness level, and a slowdown ratio (called the “fatigue factor”) that dictates how much they slow down as the race distance gets longer. 

So, under these assumptions, if you set one performance as fixed (e.g. 1600m fitness is 5:10), the only way to improve performance at a shorter distance is to give up performance at a longer distance.

Notably, the critical speed model (and my accompanying critical speed calculator) has exactly the same behavior, and so will any simple two-parameter model.

There is some reality to this assumption: typically, you really do need to give up some 10k fitness to improve your 800m time, for example. But there’s a real problem with assuming your fitness at any given distance is fixed—in all likelihood, if our 5:10/2:25 runner improves her 800m time to 2:20, her 1600m time will improve as well.

Another option would be to assume that all performances improve together. This is the assumption that underpins more traditional race prediction charts, e.g. Daniels, McMillan, etc.

But with those sorts of prediction charts, you also need to assume that every runner slows down the same amount—according to the VDOT chart, Fast-Twitch Frannie and Slow-Twitch Sally should run exactly the same time for 3200m, if you rely on their 1600m performances to predict their times.

It seems like we’re in a bit of a bind: either we give up the ability to differentiate slow-twitch-oriented and fast-twitch-oriented athletes, or we accept that improvements will create this weird teeter-totter behavior at shorter and longer distances.    

Possible solutions for better race predictions

I see two possible ways to get around these limitations—a fun hack-fix and a more ambitious approach.

A quick fix: visualize improvement as a baseline shift

Here’s one way to think about our runner’s improvement from 2:25 to 2:20 in the 800m: we just take her performance curve and shift the whole thing up by the same amount. 

For a power law model this amounts to a constant improvement of about 3% across all distances—taking the performance curve and just shifting the whole thing up or down (a so-called “baseline shift”). 

Now, this is a bit of a hack fix because we still have a somewhat questionable assumption baked in: we are assuming the athlete’s previous performances over 1600m and 800m defined a characteristic “fatigue factor” for her that stays fixed. Over the course of a season this is not a terrible assumption, but across the long arc of your career, one thing you’ll (hopefully) find is that your fatigue factor improves steadily—even at the same 1600m time, your 3200m performance increases.  

This was a pretty easy modification to make to the power law calculator: I added a slider to visualize what improvements or decrements in performance of up to 5% look like:

The idea with this slider is that you should use it to look at constant improvement for the same athlete. When comparing performances across different athletes, that’s when you’d want to enter a new set of race performances. 

With the slider, you're basically asking, if I kept the same performance profile and improved by a constant amount, how much faster would I be at each predicted race distance?

You can see the effects on any predicted performance in the table below the plot:

A more ambitious approach: pooling data across athletes

Now, the bigger question: How do you deal with the fact that your fatigue factor may not be constant over time? The only way around this problem, without running into the original problem, is to develop a way to pool information across athletes.

As we saw above, the problem with two-parameter models like power laws or critical speed is that they are “naive” to the typical behavior of performances within the same athlete (i.e. the fact that, usually, if your 800m time gets faster, so does your 1600m time). And the problem with traditional conversion charts is that they are “naive” to the individual characteristics of the athlete—they assume all athletes have the same profile across distances.  

Given how much longitudinal data we have from athletes of different levels and in different events, it should be possible to develop models that “know” something about the typical range of fatigue factors for athletes of a given level, and how quickly these fatigue factors tend to change over time. 

You could even imagine more sophisticated models that separately account for more fine-grained aspects of performance, like top speed, anaerobic energy reserves, aerobic fitness, physiological resilience, etc.

To build this kind of model, you need data—and lots of it. Conceptually, this more ambitious approach is similar to the approach I used for my CV, threshold, and VO2max calculator, which used a data-driven approach for predicting critical speed with just one performance level. 

One huge advantage of pooling data across different runners is that you can provide meaningful uncertainty estimates: for example, given that an athlete can run 5:10 for 1600m, you could provide a range of times for, say, 3200m that are plausible for a runner of this level. 

And if you were to run a new performance (say, an 800m race in 2:25), such a model could provide a better estimate and narrower uncertainty intervals. Given an improvement (another 800m, this time in 2:20), you could also project out the plausible range of improvements for 1600m and 3200m.  

I have exactly this sort of data-driven approach for race prediction on my list of long-term projects, so sign up for my email newsletter if you want to find out when I crack this problem. 

Recap

Any conversion chart or race prediction calculator for runners is built on a set of assumptions. For “individualized” models like power law or critical speed models, the assumption is that all of your race performances represent your current best effort over each distance. This assumption can have some counter-intuitive effects, like predicted performances at longer distances getting slower when your performance at shorter distances gets faster.  

Conversely, simplified charts and calculators like Daniels’ VDOT chart or the McMillan calculator assume that all runners have the same performance profile across different distances. This assumption has its own set of problems: if you are more short-distance oriented, your predicted performance in longer races will be too fast, and vice versa for more long-distance oriented runners.  

The best solution to this problem will ultimately be a calculator that pools data across different runners—accounting for (a) the typical performance profile of a runner like you, and also (b) your particular performance history and improvement trajectory. We’re not there yet, but I have some ideas about how to build such a model. 

In the meantime, check out my power law calculator and try out the new performance-improvement visualization feature! 

Learn more about the science of performance

If you want to take a scientific approach to your own training, I have two books that outline exactly how to do it. The most recent is Marathon Excellence for Everyone, which takes a full-spectrum percentage-based approach to running your best in the marathon. 

If you live outside of the United States, Marathon Excellence is also available in a Metric Edition, with all workouts in kilometers!

I also have a shorter book, Modern Training and Physiology for Middle and Long-Distance Runners, that examines a simple version of scientific training for 800m–10k training. Check it out! 

If you want to get updated when I have a new app, new book, or new article coming out, the best way to do it is to join my free email list below. I’m also on Instagram at @jdruns but by far the best way to follow my work is my email list! 

Related articles

New web app: Predicting LT1 pace and Zone 2 pace from 5k time

I’m excited to launch a new app for predicting training paces: my LT1 and Zone 2 pace calculator is now live, and provides a simple and accurate way to estimate your first lactate threshold, or LT1 pace, as well as an upper limit for your easy run pace, a.k.a. Zone 2 pace.  📲 Check out ... Read more

Designing a plyometrics program for improving bone strength in young runners

Stress fractures are an incredibly frustrating injury, especially for young runners. Even though the science behind recovery from stress fractures—more properly called “bone stress injuries”—has advanced significantly in the last several years, sustaining a bone stress injury can still completely derail your season. Of course, far better than a swift rehabilitation is simply not getting ... Read more

Three theories of tissue damage accumulation during running

Suppose you need to run 15 miles (24 km) in the next week. You can distribute this volume however you want, both within and across days. Your goal is to cover the requisite distance in a way that minimizes your risk of injury. Does it matter how you schedule out your week? This is one ... Read more
John Davis lecturing about marathon science

Lecture: The science behind modern marathon training

I just posted the video from my live lecture on the science of modern marathon training!  In the video, I uncover the science behind the modern approach to marathon training, including how VO2max, running economy, lactate threshold, and physiological resilience each contribute to marathon performance. Then, I explore the training methods that most effectively target ... Read more

What assumptions are baked into your race prediction model?

After launching my power law calculator earlier this week I got a couple emails from readers who noticed some counter-intuitive behavior. Suppose you have an athlete who has run 800m in 2:25 and 1600m in 5:10. The power law calculator predicts a 3200m performance of 11:03—pretty reasonable.  Then suppose your athlete improves their 800m time ... Read more

Tendons do not store energy for free

Very often, in training discussions online and in books—even sports science textbooks—you encounter the claim that, during running, tendons stretch out and store up energy on impact with the ground, releasing that energy later when you push off the ground. This energy storage improves your running economy, because without it, you’d have to produce that ... Read more

In windy conditions, running at a constant effort is usually better than running at a constant speed

Suppose you are running a 5k race on an out-and-back course, and there’s a strong headwind on the way out—should you aim to run at the same effort the whole way, allowing the wind to slow you down on the way out and speed you up on the way back? Or should you maintain the ... Read more

A comprehensive guide to the science of cadence for runners

Your cadence is the number of steps you take per minute while running. Simple measurement, right? But there are many questions surrounding it, including how it differs across runners, how it changes as you run faster, whether a higher cadence is more efficient, whether a lower cadence causes injury, and whether you should aim for ... Read more

About the Author

John J. Davis, Ph.D.

I have been coaching runners and writing about training and injuries for over 12 years. I've helped complete novices, NXN-qualifying high schoolers, elite-field competitors at major marathons, and runners everywhere in between. I have a Ph.D. in Human Performance, and I do scientific research focused on the biomechanics of overuse injuries in runners. My new book on marathon training, Marathon Excellence for Everyone, is now available on Amazon!

Leave a Comment

Check out my new book on marathon training!