In this article, we analyze a forecast that locked down a nation and show how easy to make small mistakes in modeling lead to big problems in outcomes.
In this analysis, we will work to validate an Intensive Care Unit (ICU) peak burden forecast. The forecast is made by national health officials. Such forecasts play a prominent role in time-constrained policy decisions that try to balance between complete societal paralysis and unnecessary loss of human life. Scrutiny of every parameter value, model, and figure is of the utmost importance. Only with scrutiny, sound decisions can be made based on data and modeling.
We will use the Finnish health authority — THL — and their ICU peak burden forecast as an example. Everything in the article can be applied to other similar forecasts.
We have made all findings, materials, and the codes allowing straightforward reproduction of the results available. You can find the link at the end of the post.
In this report, we conclude the following:
- When we correct THL’s model parameters to accord with empirical evidence, otherwise keeping the model exactly as it is, we find 45% error in the published forecast
- THL’s model assumes a curve that is already notably behind empirical time-series data
- THL’s model is based on exponential growth every 23 days, where the growth rate is closer to 10 days based on evidence
- Using a Monte Carlo simulation in an effort to reproduce THL’s forecast, we find a limited number of scenarios where the outcome is possible
Empirical evidence, comparable methods, and the apparent error in the assumptions of THL’s model show that it is unlikely that the forecasted ICU peak burden is realistically achievable. Even though in this report, we don’t present alternative forecasts, we conclude that using the methods we have published[1], it is easy to observe that higher ICU peak burden is likely. More than enough to beg an important question.
How would a higher number change the decisions that already been made based on the lower number?
The simulations we have performed are based on the Monte Carlo method and are easy to understand even for someone with no formal statistical training. Using the simpler version[2], we have gone through a large number of possible scenarios by changing doubles_in_days, max_patient_capacity, and case_fatality_rate input parameters. We have verified the results with a more elaborate[3], also Monte Carlo based simulator that allows many more input parameters and separates standard ICU burden from ventilated ICU burden.
In a 20–day simulation, we find that roughly half of the possible scenarios end up in a loss of life due to there being more demand for ICU than there is capacity.
In most scenarios, intensive care runs short on fewer than 2.6 days; in about 50 of some 1,600, on 10.4 to 13 days.
Detailed results for the simulation are outside the scope of this article and we will release it within two weeks from now in a separate article.
What is a Good Data-Driven Decision?
There are many ways to use available information to support decision-making; based on data based on current understanding (theory) and based on intuition. Decisions that are based on data can be based on recent events or historical trends. When we use data, either historical or current, the goal is to use that data to say something about the future. When theory, intuition, and data all say the same thing, it can give credibility to the forecast or finding. The finding must be reproducible; otherwise, it must be considered an error. Even if the finding otherwise meets the criteria of credibility, unless it can be reproduced, only low confidence is warranted
Still, we have to keep in mind that even when a finding has credibility, it does not mean that we can predict the future. We can try, and sometimes we get it right, but we can never use history to say something about the future certainly. The only available tool to say things about the future certainly is “a valid inference,” a tool used in philosophy. For example, we can certainly say that whenever there is combustion, there will be some heat. But as data scientists, we do not have that kind of luxury. We always work with probabilities and distributions. Never with certainties. This is a subtle but very important point in our field of work.
Pascal’s Wager
The father of the probabilistic method, Blaise Pascal, developed statistics to answer the question if God exists. After years of working on the problem, Pascal ended up concluding that even though the probability of God existing is very small, one should not take the risk of living a life of unvirtue. This is called “Pascal’s Wager,” and it is a key tenet in applying the probabilistic method. Particularly in relation to the responsibility we have in presenting the findings of our work. “Pascal’s Wager” applies to all critically important predictions, where for example loss of human life is a result of forecasting error. A meaningful way to present findings is one where a balance between the probabilities and the caveats associated with the method always go together. But particularly in the presentation phase.
The Importance of Empirical Evidence
Modeling must be underpinned by empirical evidence. Regarding COVID-19, in many areas, we are in a situation where available data paints a meaningful picture of the problem. Below we provide a summary of the available evidence we used in this report, and which can be helpful in modeling ICU peak burden. All figures are based on the most recent data available for Finland. Most countries make similar data available.
How common is the virus?
Roughly 6% of all tested cases are positive.
How many of those that tested positive require hospitalization?
More than 10% of all that tested positive require hospitalization.
How many of those hospitalized require ICU?
Roughly 40% of those hospitalized require intensive care.
How many in the ICU require ventilation?
Based on the available literature, we estimate it to be between 55%[4] and 65%[5]
What is the duration of ICU stay?
Depending on the source, we estimate it to be between 7–12[6][7] days mean duration. We note that the lower numbers are possibly due to accelerated mortality as a consequence of capacity issues. We find that among influenza patients, ICU stays historically last 12 days[8] mean duration.
How many ICU patients are there currently?
As of the 6th of April, 2020, there are 81 ICU patients.
What is the growth factor of the ICU burden at the moment?
Roughly 50% per week.
Critically Evaluating Already Made Forecast
THL has presented that the ICU burden peak will be less than 300. The peak is forecasted to take place on week 9 of the curve. Finnish politicians, including the Prime Minister, have repeatedly referred to THL’s forecast driving various measures, such as social distancing. THL’s model assumes as its starting point roughly 60 ICU patients, and a doubling rate of almost four weeks.
Week 9: about 920 in hospital, about 300 in intensive care.
- In hospital
- In intensive care
The empirical doubling rate, based on the number of ICU patients between the of March and of April[8], is about half of that. In other words, the actual growth is about twice what THL’s forecast expects.
Even a relatively small error in the growth factor can have a markable effect on the outcome. THL’s model assumes a 1.2x weekly growth rate. If the growth is instead 1.3x, the relative accuracy of the forecast is 91.7% — but the absolute ICU burden is more than a quarter higher. 1.4x weekly growth rate, which is barely 15% difference in relative accuracy, results in almost twice the absolute peak burden.
At THL’s 1.2× a week, the nine-week peak is about 140 in intensive care; at the 1.5× measured, about 370.
Another important assumption in THL’s model is the relationship between hospitalization and the need for intensive care. For example, for the second week of the curve, THL predicts 283[9] hospitalization, with 88 of the patients requiring intensive care. This means THL assumes a 31% ICU rate. Based on available data, the actual rate in Finland is 39%[10], or 21% higher than THL’s assumption. Deviation of such a degree, from empirical data, can be considered important. Moreover, THL explains the model to assume the average ICU stay to last for eight days. Based on the available literature, this is likely to be around ten days. Based on historical patient records, the average ICU stay duration is 8.5 days, but the average for influenza patients is 10.5 days[11]. Because of this, we find a 20% additional error to THL’s forecast.
By combining the error arising from the way the base assumptions deviate from empirical data, we have to assume that based on THL’s model, the peak burden will be 406 and not 280. We find that while the model itself may be useful, the current publicized forecast contains an important error. This 45% error is assuming that no other deviations can be identified i.e., this does not include any other issues in the model. In short, summary, when we correct the input parameters in THL’s model to follow empirical evidence, the published forecast based on the model contains an important error.
We also find that when comparing THL’s model with several available models, notably higher peak burden values are prominent.
It is important to understand that deviations from facts are likely; nobody in the world clearly understands the COVID-19 forecasting problem yet. Over a decade, among over 60,000 ICU stays[12], we find that patients with influenza and pneumonia appear in 187 cases. When at the moment, many hospitals handle more cases in a single day, it is clear that we are all exploring new territory here. This makes modeling and forecasting very difficult and highlights the importance of leveraging methods such as time-series forecasting with empirical data, and Monte Carlo simulation with empirical data influencing input parameter ranges. That, and close co-operation with fields where relevant skillsets — such as fintech and adtech — have been mastered.
“When on a Back of Tiger, Tread Carefully”
As Pascal’s Wager stipulates, outputs of models should be evaluated in the light of how much damage making a wrong decision can cause. For example, if lung cancer is diagnosed with negative falsely, the patient’s odds for survival dramatically drop. Likewise, if the need for ICU capacity is falsely forecasted to be lower than it will actually be, odds for patients to survive drop dramatically. This is true for COVID-19 patients and all patients that would otherwise require intensive care during peak capacity periods. Because of the way intensive care is structured to support life when it otherwise would fail, it’s safe to assume that once intensive care is denied from a patient who requires it, that patient’s chance of survival drops to zero.
Used Methods
The best way to test a claim is to reproduce it using a variety of different methods. Using the methods highlighted below, we have tried but failed to reproduce THL’s forecast. We can consistently achieve it only when we deviate away from empirically sound input parameters.
- Literature review
- Review of publicly available models
- Quantitative analysis of ICU stay data
- Regression analysis of available empirical data
We have made all findings, materials, and the codes allowing straightforward reproduction of the results through https://autonom.io/icu_burden and encourage all researchers to do the same with their work.
Can you scrutinize your country’s official forecasts in a similar way? Or the claims that we make in this article? Scrutiny is the force that propels knowledge forward. Do it, go scrutinize.
Note from the editors: Towards Data Science is a Medium publication primarily based on the study of data science and machine learning. We are not health professionals and the opinions of this article should not be interpreted as professional advice.