Handling Missing Data: A guide to Mean Imputation vs. Multiple Imputation (and why it matters for your results)

Missing data can ruin your study’s validity. Learn why Multiple Imputation is the gold standard over Mean Imputation for protecting

Handling Missing Data: A guide to Mean Imputation vs. Multiple Imputation (and why it matters for your results)

The absence of data is a killer of numerous studies. You take months in gathering surveys or in clinical trials. To your surprise, you realize that the so-called Participant 42 had not answered half of the questions. It is frustrating, but how you deal with those empty cells will dictate whether your whole study is valid or not.

When you do not pay attention to the gaps then you will run the danger of great bias. When you fill them wrong, you are producing a false certainty. The phrase quick fix in the world of professional Manuscript Editing is mostly used to refer to papers that have been rejected due to the fact that the author opted to find a quick fix to a gap in the data. These options undermine the end results.

This is the reason why your decision on whether to use Mean Imputation or Multiple Imputation is important to your scientific integrity.

The “Quick Fix”: The Mean Imputation

Mean Imputation has the easiest method of plugging a hole. Suppose that five per cent. of the hundred who reported their age did not report. You just divide the number of the other 95 participants by average age. Then you fill in the blanks using the same number.

This is, of course, logical on the surface. You retain your entire sample size and the average man remains average. There are however a number of risks associated with this approach. First, it shrinks variance. You are falsely narrowing down the scope of your data by adding several of the same values. This gives your results a more consistent appearance than they would be.

Second, it distorts correlations. When you fill in a missing point with a mean, you are disregarding the correlation that the variable has with other variables. This makes the perceived relationship between variables in your study weak. Lastly, it under-estimates error. The standard errors are reduced. This causes false positive results i.e. Type I errors.

The Gold Standard: Why Multiple Imputation Wins

Multiple Imputation (MI) is a more advanced, and an honest method. MI does not fill in a gap with a single number, instead it uses the rest of the information you have in your data. It makes the forecast of what the missing value probably was based on the current trends.

Three steps are involved in the process. The software prepares a number of complete datasets in the Imputation phase. The missing values in both versions are slightly different, depending on a regression model that takes into account uncertainty. This is followed by Analysis. You actually do your statistical tests on each of those datasets on its own. The last step is the Pooling phase, which incorporates the results into a single estimate.

Any high-impact paper is built on the basis of statistical rigor. You do not lose the natural noise and uncertainty of the data when you use Multiple Imputation. The approach gives a better approximation of the actual population parameters. It also provides broader and more realistic confidence intervals. A majority of the high-impact journals have changed their expectations and require researchers to adopt sophisticated methods such as MI as opposed to plain means replacement.

In What Cases Can Mean Imputation Be Used?

Data

Is Mean Imputation a priori bad? Not necessarily. In case you are missing less than 5 percent of your data, and the data is Missing Completely at Random (MCAR), the effect may remain unimportant. But in complicated human research, data do not disappear arbitrarily.

Individuals having lower income are likely to report their income less. The more severe patients may leave a study. It is Missing Not at Random (MNAR). Mean imputation plays on you in such instances.

How Mars Publications Can Help

The selection of an appropriate imputation strategy is not the entire battle. You should also explain that strategy to your peers and editors. We do not only edit manuscripts to correct grammar. We scrutinize your statistical reporting.

Mars Publications provides you with the clearly mentioned percentage of missing data for each variable. We have tested whether you have managed to find the mechanism of missingness. Our editors will also make sure that you describe the particular software and iterations that you applied to your Multiple Imputation.

Final Thoughts

Missing data is no excuse to betray your effort. As much as Mean Imputation is an easy method to attract the attention of a researcher, Multiple Imputation secures the truth of your work in science. Good quality research needs high-quality data processing.

In case you are unsure about the method to apply to your data, or you require assistance in justifying your decision to the journal editor, you can contact us. In Mars Publications, we are giving you the technical know-how and the professional sheen that your research requires to bring you to the finishing line.

Latest Articles