Skip to main content

Course

Handling Missing Data with Imputations in R

AdvancedSkill Level

4.7+

Updated 10/2022

Diagnose, visualize and treat missing data with a range of imputation techniques with tips to improve your results.

Start Course for Free

RData Manipulation

4 hr

13 videos

49 Exercises

4,200 XP

6,223

Statement of Accomplishment

Loved by learners at thousands of companies

Training a Team?

Try for Business

Course Description

Missing data is everywhere. The process of filling in missing values is known as imputation, and knowing how to correctly fill in missing data is an essential skill if you want to produce accurate predictions and distinguish yourself from the crowd. In this course, you’ll learn how to use visualizations and statistical tests to recognize missing data patterns and how to impute data using a collection of statistical and machine learning models. You’ll also gain decision-making skills, helping you decide which imputation method fits best in a particular situation. Finally, you’ll learn to incorporate uncertainty from imputation into your inference and predictions, making them more robust and reliable.

Prerequisites

Intermediate Regression in R Dealing With Missing Data in R

1

The Problem of Missing Data

In this chapter, you’ll find out why missing data can be a risk when analyzing a dataset. You’ll be introduced to the three missing data mechanisms and learn how to recognize them using statistical tests and visualization tools.

Missing data: what can go wrong

Linear regression with incomplete data

Analyzing regression output

Comparing models

Missing data mechanisms

Recognizing missing data mechanisms

t-test for MAR: data preparation

t-test for MAR: interpretation

Visualizing missing data patterns

Aggregation plot

Mosaic plot

2

Donor-Based Imputation

Get to know the taxonomy of imputation methods and learn three donor-based techniques: mean, hot-deck, and k-Nearest-Neighbors imputation. You’ll look under the hood to see how these methods work, before learning how to apply them to a real-world tropical weather dataset. Along the way, you’ll also learn useful tricks that you can use to make them work even better for your problems.

Mean imputation

Smelling the danger of mean imputation

Mean-imputing the temperature

Assessing imputation quality with margin plot

Hot-deck imputation

Vanilla hot-deck

Hot-deck tricks & tips I: imputing within domains

Hot-deck tricks & tips II: sorting by correlated variables

k-Nearest-Neighbors imputation

Choosing the number of neighbors

kNN tricks & tips I: weighting donors

kNN tricks & tips II: sorting variables

3

Model-Based Imputation

It’s time to learn how to use statistical and machine learning models, such as linear regression, logistic regression, and random forests, to impute missing data. In this chapter, you’ll look into how the models make their predictions and use this knowledge to draw the imputed values from conditional distributions. This is important as it ensures your imputations are more varied and plausible, making them more similar to the true data.

Model-based imputation approach

Linear regression imputation

Initializing missing values & iterating over variables

Detecting convergence

Replicating data variability

Logistic regression imputation

Drawing from conditional distribution

Model-based imputation with multiple variable types

Tree-based imputation

Imputing with random forests

Variable-wise imputation errors

Speed-accuracy trade-off

4

Uncertainty from Imputation

Imputed values are not set in stone. They are just estimates and estimates come with some uncertainty. In this final chapter, you’ll discover how bootstrapping and chained equation using the mice package can be used to incorporate imputation uncertainty into your models and analyses to make them more reliable and robust.

Multiple imputation by bootstrapping

Wrapping imputation & modeling in a function

Running the bootstrap

Bootstrapping confidence intervals

Multiple imputation by chained equations

The mice flow: mice - with - pool

Choosing default models

Using predictor matrix

Putting it all together

Analyzing missing data patterns

Imputing and inspecting outcomes

Inference with imputed data

Final remarks

Handling Missing Data with Imputations in R

Course
Complete

Earn Statement of Accomplishment

Add this credential to your LinkedIn profile, resume, or CV
Share it on social media and in your performance reviewEnroll Now

Don’t just take our word for it

*4.7

from 104 reviews

80%

16%

4%

0%

0%

Sort by

Casey

yesterday

I appreciate this course providing me crucial knowledge and skills dealing with missing data

YOVI

2 weeks ago

Renzo

2 weeks ago

JORGE LUIS ANGEL

2 weeks ago

Julien

2 weeks ago

Amy

2 weeks ago

"I appreciate this course providing me crucial knowledge and skills dealing with missing data"

Casey

YOVI

Renzo

FAQs

What imputation methods are taught in this course?

You will learn mean imputation, hot-deck imputation, k-nearest neighbors, linear regression, logistic regression, random forests, and multiple imputation using the mice package.

How does this course differ from Dealing With Missing Data in R?

Dealing With Missing Data in R is a prerequisite that covers recognition and basic handling. This advanced course focuses specifically on imputation techniques and incorporating imputation uncertainty.

What are the three missing data mechanisms and will I learn to identify them?

They are missing completely at random, missing at random, and missing not at random. Chapter 1 teaches you to recognize them using statistical tests and visualizations.

What real-world dataset is used for practice?

You will apply imputation techniques to a tropical weather dataset, practicing donor-based and model-based methods on real missing data patterns.

Does the course cover uncertainty in imputed values?

Yes. The final chapter teaches bootstrapping and chained equations with the mice package to incorporate imputation uncertainty into your analyses and make results more robust.

Join over 19 million learners and start Handling Missing Data with Imputations in R today!

Grow your data skills with DataCamp for Mobile

Make progress on the go with our mobile courses and daily 5-minute coding challenges.