Skip to main content
Mitchell Zhen avatar

Mitchell Zhen has completed

Exploring and Analyzing Data in Python

Start course For Free
4 hr
4,150 XP
Statement of Accomplishment Badge

Loved by learners at thousands of companies


Course Description

How do we get from data to answers? Exploratory data analysis is a process for exploring datasets, answering questions, and visualizing results. This course presents the tools you need to clean and validate data, to visualize distributions and relationships between variables, and to use regression models to predict and explain. You'll explore data related to demographics and health, including the National Survey of Family Growth and the General Social Survey. But the methods you learn apply to all areas of science, engineering, and business. You'll use Pandas, a powerful library for working with data, and other core Python libraries including NumPy and SciPy, StatsModels for regression, and Matplotlib for visualization. With these tools and skills, you will be prepared to work with real data, make discoveries, and present compelling results.
For Business

Training 2 or more people?

Get your team access to the full DataCamp platform, including all the features.
DataCamp for BusinessFor a bespoke solution book a demo.
  1. 1

    Read, clean, and validate

    Free

    The first step of almost any data project is to read the data, check for errors and special cases, and prepare data for analysis. This is exactly what you'll do in this chapter, while working with a dataset obtained from the National Survey of Family Growth.

    Play Chapter Now
    DataFrames and Series
    50 xp
    Read the codebook
    50 xp
    Exploring the NSFG data
    100 xp
    Clean and Validate
    50 xp
    Validate a variable
    50 xp
    Clean a variable
    100 xp
    Compute a variable
    100 xp
    Filter and visualize
    50 xp
    Make a histogram
    100 xp
    Compute birth weight
    100 xp
    Filter
    100 xp
  2. 2

    Distributions

    In the first chapter, having cleaned and validated your data, you began exploring it by using histograms to visualize distributions. In this chapter, you'll learn how to represent distributions using Probability Mass Functions (PMFs) and Cumulative Distribution Functions (CDFs). You'll learn when to use each of them, and why, while working with a new dataset obtained from the General Social Survey.

    Play Chapter Now
  3. 3

    Relationships

    Up until this point, you've only looked at one variable at a time. In this chapter, you'll explore relationships between variables two at a time, using scatter plots and other visualizations to extract insights from a new dataset obtained from the Behavioral Risk Factor Surveillance Survey (BRFSS). You'll also learn how to quantify those relationships using correlation and simple regression.

    Play Chapter Now
For Business

Training 2 or more people?

Get your team access to the full DataCamp platform, including all the features.

datasets

National Survey of Family Growth (NSFG)General Social Survey (GSS)Behavioral Risk Factor Surveillance System (BRFSS)

collaborators

Collaborator's avatar
Chester Ismay
Collaborator's avatar
Yashas Roy

prerequisites

Python Toolbox
Allen Downey HeadshotAllen Downey

Professor, Olin College

See More

Join over 18 million learners and start Exploring and Analyzing Data in Python today!

Create Your Free Account

or

By continuing, you accept our Terms of Use, our Privacy Policy and that your data is stored in the USA.