Mixture Models in R

Learn mixture models: a convenient and formal statistical framework for probabilistic clustering and classification.
Start Course for Free
Clock4 HoursPlay14 VideosCode47 ExercisesGroup3,022 Learners
Database3600 XP

Create Your Free Account

Google LinkedInFacebook
or
By continuing you accept the Terms of Use and Privacy Policy. You also accept that you are aware that your data will be stored outside of the EU and that you are above the age of 16.

Loved by learners at thousands of companies


Course Description

Mixture modeling is a way of representing populations when we are interested in their heterogeneity. Mixture models use familiar probability distributions (e.g. Gaussian, Poisson, Binomial) to provide a convenient yet formal statistical framework for clustering and classification. Unlike standard clustering approaches, we can estimate the probability of belonging to a cluster and make inference about the sub-populations. For example, in the context of marketing, you may want to cluster different customer groups and find their respective probabilities of purchasing specific products to better target them with custom promotions. When applying natural language processing to a large set of documents, you may want to cluster documents into different topics and understand how important each topic is across each document. In this course, you will learn what Mixture Models are, how they are estimated, and when it is appropriate to apply them!

  1. 1

    Introduction to Mixture Models

    Free
    In this chapter, you will be introduced to fundamental concepts in model-based clustering and how this approach differs from other clustering techniques. You will learn the generating process of Gaussian Mixture Models as well as how to visualize the clusters.
    Play Chapter Now
  2. 2

    Structure of Mixture Models and Parameters Estimation

    In this chapter, you will be introduced to the main structure of Mixture Models, how to address different data with this approach and how to estimate the parameters involved. To accomplish the estimation, you will learn an iterative method called Expectation-Maximization algorithm.
    Play Chapter Now
  3. 3

    Mixture of Gaussians with `flexmix`

    This chapter shows how to fit Gaussian Mixture Models in 1 and 2 dimensions with `flexmix` package. The data used is formed by 10.000 observations of people with their weight, height, body mass index and informed gender.
    Play Chapter Now
  4. 4

    Mixture Models Beyond Gaussians

    In this module, you will learn how Mixture Models extends to consider probability distributions different from the Gaussian and how these models are fitted with `flexmix`. The datasets used are handwritten digits images and the number of crimes in Chicago city. For the first dataset you will find clusters that summarize the handwritten digits and for the second dataset, you will find clusters of communities where is more or less dangerous to live in.
    Play Chapter Now
In the following tracks
Probability and Distributions
Collaborators
David CamposChester IsmayShon InouyeBenjamin Feder
Víctor Medina Headshot

Víctor Medina

Doctoral Researcher at The University of Edinburgh
Victor Medina is a member of the Department of Research at the Superintendency of Banks and Financial Institutions (SBIF) in Chile. At SBIF, he develops predictive risk models and mathematical tools for the Chilean Financial System with the aim of improving the extra-situ supervision of financial entities. With a passion for teaching, he is a professor of Econometrics at the Universidad de Chile and works with the School of Political Sciences at Universidad Central to create applications to extract and analyze data from social media. Victor holds an Master of Science in Statistics.
See More

What do other learners have to say?

I've used other sites—Coursera, Udacity, things like that—but DataCamp's been the one that I've stuck with.

Devon Edwards Joseph
Lloyds Banking Group

DataCamp is the top resource I recommend for learning data science.

Louis Maiden
Harvard Business School

DataCamp is by far my favorite website to learn from.

Ronald Bowers
Decision Science Analytics, USAA

Join over 6 million learners and start Mixture Models in R today!

Create Your Free Account

Google LinkedInFacebook
or
By continuing you accept the Terms of Use and Privacy Policy. You also accept that you are aware that your data will be stored outside of the EU and that you are above the age of 16.