본문으로 바로가기

강의

PySpark로 배우는 빅데이터 기초

고급기술 수준

업데이트됨 2025. 2.

PySpark를 활용한 빅데이터 작업의 기초를 익히세요.

무료로 강의 시작

SparkData Engineering

4시간

16 동영상

55 연습 문제

4,600 XP

65,217

성취 증명서

수천 개 기업의 학습자들이 사랑하는

팀을 교육하시나요?

비즈니스용으로 체험해 보세요

강의 설명

지난 몇 년간 빅데이터에 대한 관심이 꾸준히 커지면서 이제는 많은 기업에서 보편적으로 활용하고 있어요. 그렇다면 빅데이터란 정확히 무엇일까요? 이 강의는 PySpark를 통해 빅데이터의 기본을 다룹니다. Spark는 빅데이터를 위한 "번개처럼 빠른 클러스터 컴퓨팅" 프레임워크로, 범용 데이터 처리 엔진을 제공하며 메모리에서는 최대 100배, 디스크에서는 최대 10배까지 Hadoop보다 빠르게 프로그램을 실행할 수 있어요. 여러분은 Spark 프로그래밍을 위한 Python 패키지인 PySpark와 SparkSQL, MLlib(머신 러닝용) 등의 강력한 고수준 라이브러리를 사용하게 됩니다. 셰익스피어 작품을 탐색하고, Fifa 2018 데이터를 분석하며, 유전체 데이터셋에 클러스터링을 수행해 볼 거예요. 강의가 끝나면 PySpark와 일반적인 빅데이터 분석에의 적용에 대해 깊이 있게 이해하게 됩니다.

선수 조건

Introduction to Python

1

Introduction to Big Data analysis with Spark

This chapter introduces the exciting world of Big Data, as well as the various concepts and different frameworks for processing Big Data. You will understand why Apache Spark is considered the best framework for BigData.

What is Big Data?

The 3 V's of Big Data

PySpark: Spark with Python

Understanding SparkContext

Interactive Use of PySpark

Loading data in PySpark shell

Review of functional programming in Python

Use of lambda() with map()

Use of lambda() with filter()

2

Programming in PySpark RDD’s

The main abstraction Spark provides is a resilient distributed dataset (RDD), which is the fundamental and backbone data type of this engine. This chapter introduces RDDs and shows how RDDs can be created and executed using RDD Transformations and Actions.

Abstracting Data with RDDs

RDDs from Parallelized collections

RDDs from External Datasets

Partitions in your data

Basic RDD Transformations and Actions

Map and Collect

Filter and Count

Pair RDDs in PySpark

ReduceBykey and Collect

SortByKey and Collect

Advanced RDD Actions

CountingBykeys

Create a base RDD and transform it

Remove stop words and reduce the dataset

Print word frequencies

3

PySpark SQL & DataFrames

In this chapter, you'll learn about Spark SQL which is a Spark module for structured data processing. It provides a programming abstraction called DataFrames and can also act as a distributed SQL query engine. This chapter shows how Spark SQL allows you to use DataFrames in Python.

Abstracting Data with DataFrames

RDD to DataFrame

Loading CSV into DataFrame

Operating on DataFrames in PySpark

Inspecting data in PySpark DataFrame

PySpark DataFrame subsetting and cleaning

Filtering your DataFrame

Interacting with DataFrames using PySpark SQL

Running SQL Queries Programmatically

SQL queries for filtering Table

Data Visualization in PySpark using DataFrames

PySpark DataFrame visualization

Part 1: Create a DataFrame from CSV file

Part 2: SQL Queries on DataFrame

Part 3: Data visualization

4

Machine Learning with PySpark MLlib

PySpark MLlib is the Apache Spark scalable machine learning library in Python consisting of common learning algorithms and utilities. Throughout this last chapter, you'll learn important Machine Learning algorithms. You will build a movie recommendation engine and a spam filter, and use k-means clustering.

Overview of PySpark MLlib

PySpark ML libraries

PySpark MLlib algorithms

Collaborative filtering

Loading Movie Lens dataset into RDDs

Model training and predictions

Model evaluation using MSE

Classification

Loading spam and non-spam data

Feature hashing and LabelPoint

Logistic Regression model training

Loading and parsing the 5000 points data

K-means training

Visualizing clusters

Congratulations!

PySpark로 배우는 빅데이터 기초

강의
완료

수료증 획득

LinkedIn 프로필, 이력서 또는 CV에 이 인증서를 추가하세요
소셜 미디어와 성과 평가에서 공유하세요지금 등록

19백만 명 이상의 학습자와 함께 PySpark로 배우는 빅데이터 기초을(를) 시작하세요!

DataCamp for Mobile을 통해 데이터 분석 능력을 향상시키세요.

모바일 강좌와 매일 5분 코딩 챌린지를 통해 이동 중에도 학습 효과를 높이세요.