본문으로 바로가기

강의

PySpark 입문

중급기술 수준

업데이트됨 2026. 1.

PySpark를 마스터하여 빅데이터를 손쉽게 처리하세요—대규모 데이터셋을 처리하고 쿼리하며 최적화하여 강력한 분석을 수행하는 방법을 배우세요!

무료로 강의 시작

SparkData Engineering4시간11 동영상36 연습 문제2,850 XP27,034성취 증명서

무료 계정을 만드세요

또는

계속 진행하시면 당사의 이용약관, 개인정보처리방침 및 귀하의 데이터가 미국에 저장되는 것에 동의하시는 것입니다.

수천 개 기업의 학습자들이 사랑하는

2명 이상을 교육하시나요?

DataCamp for Business 체험

강의 설명

이 과정은 대규모 데이터셋을 PySpark로 다루려는 데이터 엔지니어, 데이터 사이언티스트, 그리고 Machine Learning 실무자를 위해 설계되었습니다. Apache Spark의 속도와 확장성을 살펴보고, Spark 세션을 생성하며, RDD를 다루고, 실습을 통해 DataFrame을 조작하는 방법을 배웁니다. 또한 PySpark SQL을 다루며 SQL로 데이터를 조회하고, 스키마와 복합 데이터 타입을 처리하며, 분산 환경에서 성능을 최적화하는 방법을 익힙니다. 과정을 마치면 빅데이터를 처리하고 분석하는 데 필요한 기초 역량을 갖추게 되어, Machine Learning과 빅데이터 분석과 같은 고급 응용으로 나아갈 수 있습니다.동영상에는 실시간 필기가 포함되어 있으며, 동영상 왼쪽 하단의 "Show transcript"를 클릭하면 표시할 수 있습니다. 강의 용어집은 오른쪽 리소스 섹션에서 확인할 수 있습니다. CPE 학점 취득을 위해서는 과정을 완료하고 인증 평가에서 70% 이상의 점수를 받아야 합니다. 오른쪽의 CPE 학점 안내를 클릭하면 평가로 이동할 수 있습니다.

선수 조건

Introduction to SQL Data Manipulation with pandas

1

Introduction to Apache Spark and PySpark

A General introduction to PySpark and distributed computing. This section introduces PySpark, PySpark DataFrames, and RDDs.

Introduction to PySpark

Creating a SparkSession

Loading census data

Introduction to PySpark DataFrames

Scalability and performance

Reading a CSV and performing aggregations

Filtering by company

More on Spark DataFrames

Infer and filter

Schema writeout

2

PySpark in Python

A continuation of DataFrames and complex datatypes. This section expands on what DataFrames offer in PySpark and introduces some Spark SQL concepts.

Data manipulation with DataFrames

Handling missing data with fill and drop

Column operations - creating and renaming columns

Advanced DataFrame operations

DataFrame combinations

Joining flights with their destination airports

U define it? U use it!

UDF defined

Integers in PySpark UDFs

Pandas UDFs

3

Introduction to PySpark SQL

Delve into leveraging Spark SQL and PySpark for scalable data processing, combining SQL's simplicity with PySpark's distributed computing power to handle large datasets efficiently.

Resilient distributed datasets in PySpark

Creating RDDs

Collecting RDDs

Intro to Spark SQL

Querying on a temp view

Running SQL on DataFrames

Analytics with SQL on DataFrames

PySpark aggregations

Aggregating in PySpark

Aggregating in RDDs

Complex Aggregations

PySpark at scale

Broadcasting

Bringing it all together I

Bringing it all together II

What have we learned?

PySpark 입문

강의
완료

수료증 획득

LinkedIn 프로필, 이력서 또는 CV에 이 자격증을 추가하세요
소셜 미디어와 성과 평가에서 공유하세요지금 등록

19백만 명 이상의 학습자와 함께 PySpark 입문을(를) 시작하세요!

무료 계정을 만드세요

또는

계속 진행하시면 당사의 이용약관, 개인정보처리방침 및 귀하의 데이터가 미국에 저장되는 것에 동의하시는 것입니다.

DataCamp for Mobile을 통해 데이터 분석 능력을 향상시키세요.

모바일 강좌와 매일 5분 코딩 챌린지를 통해 이동 중에도 학습 효과를 높이세요.