본문으로 바로가기

강의

PySpark로 데이터 정제하기

고급기술 수준

업데이트됨 2026. 2.

Python에서 Apache Spark로 데이터 정제 방법을 학습하세요.

무료로 강의 시작

SparkData Preparation

4시간

16 동영상

53 연습 문제

4,150 XP

33,173

성취 증명서

수천 개 기업의 학습자들이 사랑하는

팀을 교육하시나요?

비즈니스용으로 체험해 보세요

강의 설명

데이터를 다루는 일은 어렵고, 수백만에서 수십억 행을 다루면 더 복잡해집니다. 상대적으로 깔끔한 데이터로 노트북에서 작성된 데이터 처리 코드를 받으셨나요? 프로토타입 수준의 데이터 프로세스를 운영 환경으로 이전하는 일을 맡게 될 가능성이 큽니다. 누락된 필드, 기묘한 형식, 그리고 데이터 규모가 몇 자릿수나 더 큰 실제 데이터셋을 다뤄 보셨을 수도 있어요. 이런 모든 것이 처음이라도, 이 과정을 통해 Apache Spark와 Python으로 데이터 프로세스를 준비하는 데 필요한 내용을 배울 수 있습니다. 성능이 뛰어나고 유지보수 가능하며 이해하기 쉬운 데이터 처리 플랫폼을 만들기 위한 용어, 방법, 모범 사례를 익히게 됩니다.

선수 조건

Intermediate Python Introduction to PySpark

1

DataFrame details

A review of DataFrame fundamentals and the importance of data cleaning.

Intro to data cleaning with Apache Spark

Data cleaning review

Defining a schema

Immutability and lazy processing

Immutability review

Using lazy processing

Understanding Parquet

Saving a DataFrame in Parquet format

SQL and Parquet

2

Manipulating DataFrames in the real world

A look at various techniques to modify the contents of DataFrames in Spark.

DataFrame column operations

Filtering column content with Python

Filtering Question #1

Filtering Question #2

Modifying DataFrame columns

Conditional DataFrame column operations

when() example

When / Otherwise

User defined functions

Understanding user defined functions

Using user defined functions in Spark

Partitioning and lazy processing

Adding an ID Field

IDs with different partitions

More ID tricks

3

Improving Performance

Improve data cleaning tasks by increasing performance or reducing resource requirements.

Caching a DataFrame

Removing a DataFrame from cache

Improve import performance

File size optimization

File import performance

Cluster configurations

Reading Spark configurations

Writing Spark configurations

Performance improvements

Normal joins

Using broadcasting on Spark joins

Comparing broadcast vs normal joins

4

Complex processing and data pipelines

Learn how to process complex real-world data using Spark and the basics of pipelines.

Introduction to data pipelines

Quick pipeline

Pipeline data issue

Data handling techniques

Removing commented lines

Removing invalid rows

Splitting into columns

Further parsing

Data validation

Validate rows via join

Examining invalid rows

Final analysis and delivery

Dog parsing

Per image count

Percentage dog pixels

Congratulations and next steps

PySpark로 데이터 정제하기

강의
완료

수료증 획득

LinkedIn 프로필, 이력서 또는 CV에 이 인증서를 추가하세요
소셜 미디어와 성과 평가에서 공유하세요지금 등록

19백만 명 이상의 학습자와 함께 PySpark로 데이터 정제하기을(를) 시작하세요!

DataCamp for Mobile을 통해 데이터 분석 능력을 향상시키세요.

모바일 강좌와 매일 5분 코딩 챌린지를 통해 이동 중에도 학습 효과를 높이세요.