メインコンテンツへスキップ

ホーム Spark

コース

PySpark でデータをクレンジングする

上級スキルレベル

更新日 2026/02

PythonでApache Sparkを使い、データのクリーニング手法を学びます。

コースを無料で開始

SparkData Preparation

4時間

16 ビデオ

53 演習

4,150 XP

33,173

修了証明書

何千もの企業の従業員が支持

チームのトレーニングを担当していますか？

Businessをお試しください

コース説明

データを扱うのは難しいものです。ましてや数百万、数十億行規模となるとさらに大変です。きれいなデータを前提にノートPC上で書かれたデータ処理コードを受け取りましたか？おそらく、プロトタイプのデータ処理を本番へ移行する役割を任されたことがあるのではないでしょうか。欠損値や奇妙な書式、そして桁違いのデータ量を含む実世界のデータセットに取り組んだことがあるかもしれません。これが初めてでも、このコースでは、Apache Spark と Python を使ってデータ処理を準備するために必要なことを学べます。用語、手法、そして高性能で保守しやすく、理解しやすいデータ処理基盤を作るためのベストプラクティスを学習します。

前提条件

Intermediate Python Introduction to PySpark

1

DataFrame details

A review of DataFrame fundamentals and the importance of data cleaning.

Intro to data cleaning with Apache Spark

Data cleaning review

Defining a schema

Immutability and lazy processing

Immutability review

Using lazy processing

Understanding Parquet

Saving a DataFrame in Parquet format

SQL and Parquet

チャプターを開始

2

Manipulating DataFrames in the real world

A look at various techniques to modify the contents of DataFrames in Spark.

DataFrame column operations

Filtering column content with Python

Filtering Question #1

Filtering Question #2

Modifying DataFrame columns

Conditional DataFrame column operations

when() example

When / Otherwise

User defined functions

Understanding user defined functions

Using user defined functions in Spark

Partitioning and lazy processing

Adding an ID Field

IDs with different partitions

More ID tricks

チャプターを開始

3

Improving Performance

Improve data cleaning tasks by increasing performance or reducing resource requirements.

Caching a DataFrame

Removing a DataFrame from cache

Improve import performance

File size optimization

File import performance

Cluster configurations

Reading Spark configurations

Writing Spark configurations

Performance improvements

Normal joins

Using broadcasting on Spark joins

Comparing broadcast vs normal joins

チャプターを開始

4

Complex processing and data pipelines

Learn how to process complex real-world data using Spark and the basics of pipelines.

Introduction to data pipelines

Quick pipeline

Pipeline data issue

Data handling techniques

Removing commented lines

Removing invalid rows

Splitting into columns

Further parsing

Data validation

Validate rows via join

Examining invalid rows

Final analysis and delivery

Dog parsing

Per image count

Percentage dog pixels

Congratulations and next steps

チャプターを開始

PySpark でデータをクレンジングする

コース完了

修了証明書を取得

この修了書をLinkedInや履歴書、CVに追加しましょう
ソーシャルメディアや人事評価で共有しましょう今すぐ登録

19百万人を超える学習者と共にPySpark でデータをクレンジングするを始めましょう！

DataCamp for Mobileでデータスキルを磨きましょう

モバイルコースと毎日の 5 分間のコーディングチャレンジで、外出先でも進歩できます。