After doing these courses, I feel confident creating professional visualizations and dashboards
Descrizione del corso
This course teaches you to design, build, and operate batch data pipelines on Google Cloud. Topics include large-scale data transformations with Dataflow and Serverless Spark, batch data validation and cleansing, schema evolution, error handling, and pipeline orchestration with Cloud Composer.
Prerequisiti
Non ci sono prerequisiti per questo corso
Curriculum
Programma del corso
1
When to choose batch data pipelines
You will learn the critical role of a data engineer in developing and maintaining batch data pipelines, understand their core components and lifecycle, and analyze common challenges in batch data processing. You'll also identify key Google Cloud services that address these challenges.
- Batch data pipelines and their use cases50 XP
- Processing and common challenges50 XP
- Module 1 Quiz: When to choose batch data pipelines — Question 150 XP
- Module 1 Quiz: When to choose batch data pipelines — Question 250 XP
- Module 1 Quiz: When to choose batch data pipelines — Question 350 XP
- Module 1 Quiz: When to choose batch data pipelines — Question 450 XP
- Module 1 Quiz: When to choose batch data pipelines — Question 550 XP
- Module 1 Quiz: When to choose batch data pipelines — Question 650 XP
- Module 1 Quiz: When to choose batch data pipelines — Question 750 XP
- Module 1 Quiz: When to choose batch data pipelines — Question 850 XP
- Module 1 Quiz: When to choose batch data pipelines — Question 950 XP
- Module 1 Quiz: When to choose batch data pipelines — Question 1050 XP
2
Design and build batch data pipelines
You will design scalable batch data pipelines for high-volume data ingestion and transformation. You'll also optimize batch jobs for high throughput and cost-efficiency using various resource management and performance tuning techniques.
3
Control data quality in batch data pipelines
You will develop data validation rules and cleansing logic to ensure data quality within batch pipelines. You'll also implement strategies for managing schema evolution and performing data deduplication in large datasets.
4
Orchestrate and monitor batch data pipelines
You will orchestrate complex batch data pipeline workflows for efficient scheduling and lineage tracking. You'll also implement robust error handling, monitoring, and observability for batch data pipelines.
R
Build Batch Data Pipelines on Google Cloud
Corso
completato

