본문으로 건너뛰기

강의

AI Infrastructure: Networking Techniques

중급1시간

Design and deploy high-performance AI/ML solutions using Google Cloud's AI Hypercomputer, GPUs, TPUs, Compute, and Google Kubernetes Engine.

R1시간연습 문제 23개1,150 XP109수료 증명서

무료 계정 만들기

Google에서 계속 진행
또는
계속 진행하시면 다음 사항에 동의하는 것으로 간주됩니다: 이용약관개인정보 처리방침 그리고 귀하의 데이터가 미국에 저장된다는 점에 동의하는 것으로 간주됩니다.

수천 개 기업의 학습자들이 사랑하는 서비스

강의 설명

Welcome to the "AI Infrastructure: Networking Techniques" course. In this course, you'll learn to leverage Google Cloud's high-bandwidth, low-latency infrastructure to optimize data transfer and communication between all the components of your AI system. By the end, you'll grasp the critical role networking plays across the entire AI pipeline from data ingestion and training to inference and be able to apply best practices to ensure your workloads run at maximum speed.

선수 과목

이 강의에는 선수 과목이 없습니다

커리큘럼

강의 개요

1

Course overview

This module offers an overview of the course and outlines the learning objectives.
챕터 시작
2

Introduction

This module details the specialized networking requirements for AI workloads compared to traditional web applications. It covers the specific bandwidth and latency demands of each pipeline stage—from ingestion to inference—and analyzes the "rail-aligned" network architectures of Google Cloud's A3 and A4 GPU machine types designed to maximize "Goodput."
챕터 시작
3

Networking for data ingestion

This module details strategies for efficiently moving massive datasets into the cloud. It covers the use of the Cross-Cloud Network and Cloud Interconnect to establish high-bandwidth pipelines, and outlines configuration best practices—such as enabling Jumbo Frames (MTU)—to reduce protocol overhead and optimize throughput.
챕터 시작
4

Networking for AI training

This module details the critical role of low-latency networking in distributed model training. It covers the necessity of Remote Direct Memory Access (RDMA) for gradient synchronization, the benefits of Google's Titanium offload architecture in freeing up CPU resources, and the topology choices required to scale clusters without bottlenecks.
챕터 시작
5

Networking for inference

This module details the networking challenges specific to Generative AI inference, such as bursty traffic and long-lived connections. It covers optimizing Time-to-First-Token using the GKE Inference Gateway and "Queue Depth" routing, while also addressing best practices for network reliability and Identity and Access Management (IAM).
챕터 시작
6

Course Resources

Student PDF links to all modules
챕터 시작
R

AI Infrastructure: Networking Techniques

강의
완료

수료증 획득

지금 등록하기

DataCamp for Mobile로 데이터 역량을 키우세요

모바일 강의와 매일 5분 코딩 챌린지로 이동 중에도 학습을 이어가세요.