Chuyển đến nội dung chính

Khóa học

AI Infrastructure: Networking Techniques

Trung cấp1 giờ

Design and deploy high-performance AI/ML solutions using Google Cloud's AI Hypercomputer, GPUs, TPUs, Compute, and Google Kubernetes Engine.

R1 giờ23 bài tập1,150 XP99Bản xác nhận thành tích

Tạo tài khoản miễn phí của bạn

Tiếp tục với Google
hoặc
Bằng cách tiếp tục, bạn chấp nhận Điều khoản sử dụng, Chính sách quyền riêng tư và việc dữ liệu của bạn được lưu trữ tại Hoa Kỳ.

Được yêu thích bởi người học tại hàng nghìn công ty

Đào tạo một nhóm?

Dùng thử cho Doanh nghiệp

Mô tả khóa học

Welcome to the "AI Infrastructure: Networking Techniques" course. In this course, you'll learn to leverage Google Cloud's high-bandwidth, low-latency infrastructure to optimize data transfer and communication between all the components of your AI system. By the end, you'll grasp the critical role networking plays across the entire AI pipeline from data ingestion and training to inference and be able to apply best practices to ensure your workloads run at maximum speed.

Yêu cầu tiên quyết

Khóa học này không có yêu cầu tiên quyết nào

Chương trình học

Đề cương khóa học

1

Course overview

This module offers an overview of the course and outlines the learning objectives.
Bắt đầu chương
2

Introduction

This module details the specialized networking requirements for AI workloads compared to traditional web applications. It covers the specific bandwidth and latency demands of each pipeline stage—from ingestion to inference—and analyzes the "rail-aligned" network architectures of Google Cloud's A3 and A4 GPU machine types designed to maximize "Goodput."
Bắt đầu chương
3

Networking for data ingestion

This module details strategies for efficiently moving massive datasets into the cloud. It covers the use of the Cross-Cloud Network and Cloud Interconnect to establish high-bandwidth pipelines, and outlines configuration best practices—such as enabling Jumbo Frames (MTU)—to reduce protocol overhead and optimize throughput.
Bắt đầu chương
4

Networking for AI training

This module details the critical role of low-latency networking in distributed model training. It covers the necessity of Remote Direct Memory Access (RDMA) for gradient synchronization, the benefits of Google's Titanium offload architecture in freeing up CPU resources, and the topology choices required to scale clusters without bottlenecks.
Bắt đầu chương
5

Networking for inference

This module details the networking challenges specific to Generative AI inference, such as bursty traffic and long-lived connections. It covers optimizing Time-to-First-Token using the GKE Inference Gateway and "Queue Depth" routing, while also addressing best practices for network reliability and Identity and Access Management (IAM).
Bắt đầu chương
6

Course Resources

Student PDF links to all modules
Bắt đầu chương
R

AI Infrastructure: Networking Techniques

Hoàn
thành khóa học

Nhận Giấy Chứng Nhận Hoàn Thành

Đăng ký ngay

Phát triển kỹ năng dữ liệu của bạn với DataCamp for Mobile

Tiến bộ mọi lúc mọi nơi với các khóa học di động và thử thách lập trình 5 phút mỗi ngày của chúng tôi.