ข้ามไปยังเนื้อหาหลัก

หน้าหลัก Python

คอร์ส

Feature Engineering for NLP in Python

ขั้นสูงระดับทักษะ

อัปเดตแล้ว 11/2567

Learn techniques to extract useful information from text and process them into a format suitable for machine learning.

เริ่มคอร์สฟรี

PythonMachine Learning

4 ชม.

15 วิดีโอ

52 แบบฝึกหัด

4,200 XP

29,225

ใบรับรองความสำเร็จ

สร้างบัญชีฟรีของคุณ

ดำเนินการต่อด้วย Google แสดงตัวเลือกเพิ่มเติม

หรือ

เมื่อดำเนินการต่อ คุณยอมรับ ข้อกำหนดการใช้งาน ของเรา นโยบายความเป็นส่วนตัว ของเรา และยอมรับว่าข้อมูลของคุณจะถูกจัดเก็บในสหรัฐอเมริกา

เป็นที่รักของผู้เรียนในบริษัทหลายพันแห่ง

กำลังฝึกอบรมทีม?

ลองใช้สำหรับธุรกิจ

คำอธิบายคอร์ส

In this course, you will learn techniques that will allow you to extract useful information from text and process them into a format suitable for applying ML models. More specifically, you will learn about POS tagging, named entity recognition, readability scores, the n-gram and tf-idf models, and how to implement them using scikit-learn and spaCy. You will also learn to compute how similar two documents are to each other. In the process, you will predict the sentiment of movie reviews and build movie and Ted Talk recommenders. Following the course, you will be able to engineer critical features out of any text and solve some of the most challenging problems in data science!

ข้อกำหนดเบื้องต้น

Introduction to Natural Language Processing in Python Supervised Learning with scikit-learn

1

Basic features and readability scores

Learn to compute basic features such as number of words, number of characters, average word length and number of special characters (such as Twitter hashtags and mentions). You will also learn to compute readability scores and determine the amount of education required to comprehend a piece of text.

Introduction to NLP feature engineering

Data format for ML algorithms

One-hot encoding

Basic feature extraction

Character count of Russian tweets

Word count of TED talks

Hashtags and mentions in Russian tweets

Readability tests

Readability of 'The Myth of Sisyphus'

Readability of various publications

เริ่มบท

2

Text preprocessing, POS tagging and NER

In this chapter, you will learn about tokenization and lemmatization. You will then learn how to perform text cleaning, part-of-speech tagging, and named entity recognition using the spaCy library. Upon mastering these concepts, you will proceed to make the Gettysburg address machine-friendly, analyze noun usage in fake news, and identify people mentioned in a TechCrunch article.

Tokenization and Lemmatization

Identifying lemmas

Tokenizing the Gettysburg Address

Lemmatizing the Gettysburg address

Text cleaning

Cleaning a blog post

Cleaning TED talks in a dataframe

Part-of-speech tagging

POS tagging in Lord of the Flies

Counting nouns in a piece of text

Noun usage in fake news

Named entity recognition

Named entities in a sentence

Identifying people mentioned in a news article

เริ่มบท

3

N-Gram models

Learn about n-gram modeling and use it to perform sentiment analysis on movie reviews.

Building a bag of words model

Word vectors with a given vocabulary

BoW model for movie taglines

Analyzing dimensionality and preprocessing

Mapping feature indices with feature names

Building a BoW Naive Bayes classifier

BoW vectors for movie reviews

Predicting the sentiment of a movie review

Building n-gram models

n-gram models for movie tag lines

Higher order n-grams for sentiment analysis

Comparing performance of n-gram models

เริ่มบท

4

TF-IDF and similarity scores

Learn how to compute tf-idf weights and the cosine similarity score between two vectors. You will use these concepts to build a movie and a TED Talk recommender. Finally, you will also learn about word embeddings and using word vector representations, you will compute similarities between various Pink Floyd songs.

Building tf-idf document vectors

tf-idf weight of commonly occurring words

tf-idf vectors for TED talks

Cosine similarity

Range of cosine scores

Computing dot product

Cosine similarity matrix of a corpus

Building a plot line based recommender

Comparing linear_kernel and cosine_similarity

Plot recommendation engine

The recommender function

TED talk recommender

Beyond n-grams: word embeddings

Generating word vectors

Computing similarity of Pink Floyd songs

Congratulations!

เริ่มบท

Feature Engineering for NLP in Python

คอร์สเสร็จสมบูรณ์

รับใบรับรองความสำเร็จ

เพิ่มใบรับรองนี้ไปยังโปรไฟล์ LinkedIn เรซูเม่ หรือ CV ของคุณ
แชร์บน social media และในการรีวิวผลการปฏิบัติงานของคุณลงทะเบียนทันที

สำหรับธุรกิจ

ฝึกอบรม 2 คนขึ้นไปหรือไม่?

ให้ทีมของคุณเข้าถึงแพลตฟอร์ม DataCamp เต็มรูปแบบ รวมถึงฟีเจอร์ทั้งหมด

ในเส้นทางการเรียนต่อไปนี้

นักวิทยาศาสตร์การเรียนรู้ของเครื่อง ใน Python

การประมวลผลภาษาธรรมชาติ ใน Python

ผู้สอน

Rounak Banik

Rounak Banik

Data Scientist at Fractal Analytics

ผู้ร่วมงาน

คอร์ส แหล่งข้อมูล

Russian Troll Tweetsชุดข้อมูล

Movie Overviews and Taglinesชุดข้อมูล

Preprocessed Movie Reviewsชุดข้อมูล

TED Talk Transcriptsชุดข้อมูล

Real and Fake News Headlinesชุดข้อมูล

ร่วมกับผู้เรียนกว่า 19 ล้านคนและเริ่มต้น Feature Engineering for NLP in Python วันนี้!

สร้างบัญชีฟรีของคุณ

ดำเนินการต่อด้วย Google แสดงตัวเลือกเพิ่มเติม

หรือ

เมื่อดำเนินการต่อ คุณยอมรับ ข้อกำหนดการใช้งาน ของเรา นโยบายความเป็นส่วนตัว ของเรา และยอมรับว่าข้อมูลของคุณจะถูกจัดเก็บในสหรัฐอเมริกา

พัฒนาทักษะด้านข้อมูลของคุณด้วย DataCamp for Mobile

พัฒนาทักษะได้ทุกที่ทุกเวลาด้วยคอร์สเรียนบนมือถือและแบบฝึกหัดเขียนโค้ดประจำวัน 5 นาทีของเรา