Accéder au contenu principal

cours

Feature Engineering for Machine Learning in Python

Intermédiaire

Updated 12/2024

Create new features to improve the performance of your Machine Learning models.

Commencer le cours gratuitement

Inclus gratuitementPremium or Teams

PythonMachine learning4 heures16 vidéos53 exercices4,350 XP31,686Déclaration de réalisation

Créez votre compte gratuit

Google LinkedIn Facebook

ou

En continuant, vous acceptez nos Conditions d'utilisation, notre Politique de confidentialité et le fait que vos données sont stockées aux États-Unis.

Formation de 2 personnes ou plus ?

Essayer DataCamp for Business

Apprécié par les apprenants de milliers d’entreprises

Description du cours

Every day you read about the amazing breakthroughs in how the newest applications of machine learning are changing the world. Often this reporting glosses over the fact that a huge amount of data munging and feature engineering must be done before any of these fancy models can be used. In this course, you will learn how to do just that. You will work with Stack Overflow Developers survey, and historic US presidential inauguration addresses, to understand how best to preprocess and engineer features from categorical, continuous, and unstructured data. This course will give you hands-on experience on how to prepare any data for your own machine learning models.

Conditions préalables

Supervised Learning with scikit-learn

1

Creating Features

Commencer le chapitre

Why generate features?

Getting to know your data

Selecting specific data types

Dealing with categorical features

One-hot encoding and dummy variables

Dealing with uncommon categories

Numeric variables

Binarizing columns

Binning values

2

Dealing with Messy Data

Commencer le chapitre

Why do missing values exist?

How sparse is my data?

Finding the missing values

Dealing with missing values (I)

Listwise deletion

Replacing missing values with constants

Dealing with missing values (II)

Filling continuous missing values

Imputing values in predictive models

Dealing with other data issues

Dealing with stray characters (I)

Dealing with stray characters (II)

Method chaining

3

Conforming to Statistical Assumptions

Commencer le chapitre

Data distributions

What does your data look like? (I)

What does your data look like? (II)

When don't you have to transform your data?

Scaling and transformations

Normalization

Standardization

Log transformation

When can you use normalization?

Removing outliers

Percentage based outlier removal

Statistical outlier removal

Scaling and transforming new data

Train and testing transformations (I)

Train and testing transformations (II)

4

Dealing with Text Data

Commencer le chapitre

Encoding text

Cleaning up your text

High level text features

Word counts

Counting words (I)

Counting words (II)

Limiting your features

Text to DataFrame

Term frequency-inverse document frequency

Inspecting Tf-idf values

Transforming unseen data

Using longer n-grams

Finding the most common words

Feature Engineering for Machine Learning in Python

Cours
terminé

Earn Déclaration de réalisation

Ajoutez ces informations d’identification à votre profil LinkedIn, à votre CV ou à votre CV
Partagez-le sur les réseaux sociaux et dans votre évaluation de performance

Inclus avecPremium or Teams

S'inscrire maintenant

Inscrivez-vous 15 millions d’apprenants et commencer Feature Engineering for Machine Learning in Python Aujourd’hui!

Créez votre compte gratuit

Google LinkedIn Facebook

ou

En continuant, vous acceptez nos Conditions d'utilisation, notre Politique de confidentialité et le fait que vos données sont stockées aux États-Unis.