Applied MathApplied Math & Data ScienceApplied ScienceArchitecture / Masters in ArchitectureBiological EngineeringBiomedical EngineeringBusiness ManagementBusiness Management/AdministrationCivil EngineeringComputer EngineeringComputer Information SystemsComputer NetworkingComputer ScienceComputing, Cybersecurity, & Information TechnologyConstruction ManagementCybersecurityDesign/BuildElectrical EngineeringElectromechanical EngineeringEngineering

Cleaning Data for Effective Data Science: Data Ingestion, Anomaly Detection, Value Imputation, and Feature Engineering

Description

What is this course about?

The course introduces the tools and techniques needed for data ingestion, anomaly detection, value imputation, and feature engineering. Numerous ingested formats are addressed, including JSON, CSV, SQL RDBMS, HDF5, NoSQL databases, and binary serialized data structures. Instructor David Mertz outlines why some problems are peculiar to data representation, while others link to the data in itself. To address untidiness in data, learn how and when to impute missing values, detect unreliable data and statistical anomalies, and generate synthetic features that are necessary for successful data analysis and visualization goals. By the end of this course, you’ll be equipped with highly marketable and in-demand skills in data analysis, machine learning, and data integrity troubleshooting.

Note: This course was created by Pearson. We are pleased to host this training in our library.

Instructor

Who teaches this course?

David Mertz, PhD, is a data scientist, author, and former Python Foundation director and Anaconda senior trainer.

Objectives

What will I be able to do by the end of this course?

  • Analyze and process various data formats including tabular and hierarchical.
  • Detect and correct data anomalies and biases effectively.
  • Implement data ingestion across diverse formats such as JSON and CSV.
  • Apply value imputation techniques tailored to specific analytical purposes.
  • Engineer data features to enhance machine learning model performance.

Audience

Who is this course for?

  • Database administrators
  • Data scientists
  • Data analysts

Prerequisites

What do I need to know before taking this course?

  • Basic understanding of data structures and formats
  • Familiarity with data science principles and tools
Learn More