Skip to content

mishael-fav/DATA_CLEANING_AND_PREPARATION

Repository files navigation

🧹 HNG Internship Stage 1 – Data Cleaning & Preparation

This project focuses on cleaning, preparing, and optimizing a raw product dataset to ensure data quality for marketing analytics. The primary goal was to enhance the dataset's structure, handle inconsistencies, remove noise, and engineer new features for better SEO and product visibility.


🧾 Dataset Overview

  • Total Entries: 3,847
  • Final Cleaned Entries: 3,541
  • Key Features:
    • product_id: Unique product identifier
    • title: Original product title
    • bullet_points: Product features (often incomplete)
    • description: Detailed product description
    • producttype_id: Categorical identifier for product type
    • Product_Length: Numeric field (length in meters, assumed)

🎯 Objectives

  • 🗑 Remove Duplicates: Ensure unique entries using product_id
  • Handle Missing Values: Fill gaps in descriptions and bullet points
  • 🔧 Standardize Formats: Normalize column names and datatypes
  • 🚨 Detect & Treat Outliers: Apply IQR method to rectify extreme Product_Length values
  • 📝 Feature Engineering: Generate a short_title to enhance SEO
  • 📊 Visualize Improvements: Before/after comparisons using histograms and boxplots

🔧 Cleaning & Preprocessing Steps

✅ Duplicate Handling

  • Duplicates Found: 306
  • Method: .duplicated(subset='product_id')
  • Result: Cleaned dataset has 3,541 unique entries

✅ Missing Values

  • Columns affected: bullet_points, description
  • Strategy: Replaced with empty strings ""

✅ Standardization

  • Renamed columns to snake_case for consistency
  • Converted producttype_id to categorical
  • Rounded Product_Length to 3 decimal places

📏 Outlier Detection & Treatment

Product_Length

  • Max Value: 96,000 (extreme)
  • Method: IQR-based filtering
  • Action: Replaced outliers with median value
  • Result: Histogram and boxplot showed improved distribution

✂️ Title Optimization – short_title Feature

Generated a new feature short_title for better readability and SEO using:

  • Truncation to 6 words if over 12
  • Removing phrases after commas, "for", or "with"
  • Filtering special characters
  • Removing redundant or low-information words

🔠 Before vs After:

  • ⬇️ Reduced average title length
  • 🔤 Clearer, concise product names
  • 📈 Improved keyword frequency profile

📊 Visualizations Included

  • 📦 Histogram & Boxplot: Title lengths before and after cleaning Title length
  • 🗂 Bar Chart: Word frequency before vs after cleaning Frequency of words
  • 🧾 Category Distribution: Top 5 product types frequent product

About

This project focuses on cleaning, preparing, and optimizing a raw product dataset to ensure data quality for marketing analytics. The primary goal was to enhance the dataset's structure, handle inconsistencies, remove noise, and engineer new features for better SEO and product visibility.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

No releases published

Packages

 
 
 

Contributors