Skip to content

About

Application of pandas in python using the lung cancer dataset

Resources

Stars

0 stars

Watchers

1 watching

Forks

Repository files navigation

Lung Cancer Prediction Project

This project is a personal exploration of a lung cancer dataset to uncover patterns and relationships between demographic, behavioral, environmental, and health factors, and their influence on lung cancer. I undertook this project as part of my data science learning journey, applying Python and Pandas to clean, explore, and interpret the data.

Dataset Overview

The dataset comprises several features that may contribute to lung cancer risk, including:

  • Demographic Factors: Age, Gender
  • Behavioral Factors: Smoking, Alcohol Consumption, Family History of Smoking, Mental Stress
  • Environmental Factors: Exposure to Pollution
  • Health Indicators: Breathing Issues, Chest Tightness, Oxygen Saturation, Immune Weakness, Finger Discoloration, Energy Level
  • Target Variable: Lung Cancer Presence

Analysis Objectives

My goal was to dive deep into the dataset and answer the following questions:

  1. How is lung cancer prevalence distributed across different age groups?
  2. Are smoking habits significantly linked to lung cancer cases?
  3. What role do environmental factors like pollution play in lung cancer prevalence?
  4. Do health indicators show noticeable patterns in individuals with lung cancer?

Steps Performed

  1. Data Loading: Imported the dataset into a Pandas DataFrame.
  2. Data Cleaning: Addressed missing values, standardized column names for consistency, and ensured correct data types.
  3. Exploratory Data Analysis (EDA): Conducted a detailed exploration to identify trends and correlations. This included generating descriptive statistics and visualizing distributions.
  4. Feature Analysis: Analyzed each feature's relationship with lung cancer presence, identifying key contributors.
  5. Insights and Visualization: Created visualizations to better understand feature distributions and relationships with lung cancer.

Tools and Libraries

Throughout this project, I used the following tools and libraries:

  • Python
  • Pandas
  • Matplotlib
  • Seaborn

How to Run the Project

If you’d like to run the project on your own machine, follow these steps:

  1. Clone the repository:
git clone <repository_url>
  1. Install required libraries:
pip install pandas matplotlib seaborn
  1. Open the Jupyter Notebook:
jupyter notebook "Lung Cancer Prediction Project using pandas.ipynb"
  1. Run the cells sequentially to view the full analysis.

Results and Insights

Some key insights I discovered:

  • Smoking Habits: A strong correlation exists between smoking and lung cancer presence.
  • Age Groups: Older individuals show a higher prevalence of lung cancer.
  • Environmental Factors: Pollution exposure also seems to play a role, with higher exposure correlating with increased lung cancer risk.
  • Health Indicators: Breathing issues and chest tightness are frequent among individuals diagnosed with lung cancer.

Lessons Learned

This project allowed me to deepen my understanding of data cleaning and exploratory data analysis using Pandas. I also improved my skills in visualizing data to draw meaningful insights.

About

Application of pandas in python using the lung cancer dataset

Resources

Stars

0 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages