Skip to content

About

A clean and structured stand-up comedian dataset for NLP, machine learning, and data analysis projects. Suitable for sentiment analysis, humor detection, text classification, and educational research.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

 

History

3 Commits

Folders and files

Repository files navigation

standupcomedian

🎤 Stand-Up Comedy Transcript Analysis

What can data tell us about what makes us laugh?

I analyzed transcripts from 18 stand-up comedy specials (Dave Chappelle, Trevor Noah, Pete Davidson, Ricky Gervais, and more) to uncover each comedian's favorite topics, their emotional "arc" through a set, and what their word choices reveal about their comedic style.

🔍 What I Did

  1. Text Cleaning — stripped stage directions ([music playing]), punctuation, and numbers using regex
  2. Word Frequency Analysis — used CountVectorizer + custom stopwords (removing filler words like "like", "know", "just") to surface each comedian's real signature topics
  3. Sentiment Arc Analysis — split each transcript into 10 segments and tracked sentiment polarity (TextBlob) across the set, to see how mood shifts from opening to punchline to close
  4. Word Clouds — visualized each comedian's most distinctive vocabulary

📊 Key Findings

  • 🎭 Dave Chappelle's top words — black, white, life, mask, coronavirus — confirm his sets center heavily on race and social commentary
  • 😂 Trevor Noah & Vir Das maintain the most consistently positive tone throughout their sets — rarely dipping below neutral
  • 🎢 Some comedians (like Ricky Gervais) follow a classic "dip and recover" arc — starting warm, going dark mid-set, then ending on a high note
  • 🤬 Raw, high-energy comedians (Marlon Wayans, Mike Epps) show the widest sentiment swings — big peaks and big drops within a single se

🛠️ Tools Used

Python · Pandas · Matplotlib · Scikit-learn (CountVectorizer) · TextBlob · WordCloud

🚀 What's Next

Planning to extend this with topic modeling (LDA) to auto-cluster comedians by theme, and a "joke density" metric.


💬 Feedback welcome! Connect with me on LinkedIn or check out more of my projects on GitHub.

About

A clean and structured stand-up comedian dataset for NLP, machine learning, and data analysis projects. Suitable for sentiment analysis, humor detection, text classification, and educational research.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages