Data Scientist
○ Gathered and cleaned 3,500+ data points from Reddit and Twitter, effectively addressing 95% of missing values and eliminating 85% of irrelevant information.○ Executed comprehensive Exploratory Data Analysis (EDA) to uncover trends and patterns, leading to the discovery of 5 key insights that guided subsequent analysis.○ Utilized Natural Language Processing (NLP) techniques to preprocess text data, including tokenization and lemmatization, reducing text noise by 40% and enhancing model… Show more ○ Gathered and cleaned 3,500+ data points from Reddit and Twitter, effectively addressing 95% of missing values and eliminating 85% of irrelevant information.○ Executed comprehensive Exploratory Data Analysis (EDA) to uncover trends and patterns, leading to the discovery of 5 key insights that guided subsequent analysis.○ Utilized Natural Language Processing (NLP) techniques to preprocess text data, including tokenization and lemmatization, reducing text noise by 40% and enhancing model accuracy.○ Engineered an unsupervised learning approach using the BERTopic framework to cluster sentiments, achieving a coherent and interpretable topic structure, which improved topic coherence by 15%.○ Created a BERTopic pipeline integrating word embeddings, clustering, and an LLM for topic representation, reducing training time by 25%.○ Deployed the Voyage AI model for embeddings, k-means for clustering, and implemented LLaMA2 using Transformer and PyTorch. LLaMA2 was used for topic representation, boosting model efficiency by 20% and predictive accuracy by 10%.○ Achieved a coherence score of 0.58, indicating strong interpretability and relevance.○ Partnered with domain experts to validate the accuracy of clustered topics and sentiments, ensuring alignment with real-world contexts. Communicated findings and implications through detailed visualizations to stakeholders. Show less