Data Scientist
Current• Independently built a Django-based website to display client’s portfolio financial risk analysis report relative to multiple benchmarks, leveraging SQL and Python for ETL processes from Snowflake, MongoDB, and AWS S3, and utilized JavaScript libraries (Plotly, ApexCharts, Grid.js) and Tableau for dynamic data visualization• Engineered a system for daily extraction of news articles related to a PE client’s portfolio, leveraging an XGBoost model for relevance filtering, categorization, and industry classification prediction for unclassified companies• Developed automated ETL processes across disparate PE data sources to track 100+ PE funds’ near-term liquidity• Constructed a robust ETL pipeline to identify and extract new portfolio companies, standardized company identifiers across multiple data sources, and streamlined data extraction workflows from diverse databases• Developed a sentiment analysis pipeline for over 1,000 companies, integrating alternative data sourced from Google Trends and Alexa Rank APIs, and employed ARIMA and ADTK models to generate trendlines and detect anomalies