Data Scientist
CurrentElectric Vehicle Range PredictionDeveloped a data ingestion pipeline for real-time data migration from diverse sources to S3, leveraging RabbitMQ and KinesisEngineered and executed a comprehensive ETL pipeline to extract and transform data from S3 using AWS GlueConducted EDA on transformed data, addressed missing values and outliers, and gained insights through visualization in SageMaker Applied k-means clustering to discern driver behaviors, route profiling, and environmental and operational conditionsSignificantly enhanced accuracy by 10-15% with a blended machine learning model utilizing boosting algorithms like XGBoost and LightGBM, predicting electric vehicle range compared to the existing BMS model accuracy, implemented with PySpark. Data Quality AssessmentIngested a substantial volume of vehicle data from Snowflake utilizing the Snowflake Spark connectorDefined customized rules in Python to assess data quality across dimensions such as completeness, consistency, and integrity. Developed a Tableau dashboard to provide users with an overview of data quality across dimensions and rules.Introduced a refined approach for in-depth data analysis through a targeted drill-down strategy to examine violated records. Orchestrated the entire workflow in EMR Serverless, automating and deploying it seamlessly using Terraform & and GitHubActions.Lead-Acid Battery Drain AnalysisPre-processed and synchronized sparse battery data, applying time synchronization for trip summary generation using pythonConducted univariate and bivariate analyses to examine trigger occurrences on the relative and absolute vehicle data timelineCrafted a Tableau dashboard for visualizing trigger event distribution and their causal relationships, leveraging Gantt charts.