Data Scientist
Current•Developed machine learning models (regression, classification, clustering) using scikit-learn, focusing on customer behavior, churn prediction, and anomaly detection.•Managed data pipelines and ETL processes, collaborating with data engineers to optimize SQL queries for Redshift and Hive.•Built NLP-based solutions for virtual assistants and search platforms, integrating ML models with web services to improve customer satisfaction.•Leveraged Azure Data Factory Pipelines for data orchestration and processing, enabling efficient, scalable data solutions.•Conducted feature engineering (normalization, label encoding) and data imputation to enhance model accuracy.•Enhanced data visualization capabilities using Power BI and Tableau, producing dashboards for executive decision-making.•Proficient in ETL processes, collaborating with data engineers and operations teams, and optimizing SQL queries for data extraction and analysis.•Skilled in data retrieval using Hive for Hadoop clusters and SQL for RedShift, as well as data analysis with Spark SQL.•Implemented NLP/NLU, to forecast Anomaly detection DS solution development and deployment in real-world scenarios.•Implemented Machine Learning, MLOps, MLflow, Kubeflow, Python/R, SQL, Big Data, GCP, and Shell scripting.•Scaled infrastructure to support high-throughput data-intensive applications using PySpark/GPU•Integrated ML models with web services•Built next-generation AI and Search platforms for the Client, enabling smart virtual assistants across multiple channels and platforms.•Proficient in Python (numpy, scipy, pandas, scikit-learn, seaborn) and Spark 2.0 (PySpark, MLlib) for model development and algorithm implementation.•Utilized natural language processing (NLP) techniques to optimize Customer Satisfaction.•Designed data visualizations for effective data representation using Tableau and Matplotlib.