Lead Engineer - Data Scientist
Current• Designed and implemented a real-time data pipeline to process semi-structured data by integrating 250 million raw records from 30+ data sources using Pyspark.• Designed an optimized data pipeline using Databricks, S3, and Azure, reducing data processing time by 30%. This enhancement enabled the DS team to achieve a 20% increase in decision-making efficiency.• Used PySpark to distribute data processing on large streaming datasets, improving ingestion and speed by 67%.• Utilized Scala to leverage its functional programming capabilities in optimizing data transformations and processing within the pipeline, resulting in a 15% reduction in code complexity and improved maintainability.• Integrated Natural Language Processing (NLP) tools like spaCy and NLTK for improved handling of unstructured text data, resulting in a 25% enhancement in data understanding.• Applied Reinforcement Learning algorithms to dynamically optimize resource allocation, reducing wastage by 20% and enhancing overall system efficiency.