Data Engineer
Current• Responsible for setting up technical and functional requirements, data pipelines, data preparation, design, development, modelling, testing and deployment of advanced models in the AWS cloud.• Conducted exhaustive study on various available Transformers models and estimated costing, system load, model performance, cost to performance metrics to decide on the best model that would yield the best results in terms of Development, Optimization and Maintenance parameters.• Exported data from Snowflake to S3 and created on-demand tables using AWS Lambda Functions and AWS Glue using Python and PySpark.• Designed and implemented ETL pipelines on S3 files on data lake using AWS Glue.• Leveraged Pandas, NumPy and other preprocessing libraries for Data Cleaning and Feature Engineering to incorporate the same into the pipeline component of the NLP workflow.• Worked with 7TB of Unstructured Data for a classification problem and built NLP Pipelines which includes preprocessing, modeling and testing components.• Performed text preprocessing like Tokenization, Lemmatization, etc. and Feature Engineering/Extraction on multi label imbalanced data followed by applying various sampling techniques.• Built SpaCy Huggingface transformers models like BERT, ALBERT and RoBERT with de-identified transcript.• Built metrics like Confusion Matrix and Model Performance via dashboards and integrated into sisence for real time visibility.• Finetuned BERT models on de-identified client data to streamline development & deployment in cloud.• Saved $500,000 in costs to the client by developing the advanced models for setting up processing in place for CI/CD.• Achieved an overall Model Accuracy of 93% which was a 25%-point increase from old models resulting in Significant Cost Savings to the client.Environment: AWS Cloud, Python, SQL, ETL, jupyter notebook, Pyspark, Pandas, Snowflake, Sagemaker, Huggingface, SpaCy, Transfer learning, Git, JIRA, Agile, Windows and Linux.