Data Engineer
Current* Responsible for building end to end data pipelines in AWS cloud infrastructure.* Responsible for fine tuning, troubleshooting and supporting the enterprise data pipelines at production scale.* Written custom UDFs in spark to perform data encryption, date transformations and other complex business transformations.* Involved in creating external Hive tables from the files stored in the S3.* Optimized Hive tables utilizing partitions and bucketing to give better execution Hive QL queries. Used Spark-SQL to read data from hive tables and perform various data cleansing, data validations, transformations and aggregations as per downstream business team requirements.* Created views in Athena to allow secure and streamlined data analysis access to downstream business teams.* Worked extensively with the Data Science team to help productionalize machine learning models and also to build various feature datasets as needed for data analysis and modelling.* Worked extensively in automating launching of EMR clusters and terminating EMR clusters as soon as jobs are finished.* Worked on data visualization and analytics with research scientists and business stakeholders.Technologies: Spark, Kafka, AWS S3, EMR, Redshift, Athena, Glue Metastore, Hive, Java