Senior Big Data Engineer
Current• Involved in building a data pipeline and performed analytics using AWS stack (IAM,EMR, EC2, S3, RDS, Lambda, Athena, Glue, SQS, Redshift, and ECS). • Developed Spark applications utilizing Pyspark and Spark-SQL for information extraction, change, and accumulation from numerous document designs for analyzing and changing the information for client utilization designs. • Developed Spark jobs on Databricks to perform tasks like data cleansing, data validation, standardization, and then applied transformations as per the use cases.• Used AWS Redshift, S3, Spectrum and Athena services to query large amount data stored on S3 to create a Virtual Data Lake.• Created Lambda flow that created data pipeline jobs that loads data to RDS to execute stored procedures and copy the data to Redshift. • Designed and Developed ETL Processes in AWS Glue to migrate Campaign data from external sources like S3, ORC/Parquet/Text Files into AWS Redshift. • Used Spark-streaming for consuming event-based data from Kafka and joined this data set with existing Hive table data to generate performance indicators for an application.• Developed data warehouse model in Snowflake for over 100 datasets using WhereScape. Created Reports in Looker based on Snowflake Connections. • Created workflows using Airflow to automate the process of extracting weblogs into S3 Data Lake.• Responsible for running the spark jobs along with optimizing, data validation and automation. • Used AWS Lambda, running scripts/code snippets in response to events occurring in CloudWatch. • Integrated applications using Apache tomcat servers on EC2 instances and automated data pipelines into AWS using Jenkins, git, maven.• Developed and run serverless Spark based applications using AWS Lambda service and Pyspark.• Installed, configured & managed RDMS, SQL Server, MYSQL, DB2, PostgreSQL, MongoDB, Cassandra.• Managed and deployed configurations for the entire datacenter infrastructure using Terraform.