Senior Data Engineer
Current• Involved in complete Big Data flow of the application starting from data ingestion upstream to HDFS, processing the data in HDFS and analyzing the data and involved. Developed Json Scripts for deploying the Pipeline in Azure Data Factory (ADF) that process the data using the Cosmos Activity.• Developed Automation Regressing Scripts for validation of ETL process between multiple databases like AWS Redshift, Oracle, Mongo DB, T-SQL, and SQL Server using Python Collaborated with team members and stakeholders in design and development of data environment.• Experienced knowledge over designing Restful services using java-based APIs like JERSEY. Used airflow Operational Services for batch processing and scheduling workflows dynamically.• Experience in developing customized UDF’s in Python to extend Hive and Pig Latin functionality. Created Pipelines in ADF using Linked Services/Datasets/Pipeline/ to Extract, Transform, and load data from different sources like Azure SQL, Blob storage, Azure SQL Data warehouse, write-back tool and backwards.• Worked on performance tuning of data bricks delta tables. Stage the API or Kafka Data (in JSON file format) into Snowflake DB by Flattening the same for different functional services.• Responsible for estimating the cluster size, monitoring, and troubleshooting of the Spark databricks cluster. Used SSIS and T-SQL stored procedures to transfer data from OLTP databases to staging area and finally transfer into data-mart.