Senior Data Engineer
Current• Understanding the business functionality & Analysis of business logic for the ETL process• Developed Spark Streaming job to consume the data from the Kafka topic of different source systems and push the data into HDFS locations• Responsible for migrating data from legacy system to AWS Cloud which was running on XMBI Data Lake• Responsible for creating data warehouse and managing entire ETL Data Pipeline from planning to deployment Using Big Data Technologies (Hive, Sqoop, HDFS, Impala, Hue)• Hands on experience of creating tables in Hive with performance optimization using bucketing & partitioning • Creating ETL job’s and extracting data from Data Lake and loaded data into Datawarehouse as per the Business requirement.• Wrote Kafka producers to stream the data from external rest APIs to Kafka topics from legacy system to downstream the data.• Used Kafka and Kinesis for Real time streaming data from Amdocs and ingested into Data Lake.• Built real time data pipelines by developing Kafka producers and Spark streaming applications for consuming the source data as part of Xfinity Mobile• Created lambda function by using python and Downstream the data by streaming from AMDOCS.• Using CloudWatch metrics to monitor the logs for the job’s success and failure.• Each consumer will be consuming the data from Data Lake which was loaded by XMBI.• Involved in extracting the data from different data sources like FLAT files, ORC, PARQUET, CSV files • As part of development has done the cloud POC on VERTICA and ingested files from SFTP server to Data Lake of Hive External Tables.