Data Engineer
CurrentAs a Data Engineer, I bring extensive expertise in the full Software Development Life Cycle (SDLC), encompassing requirement gathering, design, development, testing, deployment, and user support. My professional journey involves optimizing data transformation processes across diverse platforms, including Oracle and Hadoop, with a primary focus on Python and Hive. I collaborate closely with Solution Architects, Principal Engineers, DevOps teams, and Business Analysts to translate business needs into actionable data products, ensuring high-quality deliverables.I have hands-on experience designing and developing ETL jobs to integrate data from Salesforce replicas into Redshift, managing data quality and integrity through advanced Big Data technologies such as Hadoop, MapReduce, Pig, Hive, Flume, Sqoop, Spark, and AWS services including S3, Lambda, EC2, EMR, RedShift, and DynamoDB. My role involves crafting and optimizing Spark SQL queries, developing custom RDDs in PySpark, and leveraging Kafka REST API for data collection and integration.I am adept at real-time data processing using Kafka and Spark Streaming, efficiently converting and storing data in Parquet format on HDFS. My experience also includes scheduling and maintaining database jobs, creating innovative solutions for data movement and transformation, and prototyping new data integration tools and techniques. Active in Agile environments, I participate in Scrum meetings, Sprint planning, and retrospectives to drive project success.