Senior System Engineer
Hyderabad, Telangana, India
• Seamlessly executed data ingestion and processing operations using Apache Spark, interfacing with Snowflake and S3, thereby streamlining data workflows and accelerating processing time for extensive datasets.• Automated large data modeling tasks using EMR Cluster with Spark, resulting in a 50% reduction in data processing time, enhancing data quality.• Used AWS Lambda functions for automating data processing, leading to a 34% improvement in efficiency.• Developed and maintained complex ETL workflows in Airflow, scheduling and monitoring data processing tasks to ensure timely and accurate data delivery across various systems.• Achieved almost 99% uptime for the database system by implementing database replication for high availability.• Skilled in developing comprehensive data models to optimize data storage and retrieval processes.• Skilled in using Spark Data Frame persistency and caching mechanisms to reduce data processing overhead and improve query performance.• Skilled in data profiling and management using SQL queries, with a focus on creating, testing, and refining queries to support data profiling and optimizing data quality and integrity