Data Engineer
Current•Built and managed AWS Data Pipelines using Lambda to facilitate data transfers between S3, DynamoDB, and Redshift for optimized ETL processes.•Developed complex T-SQL queries and stored procedures to fulfill business user requirements, integrating with Azure Synapse Analytics for real-time processing.•Integrated ERWIN and MB MDR for logical and physical data modeling, improving metadata management across projects.•Optimized data encryption processes in PySpark using hashing algorithms, ensuring client data security.•Designed and optimized Azure Data Factory pipelines for seamless data migration and ETL automation.•Implemented AWS EMR clusters for distributed processing of large datasets using Spark, Hive, and Hadoop.•Developed Python-based REST APIs for revenue tracking and analysis, automating report generation processes.•Used C++ for performance improvements in the data ingestion pipeline, optimizing low-level system operations.•Managed HBase clusters for high-performance storage and retrieval of sparse data.•Built comprehensive Tableau and Power BI dashboards for real-time KPI tracking and performance analysis.•Used Terraform to automate infrastructure provisioning for Kubernetes and Docker containers, enabling scalability across cloud platforms.