Data Engineer
Current• Work on AWS Data pipeline to configure data loads from S3 to into Redshift.• Create Tables, Stored Procedures, and extracted data using T-SQL for business users whenever required.• Perform data analysis and design, and creates and maintains large, complex logical and physical data models, and metadata repositories using ERWIN and MB MDR• Use AWS Redshift, extracted, transformed and loaded data from various heterogeneous data sources and destinations.• Assist service developers in finding relevant content in the existing reference models.• Use Access, Excel, CSV, Oracle, flat files using connectors, tasks and transformations provided by AWS Data Pipeline.• Design and Develop ETL Processes in AWS Glue to migrate Campaign data from external sources like S3, ORC/Parquet/Text Files into AWS Redshift.• Designed and built Spark/PySpark based ETL pipelines for migration of credit card transactions, account, and customer data into enterprise Hadoop Data Lake. Developed strategies in handling large datasets using partitions, Spark SQL, broadcast joins and performance tuning.• Utilize Spark SQL API in PySpark to extract and load data and perform SQL queries.• Work on developing PySpark script to encrypting the raw data by using hashing algorithms concepts on client specified columns.• Involve in design, development, and testing of the database and Developed Stored Procedures, Views, and Triggers.• Develop Python-based API (RESTful Web Service) to track revenue and perform revenue analysis.• Compile and validate data from all departments and Presenting to Director Operation.• Use KPI calculator Sheet and maintain that sheet within SharePoint.