Senior Data Engineer
Current- Creating new and maintaining existing ETL processes- Designing infrastructure for machine learning processes- Developing libraries and tools for data analysts- Performing upgrades, migrations, and optimizations of softwareand infrastructure components- Establishing data integrity and quality control systems- Configuring and integrating with external services- Improving the work environment and internal team processes- Providing technical support for team membersExamples of projects:Developer Setup Improvements: Implemented automatic project setup on Airflowlocally using Docker Compose.1. Developed an easily extensible callback processing system based on Amazon SQSqueues. Integrated it with Twilio and Sendgrid services.2. On my own initiative, optimized a process involving multiple EMR clusters usingmultithreading. Reduced the process execution time from 7 hours to 2.5 hours, andconsequently reduced AWS resource usage costs.3. Optimized EMR cluster setup by updating the EMR version, switching to newer EC2instance types, and archiving all Python dependencies. Reduced cluster startuptime from 40 minutes to 10 minutes.4. Migrated a machine learning-based prediction process from a long-runningcluster to an Airflow project. Refactored the code, added integration tests for theinput data processing stage, additional logging and notifications, andimplemented the ability to run test models to compare results with the main model.