Data Engineer
CurrentAided the migration of business and general-purpose data flows and pipelines from legacy systems to Dataiku, ensuring data integrity and operational efficiency. This involved utilizing Dataiku's plugins and custom Python scripts, enhancing data processing speed and accuracy. In my current role, I develop backend ETL processes to support various dashboards and business intelligence analyses. This work involves extensive use of Dataiku, Python, SQL, and the Google Cloud Platform, where I focus on optimizing data flow and ensuring robust data infrastructure for insightful business decision-making., at Axity.Key role activities: - Understanding business requirements in order to develop ETL processes based on them.- Creating, maintaining and automating backend data flows/processes for business dashboards. - Coordinating business requirements for my teams space in bigquery with the architecture team. - Perform ad-hoc data queries and analyses to fulfill specific business requests, such as extracting user lists or metrics. - Migrate ETL processes from legacy tools, namely Alteryx, to either of the new platforms: Dataiku or Dataflow-In this role, pyspark was used for data transformation steps on massive datasets where SQL was not a viable solution, making use of direct connections to bigquery to load and extract data, these processes lived in dataiku where they were tested for appropriate RAM usage and were then automated once the performance was deemed acceptable -Extensive use of SQL queries was made, it was found many BI processes were wasting resources and could be summarized in a single extensive query which could make use of bigquerie's optimized querying engine -Coordinate new Looker custom explores with BI users and the data architecture team as needed