Data Engineer
Current1. Orchestrate and automate azure data factory pipelines to extract documents from Veeva vault. 2. Writing notebooks using pyspark to make API call for document extraction, apply file naming convention to files in staging server and copy them into data lake storage gen2. 3. Follow medallion architecture to load data from API source into end tables and create data products for downstream consumers.4. Validate document metadata entries and create XML files to get it processed by proprietary system for archive. Finally copy the files from vault system to azure fileshare.