Jr. Data Engineer
CurrentI work creating, managing and optimizing ETL and ELT pipelines using Airflow, PySpark, Spark, SQL, Python, Bash, Docker, Polars, Pandas and Databricks. I use the Bash terminal via WSL.I build automation models using Selenium and perform web scraping using Scrapy, Selenium, Crawlee and Beautiful Soup.I also design APIs using FastAPI. My codes are versioned in Git, remotely in GitLab and organized into packages and virtual environments using Poetry and uv.I use Streamlit… Show more I work creating, managing and optimizing ETL and ELT pipelines using Airflow, PySpark, Spark, SQL, Python, Bash, Docker, Polars, Pandas and Databricks. I use the Bash terminal via WSL.I build automation models using Selenium and perform web scraping using Scrapy, Selenium, Crawlee and Beautiful Soup.I also design APIs using FastAPI. My codes are versioned in Git, remotely in GitLab and organized into packages and virtual environments using Poetry and uv.I use Streamlit, Plotly, Seaborn and Matplotlib to build dashboard and report visualizations that generate insights for decision making, fed by the pipelines, web scraping and automations mentioned above.One of my success stories was the creation of a platform that automatically generates reports from Yahoo Search using tools such as OpenAI API, Scrapy, GitHub, Apache Airflow and Streamlit. The development took 4 days and increased the agility of producing documents with reliable sources by 60%, in a clear and concise manner.I am currently implementing Azure in the company, for which I wrote the documentation and architecture, in order to ensure that the data-driven culture spreads throughout the production chains.• Table formats and file formats and extensions that I am familiar with: csv, xlsx, xls, Parquet, JSON, txt, AVRO, Delta, ORC, YAML and Iceberg. Show less