Lead Data Scientist
San Francisco, Ca, Us
Support the organization's B2B online marketplace by implementing a language agnostic data framework to rapidly deploy data applications using Docker, CRON, and AWS ECR.Maintained our Snowflake data warehouse and all warehouse functionality: data lake support, storage integration configuration, compute resource allocation, automated view building, and role permissioning.Maintained all ETL pipelines into the data warehouse using different ETL Tooling: Stitch data for the various tools we used(Salesforce, Sendgrid, etc) and AWS DMS for change data capture replication of production Postgres services.Created a recommendation engine using alternating least squares to recommend products to users based on their transfer history within a state’s seed to sale tracking system (Metrc).Created a supply chain network data set based on a state’s seed to sale tracking system (Metrc) that could allow a supplier to see current stock levels downstream of products they have sold to their customers. This data set is now a core subscription product offering called Relationships at Confident Cannabis.Created an abstract data pipeline that allowed the data team to deliver complex data objects to our production Postgres environments. This in turn allowed the backend team to build read-only Django models on top of the data objects we deployed for rapid product development.Supported the sales team by creating intelligent ETLs into Salesforce to attribute internal data insights to Salesforce related objects (Opportunities, Accounts, Events, Custom Objects, etc).Built out custom dashboard reporting within our app using Django to build Plotly JSON chart objects that could be rendered on the front end.