Data Scientist
CurrentData science and analytics developer for NeuroBlu - Holmusk's flagship product.Tech stack: Python, R, DuckDB, SQLite, Postgres, AWS, DatabricksNeuroBlu Research software developer- Develop Python/R libraries to perform data processing and statistical analysis. - Helps pharma clients and internal scientists execute high-value research projects e.g. External Comparator Arms, patient treatment journeys etc.- Improve query performance by up to 30% with PySpark, DuckDB. - Rewrote significant portion of legacy codebase using more modern relation-based APIs. - Implement robust unit-testing as part of CI/CD.QC/QA development- Implemented OHDSI (Observational Health Data Sciences and Informatics initiative) industry-standard tooling and framework for QC validation on Databricks. - Achieved 97% data quality score based on OHDSI standards, demonstrating value of Holmusk’s feature-rich data to clients and investors. - Responsible for end to end implementation, from configuring compute clusters to orchestrating workflows, all on Databricks.Data infrastructure migration- Prototyped POCs for Redshift/Snowflake/Databricks: - Benchmarked database performance based on - Query performance - Ease of implementing ETL - Scalability of compute resources - Load testing with JMeter----Holmusk, a 2019 World Economic Forum Technology Pioneer, is a data science and digital therapeutics company dedicated to establishing objective evidence as a core utility to the treatment of mental health and chronic diseases. Our proprietary technology leverages analytic and digital tools to prevent, manage and reverse treatable diseases.