Bioinformatics Analyst Ii
CurrentCo-authored ten papers in high impact journals including Nature and JAMA.Presented findings at conferences through poster and oral presentations.Designed a performant database integrating whole exome sequencing of over 230K patients, electronic health records of ~2.3 million patients, and public 'omics and clinical data sources.Developed a natural-language query interface for the complex, real-world SQL database and contributed enhanced text-to-SQL methodology to the LangChain project.Developed advanced NLP methods leveraging large language models for analysis of unstructured clinical text from EHR.Developed and released firthlogist, a Python implementation of Firth logistic regression that reduces analysis runtime by 40% over the R package.Developed and released ASkit, a Python package that achieves a >130x improvement in runtime and 17x reduction in memory usage for PheCode mapping over the R PheWAS package on large datasets.Implemented a deep learning approach to classify severity of nonalcoholic fatty liver disease and nonalcoholic steatohepatitis in liver histopathology.Implemented RNA-seq pipelines to analyze 2,658 human liver samples and discovered noveltranscripts, characterized and visualized differential gene and transcript expression, performedpathway enrichment analyses, and identified critical batch effects.Advised colleagues and collaborators in bioinformatics and genomics methodology and strategy.Mentored undergraduate bioinformatics students and interns.Maintain lab morale with a carefully curated and rapidly updated selection of photos of my cat.