Data Scientist Ii, Alexa Speech
Current- Developed a probabilistic model to estimate the impact of errors made by Alexa's automatic speech recognition (ASR) technology on Alexa’s end-to-end customer experience metrics; results were used by ASR leadership to calculate ASR’s contribution to Alexa-level goals and for future ASR-level goal setting.- Developed causal models using methods in observational causal inference to estimate causal effects of different factors on ASR accuracy, which increased confidence that the ASR org was correctly prioritizing programs to improve performance.- Developed a statistical model (unsupervised mixture model) to estimate error rates of Alexa's speaker identification models, which was used to help gate the release of Alexa speaker verification models for mobile devices that lacked human-annotated labels.- Developed a big data processing framework using PySpark to identify 300k+ entities that caused the highest rates of customer friction across all Alexa traffic in Japan, which helped augment training sets to improve Alexa ASR models, and also added a new capability to our data science team to process big data.- Conducted a data analysis using hypothesis testing, regression analysis, and hierarchical clustering to identify customer cohorts with low ASR accuracy, which led to program initiatives to improve accuracy.- Devised a new method to bridge ASR errors with monetary impact, and estimated the 365-day downstream profit loss caused by ASR errors in each key segment of Alexa traffic, which enabled the ASR org to identify the low-performing segments associated with highest profit loss.