Computational Linguist (Text-To-Speech) @ Meta
Current- Developed tailored text normalization (TN) test sets for Spanish locales to optimize TTS voice performance, addressing AI performance challenges in data annotation.- Implemented and refined data-driven TN models using automated techniques, resulting in significant increase in overall accuracy.- Utilized diverse methods, including generative AI prompting and corpora scraping, to optimize training data for ML models.- Maintained rule-based systems for scalable data-driven models, ensuring sustainable maintenance and optimization.- Conducted targeted data quality checks and qualitative analysis to enhance auto-annotation quality.- Authored annotation guidelines for crowdsourced datasets, supporting human judgment pipelines for model refinement.- Drove effective communication of progress and findings to stakeholders, facilitating informed decision-making and collaboration.