Ai Engineer
CurrentResearch and develop Talking Head Generation for Virtual Assistant. + Research, develop, and evaluate Generative Model: GAN, Diffusion, Diffusion Transformer. + Research multiple techniques to improve the quality of generated face: PatchGAN Discriminator, VGG Loss, Diffusion. + Research models to control pose of face based on audio condition: Syncnet. + Enhance resolution of generated face: VAE, VQGAN Encoder, Pix2Pix Architecture. + Built streaming microservices for Virtual Assistant: gRPC, Docker. + Built data pipeline to download and preprocess data.Research and develop Automatic Speech Recognition models for Vietnamese Virtual Assistants. + Built, develop, and evaluate End-to-End ASR models such as Wav2vec2.0, HuBERT, and Speech2c using Fairseq, Transformers, HuggingFace, and Pytorch framework. + Built a Data Labeling system using Label Studio (Python, ReactJS). The system has supported labeling for approximately 500 hours of audio. + Research Training Strategies for the pre-training and finetuning phases of Wav2vec2.0 and Speech2c. + Improve ASR model using Masked-CTC, Intermediate-CTC, CNN, and Attention.