Ai Research Intern
During my internship from March to June 2024, I engaged in two distinct projects, each focusing on advanced applications of Large Language Models.Project 1: The Grading Capabilities of Large Language ModelsThis project involved a comparative study of OpenAI and open-source LLMs across Python and short answer assessments using rubrics. Key highlights and contributions include:-Developed an automated scoring pipeline that seamlessly integrated multiple Large Language Model platforms through advanced prompt engineering.-Authored a comprehensive framework to evaluate LLM performance, utilising a variety of rubrics and models.-Executed detailed evaluations of various combinations of LLMs and rubrics, diligently reporting the findings to assess the efficacy of automated grading systems.Project 2: Retrieval-Augmented Generation (RAG) EvaluationIn this project, I carried out a comparative analysis of different locally implemented RAG workflows. My work focused on assessing the efficacy of both open and closed-source RAG implementations, with significant outcomes including:-Conducted thorough evaluations of five distinct Retrieval-Augmented Generation (RAG) implementations, assessing their ability to retrieve and generate accurate information effectively.-Compiled and presented findings that provided essential insights into the operational performance of RAG systems.-Investigated and identified potential improvements to enhance the performance and efficiency of locally implemented RAG applications.