Ai Compute Infrastructure Intern
CurrentOptimized the maintenance Cerebras’ AI compute infrastructure by developing Grafana dashboards to monitor cluster health, reducing system downtime.Ensured functionality of 50+ Cerebras appliances through sanity workload testing and performance evaluations. Automated critical tasks such as applying wafer repairs and built email reporting tools, improving overall system reliability and productivity for users / maintainers of the cluster.