Software Development Engineer Iii, Aws Bedrock (Generative Ai)
Current- Bedrock core leads - Partner with Anthropic AI hosting the Claude model families and launching new intelligent models: https://aws.amazon.com/bedrock/claude/ - LLM inference runtime (On-Demand Throughput, Provisioned Throughput). Providing a high-scale, high-throughput, high-performance, multi-tenant distributed model serving infrastructure. Handle substantial input and output token processing rates, ensuring low latency, high resilience and efficiency. Key features include auto-scaling to dynamically adjust resources, robust hardware failure handling, and optimized capacity efficiency for model hosting functionalities. Make it supports complex model routing, accommodates models with long context lengths of up to 200k tokens, and manages concurrent requests effectively. Additionally, the infrastructure supports multimodal models, including both text and image processing, while guaranteeing compliance with trust and safety standards. - Claude on Trainium, Computer Use, Prompt Caching, Benchmarking, and more privileged & confidential..Read more: https://artificialanalysis.ai/providers/amazon_bedrock