SDE II, ML Infra Services, Annapurna Labs
Amazon
Seattle, Washington
Posted 1 weeks ago
Responsibilities
Primary Duties
- Lead the design and implementation of ML infrastructure platform, building systems for capacity management, workload scheduling, and fleet orchestration across ML accelerators.
- Work with ML scientists, training infrastructure engineers, hardware teams, and internal customers to ensure the ML Infra service delivers seamless ML Accelerator access with low wait times, high utilization, and zero-config deployment from various environments.
Additional Duties
- Design and code solutions to help our team drive efficiencies in software architecture.
- Create metrics, implement automation and other improvements, and resolve the root cause of software defects.
- Build high-impact solutions to deliver to our large customer base.
- Participate in design discussions, code review, and communicate with internal and external stakeholders.
- Work cross-functionally to help drive business decisions with your technical input.
- Work in a startup-like development environment, where you're always working on the most important stuff.
Experience Requirements
Required
Deep knowledge of profiling and optimization, resource management, scheduling, code generation are needed. The ideal candidate will have worked on new instruction set architectures, which may include CPU, NPU, GPU and other forms of compute.
Full Job Description
Software Engineer
Annapurna Labs was a startup company acquired by AWS in 2015, and is now fully integrated. If AWS is an infrastructure company, then think Annapurna Labs as the infrastructure provider of AWS. Our org covers multiple disciplines including silicon engineering, hardware design and verification, software, and operations. AWS Nitro, ENA, EFA, Graviton and F1 EC2 Instances, AWS Neuron, Inferentia and Trainium ML Accelerators, and in storage with scalable NVMe, are some of the products we have delivered, over the last few years. AWS Neuron is the complete software stack for the AWS Inferentia and Trainium cloud-scale machine learning accelerators and the Trn1 and Inf1 servers that use them.
This position is for a Software Engineer that will lead the development of machine learning tools to run, optimize, and analyze machine learning workloads. This candidate must have had experience leading machine learning tool projects, preferably starting from architecture through several generations of delivery to customers. Deep knowledge of profiling and optimization, resource management, scheduling, code generation are needed. The ideal candidate will have worked on new instruction set architectures, which may include CPU, NPU, GPU and other forms of compute.
Key job responsibilities:
A day in the life:
About the team:
Diverse Experiences: We value diverse experiences and non-traditional career paths. If your career is just starting or includes alternative experiences, we encourage you to apply.
Inclusive Team Culture: Our employee-led affinity groups foster inclusion. Events like CORE and AmazeCon inspire us to embrace our uniqueness.
Work/Life Balance: We strive for flexibility as part of our working culture, supporting you both at work and at home.
Mentorship & Career Growth: We offer knowledge-sharing, mentorship, and one-on-one code reviews to help you grow as a professional.
Annapurna Labs was a startup company acquired by AWS in 2015, and is now fully integrated. If AWS is an infrastructure company, then think Annapurna Labs as the infrastructure provider of AWS. Our org covers multiple disciplines including silicon engineering, hardware design and verification, software, and operations. AWS Nitro, ENA, EFA, Graviton and F1 EC2 Instances, AWS Neuron, Inferentia and Trainium ML Accelerators, and in storage with scalable NVMe, are some of the products we have delivered, over the last few years. AWS Neuron is the complete software stack for the AWS Inferentia and Trainium cloud-scale machine learning accelerators and the Trn1 and Inf1 servers that use them.
This position is for a Software Engineer that will lead the development of machine learning tools to run, optimize, and analyze machine learning workloads. This candidate must have had experience leading machine learning tool projects, preferably starting from architecture through several generations of delivery to customers. Deep knowledge of profiling and optimization, resource management, scheduling, code generation are needed. The ideal candidate will have worked on new instruction set architectures, which may include CPU, NPU, GPU and other forms of compute.
Key job responsibilities:
- Lead the design and implementation of ML infrastructure platform, building systems for capacity management, workload scheduling, and fleet orchestration across ML accelerators.
- Work with ML scientists, training infrastructure engineers, hardware teams, and internal customers to ensure the ML Infra service delivers seamless ML Accelerator access with low wait times, high utilization, and zero-config deployment from various environments.
A day in the life:
- Design and code solutions to help our team drive efficiencies in software architecture.
- Create metrics, implement automation and other improvements, and resolve the root cause of software defects.
- Build high-impact solutions to deliver to our large customer base.
- Participate in design discussions, code review, and communicate with internal and external stakeholders.
- Work cross-functionally to help drive business decisions with your technical input.
- Work in a startup-like development environment, where you're always working on the most important stuff.
About the team:
- High-impact, high-visibility: You'll directly accelerate every Neuron team's ability to ship — your work multiplies the output of 100+ engineers.
- Greenfield opportunities: We're actively building new capabilities with significant design ownership for SDEs.
- Small, senior team: where every person owns major components and drives architectural decisions.
- AI infrastructure: Work at the intersection of Kubernetes, custom silicon, and large-scale ML workloads.
Diverse Experiences: We value diverse experiences and non-traditional career paths. If your career is just starting or includes alternative experiences, we encourage you to apply.
Inclusive Team Culture: Our employee-led affinity groups foster inclusion. Events like CORE and AmazeCon inspire us to embrace our uniqueness.
Work/Life Balance: We strive for flexibility as part of our working culture, supporting you both at work and at home.
Mentorship & Career Growth: We offer knowledge-sharing, mentorship, and one-on-one code reviews to help you grow as a professional.
Company Culture
Core Values
Diverse Experiences: We value diverse experiences and non-traditional career paths.Inclusive Team Culture: Our employee-led affinity groups foster inclusion.Work/Life Balance: We strive for flexibility as part of our working culture.Mentorship & Career Growth: We offer knowledge-sharing, mentorship, and one-on-one code reviews to help you grow as a professional.
How to Apply
$90
/ hour
Amazon pays $90 for Software Engineer in Seattle, Washington, with most salaries ranging from $70 to $121. Pay can vary based on role, experience, and local cost of living.
Median
$90
Low
$70
High
$121
Companies Similar to Amazon for Jobs
Share This Job
Figures represent approximate ranges and may vary based on experience, location, and other factors. For the most accurate information, please consult the employer directly. Contact us to suggest updates to this information.





