Jump to content

Software Engineer- AI/ML, Amazon Neuron Training - Cupertino | Amazon Job


 Share

Job Opportunity Details

Type

Full Time

Salary

Not Telling

Work from home

No

Weekly Working Hours

Not Telling

Positions

Not Telling

Working Location

Cupertino, Cupertino, CA, United States   [ View map ]

Job Description

The Annapurna Labs team at Amazon Web Services (AWS) builds AWS Neuron, the software development kit used to accelerate deep learning and GenAI workloads on AWS Trainium, Amazon's custom machine learning accelerator. Neuron includes an ML compiler, runtime, collectives library, and application framework that integrate with PyTorch and JAX, so customers can train frontier-scale models on Trainium without rewriting their stack.

The Distributed Training team enables the training of a wide range of models, from large-scale pretraining through post-training and reinforcement learning, on AWS's custom ML accelerators. As more customer workloads shift toward RLHF, PPO/GRPO, and other fine-tuning methods, we are building the distributed training infrastructure, parallelism techniques, numerics, and high-performance kernels that these methods depend on. As part of the broader Neuron organization, we work across frameworks, kernels, compiler, runtime, and collectives — a true hardware and software co-design in practice. We not only optimize current performance but also contribute to future architecture designs, since the gaps we characterize today become requirements for the next generation of Trainium.

We are looking for software engineers to help build and fine tune these distributed training solutions. This role offers a rare opportunity to work at the intersection of machine learning, high-performance computing, and distributed systems, where you will help shape the direction of AI acceleration technology.


Key job responsibilities
You’ll implement and tune components of our distributed training stack for large-scale training, post-training, and reinforcement learning workloads on the latest Trainium instances, working across PyTorch and the Neuron software stack. You'll contribute to parallelism strategies such as data, tensor, and pipeline parallelism and apply reduced-precision formats under the guidance of senior team members. You'll profile workloads to help determine whether a bottleneck sits in compute, memory, collectives, or host overhead, and work with compiler, runtime, and collectives engineers to help land the fix. You'll build and maintain internal tooling, benchmarks, and tests that keep the team's performance work reproducible, and take on increasing ownership as you grow in the role.

About the team
Inclusive Team Culture
Here at Amazon, we embrace our differences. We are committed to furthering our culture of inclusion. We have ten employee-led affinity groups, reaching 40,000 employees in over 190 chapters globally. We have innovative benefit offerings, and host annual and ongoing learning experiences, including our Conversations on Race and Ethnicity (CORE) and AmazeCon (gender diversity) conferences. Amazon’s culture of inclusion is reinforced within our 16 Leadership Principles, which remind team members to seek diverse perspectives, learn and be curious, and earn trust.

Work/Life Balance
Our team puts a high value on work-life balance. It isn’t about how many hours you spend at home or at work; it’s about the flow you establish that brings energy to both parts of your life. We believe striking the right balance between your personal and professional life is critical to life-long happiness and fulfillment. We offer flexibility in working hours and encourage you to find your own balance between your work and personal lives.

Basic Qualifications:

- Bachelor's degree or above in computer science or equivalent
- 3+ years of experience with full software development life cycle in production
- 3+ years of experience with at least one programming language such as Python, C/C++, or a similar language
- 2+ years of experience with system design (design patterns, reliability and scaling) of new or existing systems
- Familiarity with LLM/transformer fundamentals such as attention mechanisms, autoregressive decoding, KV-cache behavior, and various forms of parallelism

Preferred Qualifications:

- Master's degree or above in computer science or equivalent
- Experience in developing and deploying LLMs in production on GPUs, Neuron, TPU or other AI acceleration hardware, or experience in computer architecture
- Experience with ML frameworks such as Pytorch/Jax, Distributed libraries and Frameworks, RL frameworks or End-to-end Model Training
- Experience with performance engineering: workload profiling, characterization (compute bound, memory bound, network bound), and optimization
- Experience with open source projects

Amazon is an equal opportunity employer and does not discriminate on the basis of protected veteran status, disability, or other legally protected status.

Los Angeles County applicants: Job duties for this position include: work safely and cooperatively with other employees, supervisors, and staff; adhere to standards of excellence despite stressful conditions; communicate effectively and respectfully with employees, supervisors, and staff to ensure exceptional customer service; and follow all federal, state, and local laws and Company policies. Criminal history may have a direct, adverse, and negative relationship with some of the material job duties of this position. These include the duties and responsibilities listed above, as well as the abilities to adhere to company policies, exercise sound judgment, effectively manage stress and work safely and respectfully with others, exhibit trustworthiness and professionalism, and safeguard business operations and the Company’s reputation. Pursuant to the Los Angeles County Fair Chance Ordinance, we will consider for employment qualified applicants with arrest and conviction records.

Our inclusive culture empowers Amazonians to deliver the best results for our customers. If you have a disability and need a workplace accommodation or adjustment during the application and hiring process, including support for the interview or onboarding process, please visit https://amazon.jobs/content/en/how-we-hire/accommodations for more information. If the country/region you’re applying in isn’t listed, please contact your Recruiting Partner.

The base salary range for this position is listed below. Your Amazon package will include sign-on payments and restricted stock units (RSUs). Final compensation will be determined based on factors including experience, qualifications, and location. Amazon also offers comprehensive benefits including health insurance (medical, dental, vision, prescription, Basic Life & AD&D insurance and option for Supplemental life plans, EAP, Mental Health Support, Medical Advice Line, Flexible Spending Accounts, Adoption and Surrogacy Reimbursement coverage), 401(k) matching, paid time off, and parental leave. Learn more about our benefits at https://amazon.jobs/en/benefits.



USA, CA, Cupertino - 165,200.00 - 223,600.00 USD annually


More Information

Application Details

  • Organization Details
    Annapurna Labs (U.S.) Inc.
 Share


User Feedback

Recommended Comments

There are no comments to display.

Join the conversation

You are posting as a guest. If you have an account, sign in now to post with your account.
Note: Your post will require moderator approval before it will be visible.

Guest
Add a comment...

×  Pasted as rich text.   Paste as plain text instead

  Only 75 emoji are allowed.

×  Your link has been automatically embedded.   Display as a link instead

×  Your previous content has been restored.   Clear editor

×  You cannot paste images directly. Upload or insert images from URL.

Loading...
×
×
  • Create New...