Sr. Software Engineer- AI/ML, AWS Neuron Apps
Shape the Future of AI Accelerators at AWS Neuron. Join the elite team behind AWS Neuron—the software stack powering AWS’s next-generation AI accelerators Inferentia and Trainium. As a Senior Software Engineer in our Machine Learning Applications team, you’ll be at the forefront of deploying and optimizing some of the world’s most sophisticated AI models at unprecedented scale.
What You’ll Impact
Pioneer distributed inference solutions for industry‑leading LLMs such as GPT, Llama, Qwen
Optimize breakthrough language and vision generative AI models
Collaborate directly with silicon architects and compiler teams to push the boundaries of AI acceleration
Drive performance benchmarking and tuning that directly impacts millions of inference calls globally
Technical Impact You’ll Drive
Spearhead distributed inference architecture for PyTorch and JAX using XLA
Engineer breakthrough performance optimizations for AWS Trainium and Inferentia
Develop ML tools to enhance LLM accuracy and efficiency
Transform complex tensor operations into highly optimized hardware implementations
Pioneer benchmarking methodologies that shape next‑gen AI accelerator design
What Makes This Role Unique
Direct influence on AWS’s AI infrastructure used by thousands of ML applications
Full‑stack optimization from high‑level frameworks to hardware‑specific primitives
Creation of tools and frameworks that define industry standards for ML deployment
Collaboration with both open‑source ML communities and hardware architecture teams
Your Technical Arsenal Should Include
Deep expertise in Python and ML framework internals
Strong understanding of distributed systems and ML optimization
Passion for performance tuning and system architecture
Basic Qualifications
5+ years of full software development life cycle, including coding standards, code reviews, source control management, build processes, testing, and operations experience
5+ years of programming experience using Python or C++ and PyTorch
Experience with AI acceleration via quantization, parallelism, model compression, batching, KV caching, vllm serving
Experience with accuracy debugging & tooling, performance benchmarking of AI accelerators
Fundamentals of machine learning and deep learning models, their architecture, training and inference lifecycles, and experience on optimizations for improving model execution
Preferred Qualifications
Master’s degree in computer science or equivalent
Master’s degree in machine learning or equivalent
Experience with accuracy debugging & tooling, performance benchmarking of AI accelerators
Experience in developing CUDA kernels, HPC and inference optimization, tensor operations
Benefits
The base salary range for this position is 168,100.00 to 227,400.00 USD annually. Your Amazon package will include sign‑on payments and restricted stock units (RSUs). Final compensation will be determined based on factors including experience, qualifications, and location. Amazon also offers comprehensive benefits including health insurance (medical, dental, vision, prescription, Basic Life & AD&D insurance and optional Supplemental life plans, EAP, Mental Health Support, Medical Advice Line, Flexible Spending Accounts, Adoption and Surrogacy Reimbursement coverage), 401(k) matching, paid time off, and parental leave.
Equal Opportunity Employer
Amazon is an equal opportunity employer and does not discriminate on the basis of protected veteran status, disability, or other legally protected status. Our inclusive culture empowers Amazonians to deliver the best results for our customers. If you have a disability and need a workplace accommodation or adjustment during the application and hiring process, including support for the interview or onboarding process, please visit for more information. If the country/region you’re applying in isn’t listed, please contact your Recruiting Partner.
#J-18808-Ljbffr