Talent Apply
Log in
All jobs
A

Senior Software Engineer - Machine Learning Inference Applications

Amazon
onsiteUSD 168,100 - 227,400 / year
Seattle, WA

About this role

Description

  • AWS Neuron is the complete software stack for the AWS Inferentia and Trainium cloud-scale machinelearning accelerators. This role is for a senior software engineer in the Machine Learning Inference Applications team. This role is responsible for development and performance optimization of core building blocks of LLM Inference - Attention, MLP, Quantization, Speculative Decoding, Mixture of Experts, etc. The team works side by side with chip architects, compiler engineers and runtime engineers to deliver performance and accuracy on Neuron devices across a range of models. Key job responsibilities Responsibilities of this role include adapting latest research in LLM optimization to Neuron chips to extract best performance from both open source as well as internally developed models. Working across teams and organizations is key.

Read the full description on TalentApply

Create a free account to see the complete job description, how well your CV matches this role, and apply in one click.

Clean up your CV

AI rewrites and formats your CV so it reads well and gets past screeners.

See how you score

Get your match percentage for this exact role before you spend time applying.

Apply professionally

Send a polished application in one click — no retyping the same details.

Track it easily

Follow every application in one place instead of digging through your inbox.

Free account · No card required

Your next opportunity starts here

Prepare, apply, track, interview and get hired — all from one platform, with AI in your corner.

Download app