AI infrastructure roles, filterable by the stack you actually work on.

Distributed training · inference serving · GPU fleets · network fabric — aggregated straight from company boards, never a copy of a copy.

1297 open roles · 85 companies · last verified today

410 roles

ml-platform

Amazon (AWS)hyperscaler

Austin, Texas, USA · onsite · unknown

linux-kernelsoftware-engineertrainiumml-platform

posted 3mo ago · verified today

Amazon (AWS)hyperscaler

Seattle, Washington, USA · onsite · unknown

software-engineertrainiumcpp-langlinux-kernelml-platform

posted 5mo ago · verified today

Amazon (AWS)hyperscaler

Cupertino, California, USA · Seattle, Washington, USA · onsite · senior

performance-engineersoftware-engineertrainiumgpu-kernelsml-platform

posted 11mo ago · verified today

Amazon (AWS)hyperscaler

Santa Clara, California, USA · onsite · unknown

gpu-kernelstraining-frameworksml-platformresearch-engineersoftware-engineer

posted 9mo ago · verified today

Amazon (AWS)hyperscaler

Seattle, Washington, USA · onsite · unknown

cpp-langperformance-engineersoftware-engineertrainiumml-platform

posted 8w ago · verified today

Amazon (AWS)hyperscaler

New York, New York, USA · onsite · mid

software-engineerml-platform

posted 13mo ago · verified today

Amazon (AWS)hyperscaler

New York, New York, USA · onsite · senior

software-engineertrainiumml-platform

posted 7mo ago · verified today

Amazon (AWS)hyperscaler

Cupertino, California, USA · Seattle, Washington, USA · onsite · mid

trainiumperformance-engineersoftware-engineercollectivesgpu-generic

posted 7mo ago · verified today

Amazon (AWS)hyperscaler

New York, New York, USA · onsite · senior

software-engineertrainiumml-platform

posted 6w ago · verified today

Amazon (AWS)hyperscaler

Cupertino, California, USA · Seattle, Washington, USA · onsite · senior

software-engineertrainiummodel-parallelismpre-trainingtraining-frameworks

posted 5w ago · verified today

Amazon (AWS)hyperscaler

Cupertino, California, USA · onsite · manager

eng-managertrainiumdeepspeed-libfsdpmodel-parallelism

posted 3mo ago · verified today

Amazon (AWS)hyperscaler

Toronto, Ontario, CAN · onsite · manager

eng-managertrainiummlir-llvmml-platform

posted 7mo ago · verified today

Amazon (AWS)hyperscaler

Seattle, Washington, USA · onsite · mid

software-engineertrainiumml-platformscheduling-orchestrationgo-lang

posted 11w ago · verified today

Amazon (AWS)hyperscaler

Cupertino, California, USA · onsite · senior

gpu-kernelsvllm-enginemodel-parallelismdistributed-inferenceml-platform

posted 6w ago · verified today

Amazon (AWS)hyperscaler

Cupertino, California, USA · onsite · mid

gpu-kernelsinferenceinference-enginesdistributed-inferenceml-platform

posted 6w ago · verified today

Amazon (AWS)hyperscaler

Tel Aviv-Yafo, Tel Aviv, ISR · onsite · mid

cpp-langsoftware-engineercustom-asicdistributed-inferencegpu-kernels

posted 5w ago · verified today

Amazon (AWS)hyperscaler

Bellevue, Washington, USA · onsite · mid

software-engineerml-platformtrainiumscheduling-orchestrationfsdp

posted 6w ago · verified today

Amazon (AWS)hyperscaler

Boston, Massachusetts, USA · Seattle, Washington, USA +1 more · onsite · senior

cudacutlass-cutedistributed-inferenceevaluationflash-attention

posted 4w ago · verified today

FluidStackneocloud

Austin, TX · onsite · $202K–$241K/year · unknown

cluster-datacenterdatacenter-engineergpu-genericnetwork-fabricroce-net

posted 2d ago · verified today

Fireworks AIinference provider

New York · San Mateo · hybrid · $200K–$290K/year · senior

reliability-sresreobservabilitycpp-langgo-lang

posted 4w ago · verified today