AI infrastructure roles, filterable by the stack you actually work on.
Distributed training · inference serving · GPU fleets · network fabric — aggregated straight from company boards, never a copy of a copy.
1296 open roles · 85 companies · last verified today
453 roles
Lambdaneocloud
Bellevue Office · San Francisco Office (Fremont St) +1 more · remote · $297K–$440K · manager
eng-managercluster-datacentergpu-genericreliability-srescheduling-orchestration
posted 3w ago · verified today
Anthropicfrontier lab
Remote-Friendly (Travel Required) · Remote-Friendly US (Travel Required) +1 more · remote · $320K–$405K · manager
cluster-datacenterdatacenter-engineerreliability-sregpu-generic
posted 3w ago · verified today
Amazon (AWS)hyperscaler
Seattle, Washington, USA · onsite · unknown
software-engineertrainiumcluster-datacentercpp-langscheduling-orchestration
posted 3w ago · verified today
Lambdaneocloud
Elk Grove Village, IL - Data Center · remote · $137K–$183K · manager
cluster-datacenterdatacenter-engineereng-managerinfiniband-opsnetwork-fabric
posted 3w ago · verified today
Amazon (AWS)hyperscaler
Arlington, Virginia, USA · Denver, Colorado, USA +1 more · onsite · senior
gpu-generickubernetes-opsscheduling-orchestrationsolutions-architectml-platform
posted 4w ago · verified today
Amazon (AWS)hyperscaler
Arlington, Virginia, USA · Denver, Colorado, USA +1 more · onsite · senior
cluster-datacentergpu-generickubernetes-opsnvidiascheduling-orchestration
posted 4w ago · verified today
SambaNovachip vendor
Bengaluru, India · Bengaluru, Karnataka, India · 10300K–12500K INR · staff plus
inference-enginessoftware-engineergo-langinferencekubernetes-ops
posted 4w ago · verified today
SambaNovachip vendor
Bengaluru, India · Bengaluru, Karnataka, India · 10300K–12500K INR · manager
cluster-datacenterobservabilitysoftware-engineerml-platformreliability-sre
posted 3w ago · verified today
SambaNovachip vendor
Bengaluru, India · Bengaluru, Karnataka, India · 8900K–10800K INR · staff plus
cluster-datacentercpp-langgo-langobservabilitypython-lang
posted 3w ago · verified today
SambaNovachip vendor
Bengaluru, India · Bengaluru, Karnataka, India · 6000K–8000K INR · staff plus
software-engineercluster-datacentercpp-langgo-langobservability
posted 3w ago · verified today
Nebiusneocloud
Minnesota · Minnesota, United States · senior
datacenter-engineercluster-datacenterreliability-sregpu-generic
posted 6mo ago · verified today
Together AIneocloud
Amsterdam · London · staff plus
gpu-genericinferencekubernetes-opsscheduling-orchestrationsoftware-engineer
posted 3w ago · verified today
Amazon (AWS)hyperscaler
Austin, Texas, USA · Cupertino, California, USA +1 more · onsite · unknown
cluster-datacentercpp-langgpu-genericlinux-kernelpython-lang
posted 15d ago · verified today
Amazon (AWS)hyperscaler
Tel Aviv-Yafo, Tel Aviv, ISR · onsite · junior
python-langsoftware-engineerml-platformreliability-sreobservability
posted 4w ago · verified today
Together AIneocloud
India · Remote · unknown
go-langgpu-generickubernetes-opspython-langrust-lang
posted 4w ago · verified today
Together AIneocloud
Amsterdam · unknown
cluster-datacentersoftware-engineerscheduling-orchestrationgpu-genericreliability-sre
posted 4w ago · verified today
SambaNovachip vendor
Austin, Texas, United States · Austin, TX +2 more · $210K–$280K · staff plus
inferenceml-platformobservabilityreliability-sresre
posted 9mo ago · verified today
Crusoeneocloud
Dublin - IE · onsite · senior
network-fabricobservabilitynetwork-engineerpython-langcluster-datacenter
posted 4w ago · verified today
Amazon (AWS)hyperscaler
Tel Aviv-Yafo, Tel Aviv, ISR · onsite · senior
cluster-datacenterpython-langsoftware-engineerreliability-srecustom-asic
posted 4mo ago · verified today
Amazon (AWS)hyperscaler
Seattle, Washington, USA · onsite · manager
cudaeng-managerml-platformperformance-engineercluster-datacenter
posted 3mo ago · verified today
Amazon (AWS)hyperscaler
San Francisco, California, USA · onsite · senior
cudagpu-kernelsnvidiapython-langresearch-engineer
posted 3mo ago · verified today
Amazon (AWS)hyperscaler
Austin, Texas, USA · Cupertino, California, USA · onsite · senior
ml-platformpython-langsoftware-engineertrainiumobservability
posted 5mo ago · verified today
Amazon (AWS)hyperscaler
Austin, Texas, USA · onsite · senior
datacenter-engineertrainiumcluster-datacenterreliability-sre
posted 5mo ago · verified today
Amazon (AWS)hyperscaler
Seattle, Washington, USA · onsite · unknown
software-engineertrainiumcpp-langlinux-kernelml-platform
posted 5mo ago · verified today
Amazon (AWS)hyperscaler
Austin, Texas, USA · Cupertino, California, USA +1 more · onsite · manager
eng-managerreliability-sretrainiumcluster-datacenter
posted 11w ago · verified today
Amazon (AWS)hyperscaler
Seattle, Washington, USA · onsite · mid
software-engineerml-platformpost-trainingreinforcement-learningtraining-frameworks
posted 3mo ago · verified today
Amazon (AWS)hyperscaler
Cupertino, California, USA · onsite · mid
gpu-kernelsinferenceinference-enginesdistributed-inferenceml-platform
posted 6w ago · verified today
Amazon (AWS)hyperscaler
Cupertino, California, USA · Denver, Colorado, USA +1 more · onsite · mid
datacenter-engineergpu-genericcluster-datacenterreliability-sre
posted 8w ago · verified today
Amazon (AWS)hyperscaler
Austin, Texas, USA · Cupertino, California, USA +1 more · onsite · senior
reliability-srelinux-kernelsoftware-engineergpu-genericpython-lang
posted 3mo ago · verified today
Amazon (AWS)hyperscaler
Cupertino, California, USA · Seattle, Washington, USA · onsite · senior
gpu-genericnvidiacluster-datacenterreliability-sre
posted 3mo ago · verified today