AI infrastructure roles, filterable by the stack you actually work on.
Distributed training · inference serving · GPU fleets · network fabric — aggregated straight from company boards, never a copy of a copy.
1296 open roles · 85 companies · last verified today
116 roles
Anthropicfrontier lab
New York City, NY · San Francisco, CA +1 more · $320K–$485K · staff plus
cpp-langebpfgo-langgpu-generickubernetes-ops
posted 4mo ago · verified today
Anthropicfrontier lab
London, UK · £325K–£485K · staff plus
kubernetes-opsscheduling-orchestrationsoftware-engineerml-platformebpf
posted 4mo ago · verified today
Anthropicfrontier lab
New York City, NY · San Francisco, CA +1 more · $280K–$850K · unknown
cudagpu-genericgpu-kernelsperformance-engineercollectives
posted 11mo ago · verified today
Nebiusneocloud
Remote - Europe · Remote - United States · remote · $170K–$300K/year est. · senior
cluster-datacentercpp-langgpu-genericinfiniband-opsnetwork-fabric
posted 5mo ago · verified today
Nebiusneocloud
Amsterdam, Netherlands · Remote - Europe +1 more · remote · senior
performance-engineercollectivescpp-langgo-langgpu-generic
posted 4mo ago · verified today
Nebiusneocloud
Palo Alto · Palo Alto, California, United States · $195K–$262K · senior
model-parallelismpost-trainingpytorch-distreinforcement-learningtraining-frameworks
posted 8w ago · verified today
Nebiusneocloud
Amsterdam, Netherlands · Berlin, Germany +2 more · remote · senior
cluster-datacentercpp-langgpu-genericinfiniband-opsnetwork-fabric
posted 25mo ago · verified today
Nebiusneocloud
Remote - United States · United States · $180K–$224K/year est. · senior
gpu-genericnvidiacluster-datacenterlinux-kerneldatacenter-engineer
posted 7mo ago · verified today
Nebiusneocloud
Amsterdam, Netherlands · Czech Republic +5 more · remote · unknown
cudagpu-genericgpu-kernelsnccl-libperformance-engineer
posted 4mo ago · verified today
Nebiusneocloud
Remote - United States · United States · remote · $170K–$300K · manager
gpu-genericperformance-engineercluster-datacentercpp-langinfiniband-ops
posted 4mo ago · verified today
Crusoeneocloud
San Francisco, CA - US · onsite · mid
gpu-genericcluster-datacenternvidiareliability-srenccl-lib
posted 6w ago · verified today
Crusoeneocloud
San Francisco, CA - US · Sunnyvale, CA - US · onsite · staff plus
gpu-kernelssoftware-engineercluster-datacenternccl-libobservability
posted 5w ago · verified today
Crusoeneocloud
San Francisco, CA - US · Sunnyvale, CA - US · onsite · staff plus
collectivesgpu-genericinfiniband-opskubernetes-opsnccl-lib
posted 7mo ago · verified today
Lambdaneocloud
Bellevue Office · San Francisco Office (Fremont St) · remote · $240K–$356K · senior
reliability-sresreansiblecluster-datacenterinfiniband-ops
posted 15d ago · verified today
Lambdaneocloud
Remote, USA · remote · $122K–$162K · senior
collectivescudagpu-genericinfiniband-opskubernetes-ops
posted 3w ago · verified today
Lambdaneocloud
Bellevue Office · San Francisco Office (Fremont St) +1 more · remote · $314K–$465K · staff plus
go-langgpu-generickubernetes-opsml-platformnvidia
posted 5w ago · verified today
Lambdaneocloud
Bellevue Office · San Francisco Office (Fremont St) +1 more · remote · $266K–$395K · senior
go-langgpu-generickubernetes-opsml-platformnvidia
posted 4w ago · verified today
Together AIneocloud
San Francisco · $240K–$280K/year est. · staff plus
kubernetes-opssoftware-engineercluster-datacentergo-langgpu-generic
posted 8w ago · verified today
Together AIneocloud
Amsterdam · senior
nvidiasoftware-engineercluster-datacenterinfiniband-opskubernetes-ops
posted 7mo ago · verified today
Together AIneocloud
San Francisco · $220K–$290K/year est. · senior
cluster-datacentergpu-genericnvidiasoftware-engineeransible
posted 15mo ago · verified today
Together AIneocloud
San Francisco · $200K–$290K/year est. · senior
inferenceinference-enginesnvidiasoftware-engineerdistributed-inference
posted 13mo ago · verified today
Together AIneocloud
San Francisco · $200K–$290K/year est. · unknown
cudafine-tuningml-platformnccl-libnvidia
posted 6w ago · verified today
Together AIneocloud
San Francisco · $200K–$290K/year est. · unknown
kubernetes-opsml-platformsoftware-engineeransiblego-lang
posted 14mo ago · verified today
Together AIneocloud
Amsterdam · unknown
cluster-datacentergpu-genericsoftware-engineercudago-lang
posted 16mo ago · verified today
Together AIneocloud
San Francisco · $190K–$270K/year est. · mid
cluster-datacentergpu-genericpython-langreliability-sresoftware-engineer
posted 4mo ago · verified today
Together AIneocloud
Bangalore, India · Remote · mid
gpu-generickubernetes-opspython-langsoftware-engineeransible
posted 9w ago · verified today