AI infrastructure roles, filterable by the stack you actually work on.
Distributed training · inference serving · GPU fleets · network fabric — aggregated straight from company boards, never a copy of a copy.
1297 open roles · 85 companies · last verified today
196 roles
FluidStackneocloud
Austin, TX · CA +4 more · onsite · $208K–$263K/year · manager
datacenter-engineernvidiacluster-datacentergpu-genericnetwork-fabric
posted 3w ago · verified today
Amazon (AWS)hyperscaler
Arlington, Virginia, USA · Denver, Colorado, USA +1 more · onsite · unknown
kubernetes-opsnvidiascheduling-orchestrationgpu-genericsolutions-architect
posted 3w ago · verified today
Nebiusneocloud
South Korea · staff plus
solutions-architectcudagpu-genericansiblekubernetes-ops
posted 3w ago · verified today
Nebiusneocloud
New York City, New York, United States · Remote - United States +1 more · remote · $200K–$245K · senior
gpu-genericnvidiasolutions-architectinferenceinference-engines
posted 7mo ago · verified today
Lambdaneocloud
Elk Grove Village, IL - Data Center · remote · $137K–$183K · manager
cluster-datacenterdatacenter-engineereng-managerinfiniband-opsnetwork-fabric
posted 3w ago · verified today
Amazon (AWS)hyperscaler
Arlington, Virginia, USA · Denver, Colorado, USA +1 more · onsite · senior
gpu-generickubernetes-opsscheduling-orchestrationsolutions-architectml-platform
posted 4w ago · verified today
Amazon (AWS)hyperscaler
Arlington, Virginia, USA · Denver, Colorado, USA +1 more · onsite · senior
cluster-datacentergpu-generickubernetes-opsnvidiascheduling-orchestration
posted 4w ago · verified today
Nebiusneocloud
Remote - United States · United States · remote · $180K–$220K/year est. · senior
gpu-genericnvidiasolutions-architectcluster-datacenterperformance-engineer
posted 5mo ago · verified today
Nebiusneocloud
Remote - Europe · remote · mid
solutions-architectgpu-generickubernetes-opspython-langterraform-iac
posted 3w ago · verified today
SambaNovachip vendor
Austin, Texas, United States · Austin, TX +2 more · $210K–$280K · staff plus
inferenceml-platformobservabilityreliability-sresre
posted 9mo ago · verified today
CoreWeaveneocloud
Bellevue, WA · San Francisco, CA +1 more · $198K–$264K/year est. · staff plus
amdcluster-datacenterdeepspeed-libdistributed-inferencefsdp
posted 4w ago · verified today
Amazon (AWS)hyperscaler
Seattle, Washington, USA · onsite · manager
eng-managernvidiacluster-datacentercpp-langml-platform
posted 5w ago · verified today
Amazon (AWS)hyperscaler
Seattle, Washington, USA · onsite · manager
eng-managernetwork-fabricrdma-verbsnvidia
posted 5w ago · verified today
Amazon (AWS)hyperscaler
Seattle, Washington, USA · onsite · manager
eng-managernetwork-fabricrdma-verbsnvidia
posted 6w ago · verified today
Amazon (AWS)hyperscaler
San Francisco, California, USA · onsite · senior
cudagpu-kernelsnvidiapython-langresearch-engineer
posted 3mo ago · verified today
Amazon (AWS)hyperscaler
Cupertino, California, USA · Seattle, Washington, USA · onsite · manager
collectivescpp-langcudagpu-kernelsnccl-lib
posted 6w ago · verified today
Amazon (AWS)hyperscaler
Seattle, Washington, USA · onsite · mid
nvidiasoftware-engineercluster-datacenternetwork-fabric
posted 5mo ago · verified today
Amazon (AWS)hyperscaler
Seattle, Washington, USA · onsite · mid
software-engineernvidiacluster-datacentergpu-genericnetwork-fabric
posted 3mo ago · verified today
Amazon (AWS)hyperscaler
Cupertino, California, USA · Seattle, Washington, USA · onsite · senior
gpu-genericnvidiacluster-datacenterreliability-sre
posted 3mo ago · verified today
Amazon (AWS)hyperscaler
Seattle, Washington, USA · onsite · mid
gpu-generickubernetes-opssoftware-engineernvidiaobservability
posted 9w ago · verified today
Amazon (AWS)hyperscaler
Boston, Massachusetts, USA · Seattle, Washington, USA +1 more · onsite · senior
cudacutlass-cutedistributed-inferenceevaluationflash-attention
posted 4w ago · verified today
Amazon (AWS)hyperscaler
Herndon, Virginia, USA · New York, New York, USA +1 more · onsite · senior
solutions-architectgpu-genericinference-enginesmodel-parallelismtraining-frameworks
posted 3mo ago · verified today
Amazon (AWS)hyperscaler
Toronto, Ontario, CAN · onsite · senior
gpu-kernelsperformance-engineertrainiumcudatriton-lang
posted 15d ago · verified today
Nebiusneocloud
Reykjanesbær · Reykjanesbær, Iceland · junior
cluster-datacenterdatacenter-engineernvidia
posted 4w ago · verified today
Perplexityai startup
London · unknown
cudagpu-kernelsinferenceinference-enginesgpu-generic
posted 5mo ago · verified today
Perplexityai startup
New York City · Palo Alto +1 more · $220K–$485K/year · mid
cudagpu-kernelsinferenceinference-enginescutlass-cute
posted 5mo ago · verified today
Baseteninference provider
San Francisco · hybrid · $165K–$330K/year · manager
eng-managercpp-langgo-langgpu-genericinference
posted 3mo ago · verified today
Baseteninference provider
Montreal · New York +2 more · hybrid · $165K–$330K/year · unknown
cpp-langnetwork-fabricnvidiasoftware-engineercollectives
posted 6mo ago · verified today
Baseteninference provider
Montreal · New York +2 more · hybrid · $165K–$330K/year · senior
evaluationgpu-genericperformance-engineerpython-langinference
posted 8mo ago · verified today
Baseteninference provider
Montreal · New York +3 more · hybrid · $165K–$330K/year · unknown
inferencekubernetes-opsml-platformpython-langsoftware-engineer
posted 18mo ago · verified today