AI infrastructure roles, filterable by the stack you actually work on.

Distributed training · inference serving · GPU fleets · network fabric — aggregated straight from company boards, never a copy of a copy.

1297 open roles · 85 companies · last verified today

196 roles

nvidia

FluidStackneocloud

Austin, TX · CA +4 more · onsite · $208K–$263K/year · manager

datacenter-engineernvidiacluster-datacentergpu-genericnetwork-fabric

posted 3w ago · verified today

Amazon (AWS)hyperscaler

Arlington, Virginia, USA · Denver, Colorado, USA +1 more · onsite · unknown

kubernetes-opsnvidiascheduling-orchestrationgpu-genericsolutions-architect

posted 3w ago · verified today

Nebiusneocloud

New York City, New York, United States · Remote - United States +1 more · remote · $200K–$245K · senior

gpu-genericnvidiasolutions-architectinferenceinference-engines

posted 7mo ago · verified today

Lambdaneocloud

Elk Grove Village, IL - Data Center · remote · $137K–$183K · manager

cluster-datacenterdatacenter-engineereng-managerinfiniband-opsnetwork-fabric

posted 3w ago · verified today

Amazon (AWS)hyperscaler

Arlington, Virginia, USA · Denver, Colorado, USA +1 more · onsite · senior

gpu-generickubernetes-opsscheduling-orchestrationsolutions-architectml-platform

posted 4w ago · verified today

Amazon (AWS)hyperscaler

Arlington, Virginia, USA · Denver, Colorado, USA +1 more · onsite · senior

cluster-datacentergpu-generickubernetes-opsnvidiascheduling-orchestration

posted 4w ago · verified today

Nebiusneocloud

Remote - United States · United States · remote · $180K–$220K/year est. · senior

gpu-genericnvidiasolutions-architectcluster-datacenterperformance-engineer

posted 5mo ago · verified today

SambaNovachip vendor

Austin, Texas, United States · Austin, TX +2 more · $210K–$280K · staff plus

inferenceml-platformobservabilityreliability-sresre

posted 9mo ago · verified today

CoreWeaveneocloud

Bellevue, WA · San Francisco, CA +1 more · $198K–$264K/year est. · staff plus

amdcluster-datacenterdeepspeed-libdistributed-inferencefsdp

posted 4w ago · verified today

Amazon (AWS)hyperscaler

Seattle, Washington, USA · onsite · manager

eng-managernvidiacluster-datacentercpp-langml-platform

posted 5w ago · verified today

Amazon (AWS)hyperscaler

Seattle, Washington, USA · onsite · manager

eng-managernetwork-fabricrdma-verbsnvidia

posted 5w ago · verified today

Amazon (AWS)hyperscaler

Seattle, Washington, USA · onsite · manager

eng-managernetwork-fabricrdma-verbsnvidia

posted 6w ago · verified today

Amazon (AWS)hyperscaler

San Francisco, California, USA · onsite · senior

cudagpu-kernelsnvidiapython-langresearch-engineer

posted 3mo ago · verified today

Amazon (AWS)hyperscaler

Cupertino, California, USA · Seattle, Washington, USA · onsite · manager

collectivescpp-langcudagpu-kernelsnccl-lib

posted 6w ago · verified today

Amazon (AWS)hyperscaler

Cupertino, California, USA · Seattle, Washington, USA · onsite · senior

gpu-genericnvidiacluster-datacenterreliability-sre

posted 3mo ago · verified today

Amazon (AWS)hyperscaler

Seattle, Washington, USA · onsite · mid

gpu-generickubernetes-opssoftware-engineernvidiaobservability

posted 9w ago · verified today

Amazon (AWS)hyperscaler

Boston, Massachusetts, USA · Seattle, Washington, USA +1 more · onsite · senior

cudacutlass-cutedistributed-inferenceevaluationflash-attention

posted 4w ago · verified today

Amazon (AWS)hyperscaler

Herndon, Virginia, USA · New York, New York, USA +1 more · onsite · senior

solutions-architectgpu-genericinference-enginesmodel-parallelismtraining-frameworks

posted 3mo ago · verified today

Amazon (AWS)hyperscaler

Toronto, Ontario, CAN · onsite · senior

gpu-kernelsperformance-engineertrainiumcudatriton-lang

posted 15d ago · verified today

Nebiusneocloud

Reykjanesbær · Reykjanesbær, Iceland · junior

cluster-datacenterdatacenter-engineernvidia

posted 4w ago · verified today

Perplexityai startup

New York City · Palo Alto +1 more · $220K–$485K/year · mid

cudagpu-kernelsinferenceinference-enginescutlass-cute

posted 5mo ago · verified today

Baseteninference provider

San Francisco · hybrid · $165K–$330K/year · manager

eng-managercpp-langgo-langgpu-genericinference

posted 3mo ago · verified today

Baseteninference provider

Montreal · New York +2 more · hybrid · $165K–$330K/year · unknown

cpp-langnetwork-fabricnvidiasoftware-engineercollectives

posted 6mo ago · verified today

Baseteninference provider

Montreal · New York +2 more · hybrid · $165K–$330K/year · senior

evaluationgpu-genericperformance-engineerpython-langinference

posted 8mo ago · verified today

Baseteninference provider

Montreal · New York +3 more · hybrid · $165K–$330K/year · unknown

inferencekubernetes-opsml-platformpython-langsoftware-engineer

posted 18mo ago · verified today