AI infrastructure roles, filterable by the stack you actually work on.

Distributed training · inference serving · GPU fleets · network fabric — aggregated straight from company boards, never a copy of a copy.

1293 open roles · 85 companies · last verified today

453 roles

reliability-sre

Nscaleneocloud

Houston · San Francisco +1 more · $100K–$140K · unknown

gpu-genericreliability-sresrecluster-datacenternvidia

posted 3mo ago · verified today

Nscaleneocloud

Houston · New York +3 more · $200K–$300K · manager

cluster-datacenterdatacenter-engineerreliability-sregpu-genericeng-manager

posted 5mo ago · verified today

Nscaleneocloud

Houston · New York +2 more · $230K–$343K · manager

eng-managercluster-datacenterscheduling-orchestrationslurm-admingpu-generic

posted 5w ago · verified today

Amazon (AWS)hyperscaler

Arlington, Virginia, USA · Denver, Colorado, USA +1 more · onsite · senior

solutions-architectkubernetes-opsscheduling-orchestrationcluster-datacentergpu-generic

posted 16d ago · verified today

Anthropicfrontier lab

New York City, NY · San Francisco, CA +1 more · $405K–$625K · manager

eng-managerinferencescheduling-orchestrationinference-enginescluster-datacenter

posted 15d ago · verified today

CoreWeaveneocloud

Livingston, NJ · Manhattan, NY +2 more · $153K–$204K/year est. · senior

network-fabricobservabilitygo-langpython-langreliability-sre

posted 19d ago · verified today

SambaNovachip vendor

Stockholm, Sweden · staff plus

cpp-langsoftware-engineerinferencelinux-kernelnetwork-fabric

posted 19d ago · verified today

OpenAIfrontier lab

San Francisco · hybrid · $177K–$300K/year · senior

datacenter-engineercluster-datacenterpython-langreliability-sre

posted 20d ago · verified today

Amazon (AWS)hyperscaler

Cupertino, California, USA · Denver, Colorado, USA +1 more · onsite · senior

datacenter-engineercluster-datacentergpu-genericreliability-sre

posted 3w ago · verified today

Amazon (AWS)hyperscaler

Cupertino, California, USA · Denver, Colorado, USA +1 more · onsite · senior

datacenter-engineergpu-genericreliability-srecluster-datacenter

posted 3w ago · verified today

Amazon (AWS)hyperscaler

Cupertino, California, USA · Denver, Colorado, USA +1 more · onsite · mid

cluster-datacentergpu-genericreliability-sredatacenter-engineerpython-lang

posted 3w ago · verified today

Mistral AIfrontier lab

Montréal · New York +2 more · remote · unknown

kubernetes-opsreliability-srescheduling-orchestrationslurm-adminsre

posted 10w ago · verified today

Mistral AIfrontier lab

Amsterdam · Berlin +4 more · remote · unknown

kubernetes-opsreliability-srescheduling-orchestrationslurm-admincluster-datacenter

posted 10w ago · verified today

Amazon (AWS)hyperscaler

Arlington, Virginia, USA · Denver, Colorado, USA +1 more · onsite · unknown

kubernetes-opsnvidiascheduling-orchestrationgpu-genericsolutions-architect

posted 3w ago · verified today

Amazon (AWS)hyperscaler

New York, New York, USA · Sunnyvale, California, USA · onsite · senior

kubernetes-opsml-platforminferenceinference-enginesquantization

posted 3w ago · verified today

OpenAIfrontier lab

San Francisco · hybrid · $266K–$445K/year · unknown

software-engineercpp-langpython-langobservabilityperformance-engineer

posted 3w ago · verified today

Anthropicfrontier lab

New York City, NY · San Francisco, CA +1 more · $320K–$485K · staff plus

inferenceinference-enginesscheduling-orchestrationsoftware-engineerdistributed-inference

posted 3w ago · verified today

Crusoeneocloud

San Francisco, CA - US · Sunnyvale, CA - US · onsite · manager

eng-managerkubernetes-opsslurm-admincluster-datacenterscheduling-orchestration

posted 3w ago · verified today

Lambdaneocloud

Bellevue Office · San Francisco Office (Fremont St) +1 more · remote · $297K–$440K · manager

eng-managercluster-datacentergpu-genericreliability-srescheduling-orchestration

posted 3w ago · verified today