AI infrastructure roles, filterable by the stack you actually work on.

Distributed training · inference serving · GPU fleets · network fabric — aggregated straight from company boards, never a copy of a copy.

1294 open roles · 85 companies · last verified today

Crusoeneocloud

San Francisco, CA - US · Sunnyvale, CA - US · onsite · manager

eng-managerkubernetes-opsslurm-admincluster-datacenterscheduling-orchestration

posted 3w ago · verified today

OpenAIfrontier lab

San Francisco · hybrid · $266K–$445K/year · unknown

software-engineerinference-enginesinferencekv-cache-systemsscheduling-orchestration

posted 3w ago · verified today

Lambdaneocloud

Bellevue Office · San Francisco Office (Fremont St) +1 more · remote · $297K–$440K · manager

eng-managercluster-datacentergpu-genericreliability-srescheduling-orchestration

posted 3w ago · verified today

Nebiusneocloud

New York City, New York, United States · Remote - United States +1 more · remote · $200K–$245K · senior

gpu-genericnvidiasolutions-architectinferenceinference-engines

posted 7mo ago · verified today

Amazon (AWS)hyperscaler

Arlington, Virginia, USA · Denver, Colorado, USA +1 more · onsite · senior

gpu-generickubernetes-opsscheduling-orchestrationsolutions-architectml-platform

posted 4w ago · verified today

Amazon (AWS)hyperscaler

Arlington, Virginia, USA · Denver, Colorado, USA +1 more · onsite · senior

cluster-datacentergpu-generickubernetes-opsnvidiascheduling-orchestration

posted 4w ago · verified today

SambaNovachip vendor

Bengaluru, India · Bengaluru, Karnataka, India · 10300K–12500K INR · staff plus

inference-enginessoftware-engineergo-langinferencekubernetes-ops

posted 4w ago · verified today

SambaNovachip vendor

Austin, Texas, United States · Austin, TX +2 more · $210K–$280K · staff plus

inferenceml-platformobservabilityreliability-sresre

posted 9mo ago · verified today

CoreWeaveneocloud

Bellevue, WA · San Francisco, CA +1 more · $198K–$264K/year est. · staff plus

amdcluster-datacenterdeepspeed-libdistributed-inferencefsdp

posted 4w ago · verified today

Amazon (AWS)hyperscaler

Seattle, Washington, USA · onsite · mid

software-engineertrainiumml-platformscheduling-orchestrationgo-lang

posted 11w ago · verified today

Amazon (AWS)hyperscaler

Bellevue, Washington, USA · onsite · mid

software-engineerml-platformtrainiumscheduling-orchestrationfsdp

posted 6w ago · verified today

Amazon (AWS)hyperscaler

Seattle, Washington, USA · onsite · mid

gpu-generickubernetes-opssoftware-engineernvidiaobservability

posted 9w ago · verified today

Amazon (AWS)hyperscaler

Boston, Massachusetts, USA · Seattle, Washington, USA +1 more · onsite · senior

cudacutlass-cutedistributed-inferenceevaluationflash-attention

posted 4w ago · verified today

CoreWeaveneocloud

Bellevue, WA · Livingston, NJ +4 more · $139K–$242K/year est. · senior

go-langgpu-generickubernetes-opssoftware-engineerlinux-kernel

posted 6mo ago · verified today

FluidStackneocloud

Austin, TX · New York, NY +2 more · onsite · $224K–$344K/year · manager

cluster-datacentereng-managergpu-virtualizationkubernetes-opsreliability-sre

posted 8w ago · verified today

FluidStackneocloud

Austin, TX · New York, NY +2 more · onsite · $173K–$224K/year · unknown

reliability-sresreinfiniband-opskubernetes-opsnccl-lib

posted 8w ago · verified today

FluidStackneocloud

Austin, TX · New York, NY +2 more · onsite · $208K–$269K/year · unknown

observabilityreliability-sresrecluster-datacentergpu-generic

posted 10w ago · verified today

FluidStackneocloud

Austin, TX · New York, NY +2 more · onsite · $208K–$269K/year · unknown

reliability-sresregpu-genericobservabilitycluster-datacenter

posted 8w ago · verified today

FluidStackneocloud

Austin, TX · New York, NY +2 more · onsite · $208K–$269K/year · unknown

reliability-sresrecluster-datacentergpu-genericobservability

posted 3mo ago · verified today