AI infrastructure roles, filterable by the stack you actually work on.
Distributed training · inference serving · GPU fleets · network fabric — aggregated straight from company boards, never a copy of a copy.
1294 open roles · 85 companies · last verified today
685 roles
CoreWeaveneocloud
Bellevue, WA · Livingston, NJ +4 more · $182K–$242K/year est. · senior
gpu-generickubernetes-opssoftware-engineergo-langlinux-kernel
posted 19d ago · verified today
Crusoeneocloud
Shakopee, MN - US · onsite · manager
cluster-datacenterdatacenter-engineereng-managergpu-generic
posted 19d ago · verified today
Lambdaneocloud
Atlanta, GA - Data Center · $89K–$119K · unknown
cluster-datacenterdatacenter-engineernvidia
posted 19d ago · verified today
Anthropicfrontier lab
Remote-Friendly, United States · Remote-Friendly US (Travel Required) · remote · $320K–$405K · senior
cluster-datacenterdatacenter-engineer
posted 19d ago · verified today
SambaNovachip vendor
Stockholm, Sweden · staff plus
cpp-langsoftware-engineerinferencelinux-kernelnetwork-fabric
posted 19d ago · verified today
Amazon (AWS)hyperscaler
Austin, Texas, USA · Cupertino, California, USA · onsite · junior
custom-asicdatacenter-engineerpython-langtrainiumsoftware-engineer
posted 3w ago · verified today
xAIfrontier lab
Memphis, TN · Southaven, MS · unknown
network-engineercluster-datacenternetwork-fabricinfiniband-opsroce-net
posted 20d ago · verified today
Crusoeneocloud
San Francisco, CA - US · Sunnyvale, CA - US · onsite · manager
network-fabriccluster-datacenterinfiniband-opsroce-netgpu-generic
posted 20d ago · verified today
OpenAIfrontier lab
San Francisco · hybrid · $177K–$300K/year · senior
datacenter-engineercluster-datacenterpython-langreliability-sre
posted 20d ago · verified today
Amazon (AWS)hyperscaler
Cupertino, California, USA · Denver, Colorado, USA +1 more · onsite · senior
datacenter-engineercluster-datacentergpu-genericreliability-sre
posted 3w ago · verified today
Amazon (AWS)hyperscaler
Cupertino, California, USA · Denver, Colorado, USA +1 more · onsite · senior
datacenter-engineergpu-genericreliability-srecluster-datacenter
posted 3w ago · verified today
Amazon (AWS)hyperscaler
Cupertino, California, USA · Denver, Colorado, USA +1 more · onsite · mid
cluster-datacentergpu-genericreliability-sredatacenter-engineerpython-lang
posted 3w ago · verified today
FluidStackneocloud
Austin, TX · CA +4 more · onsite · $208K–$263K/year · manager
datacenter-engineernvidiacluster-datacentergpu-genericnetwork-fabric
posted 3w ago · verified today
Amazon (AWS)hyperscaler
Cupertino, California, USA · onsite · senior
training-frameworksdistributed-inferencecluster-datacentercollectivesdatacenter-engineer
posted 3w ago · verified today
Mistral AIfrontier lab
Montréal · New York +2 more · remote · unknown
kubernetes-opsreliability-srescheduling-orchestrationslurm-adminsre
posted 10w ago · verified today
Mistral AIfrontier lab
Amsterdam · Berlin +4 more · remote · unknown
kubernetes-opsreliability-srescheduling-orchestrationslurm-admincluster-datacenter
posted 10w ago · verified today
Amazon (AWS)hyperscaler
Arlington, Virginia, USA · Denver, Colorado, USA +1 more · onsite · unknown
kubernetes-opsnvidiascheduling-orchestrationgpu-genericsolutions-architect
posted 3w ago · verified today
Nebiusneocloud
Remote - Europe · remote · senior
datacenter-engineercluster-datacenternetwork-fabric
posted 3w ago · verified today
Amazon (AWS)hyperscaler
New York, New York, USA · Sunnyvale, California, USA · onsite · senior
kubernetes-opsml-platforminferenceinference-enginesquantization
posted 3w ago · verified today
FluidStackneocloud
Austin, TX · New York, NY +2 more · hybrid · $162K–$202K/year · unknown
network-engineernetwork-fabriccluster-datacenter
posted 14d ago · verified today
Anthropicfrontier lab
New York City, NY · San Francisco, CA +1 more · $320K–$485K · staff plus
inferenceinference-enginesscheduling-orchestrationsoftware-engineerdistributed-inference
posted 3w ago · verified today
Crusoeneocloud
San Francisco, CA - US · Sunnyvale, CA - US · onsite · manager
eng-managerkubernetes-opsslurm-admincluster-datacenterscheduling-orchestration
posted 3w ago · verified today
Lambdaneocloud
Bellevue Office · San Francisco Office (Fremont St) +1 more · remote · $297K–$440K · manager
eng-managercluster-datacentergpu-genericreliability-srescheduling-orchestration
posted 3w ago · verified today
Anthropicfrontier lab
Remote-Friendly (Travel Required) · Remote-Friendly US (Travel Required) +1 more · remote · $320K–$405K · manager
cluster-datacenterdatacenter-engineerreliability-sregpu-generic
posted 3w ago · verified today
Amazon (AWS)hyperscaler
Seattle, Washington, USA · onsite · unknown
software-engineertrainiumcluster-datacentercpp-langscheduling-orchestration
posted 3w ago · verified today
xAIfrontier lab
Memphis, TN · Southaven, MS · senior
cluster-datacenterdatacenter-engineer
posted 3w ago · verified today
CoreWeaveneocloud
Bellevue, WA · Las Vegas, NV - DC +4 more · $120K–$145K/year est. · manager
cluster-datacenterdatacenter-engineereng-manager
posted 3w ago · verified today
Lambdaneocloud
Elk Grove Village, IL - Data Center · remote · $137K–$183K · manager
cluster-datacenterdatacenter-engineereng-managerinfiniband-opsnetwork-fabric
posted 3w ago · verified today
Amazon (AWS)hyperscaler
Arlington, Virginia, USA · Denver, Colorado, USA +1 more · onsite · senior
gpu-generickubernetes-opsscheduling-orchestrationsolutions-architectml-platform
posted 4w ago · verified today
Amazon (AWS)hyperscaler
Arlington, Virginia, USA · Denver, Colorado, USA +1 more · onsite · senior
cluster-datacentergpu-generickubernetes-opsnvidiascheduling-orchestration
posted 4w ago · verified today