AI infrastructure roles, filterable by the stack you actually work on.
Distributed training · inference serving · GPU fleets · network fabric — aggregated straight from company boards, never a copy of a copy.
1297 open roles · 85 companies · last verified today
449 roles
DigitalOceanneocloud
Atlanta · *United States · manager
cluster-datacenterdatacenter-engineergpu-genericreliability-sre
posted 20d ago · verified today
DigitalOceanneocloud
Seattle Metro · Wenatchee · manager
cluster-datacenterdatacenter-engineergpu-genericreliability-sre
posted 20d ago · verified today
DigitalOceanneocloud
Austin Metro · Seattle · $107K–$134K/year est. · mid
reliability-sresregpu-generickubernetes-opsgo-lang
posted 7mo ago · verified today
Scalewayneocloud
Paris · hybrid · manager
eng-managercluster-datacenterkubernetes-opsnvidiareliability-sre
posted 3w ago · verified today
DeepInfrainference provider
Bulgaria - Remote · remote · mid
cpp-langcudagpu-genericinferenceinference-engines
posted 5mo ago · verified today
DeepInfrainference provider
Palo Alto, United States · onsite · mid
cudainferencepython-langsoftware-engineercpp-lang
posted 5mo ago · verified today
DeepInfrainference provider
Palo Alto, United States · onsite · $140K–$150K/year est. · junior
inference-enginessoftware-engineercpp-langcudaml-platform
posted 8mo ago · verified today
DeepInfrainference provider
Bulgaria - Remote · remote · junior
cpp-langcudainferenceinference-enginespython-lang
posted 11mo ago · verified today
Sakana AIfrontier lab
Tokyo · unknown
reliability-srenvidiainferencesregpu-generic
added 7d ago · verified today
Fish Audio (39 AI)ai startup
Location not specified · senior
cluster-datacentergpu-genericreliability-sresrekubernetes-ops
added 7d ago · verified today
Waferinference provider
San Francisco · onsite · $200K–$300K/year · unknown
gpu-kernelsinferenceinference-enginescluster-datacenterperformance-engineer
posted 8w ago · verified today
Runwareinference provider
Remote · remote · senior
reliability-sresrego-langkubernetes-opsobservability
posted 8w ago · verified today
Runwareinference provider
United Kingdom · remote · senior
gpu-genericinferencenvidiareliability-sresre
posted 4mo ago · verified today
Relaceinference provider
San Francisco · onsite · mid
software-engineercluster-datacenterscheduling-orchestrationinferenceml-platform
posted 10mo ago · verified today
Exaai startup
Singapore · onsite · 90K–300K SGD/year · unknown
cluster-datacenterkubernetes-opsscheduling-orchestrationsoftware-engineerdistributed-inference
posted 6mo ago · verified today
Inceptionfrontier lab
San Mateo, United States · onsite · senior
inferenceinference-enginespython-langsoftware-engineergpu-generic
posted 6mo ago · verified today
Inceptionfrontier lab
San Mateo, United States · onsite · unknown
reinforcement-learningml-platformpost-trainingray-distributedsoftware-engineer
posted 6mo ago · verified today
Inceptionfrontier lab
San Mateo, United States · onsite · unknown
cudagpu-genericinferenceinference-engineskubernetes-ops
posted 6mo ago · verified today
Arcee AIfrontier lab
San Francisco, CA · SF, CA · unknown
cluster-datacentergpu-generickubernetes-opsreliability-sresre
posted 3mo ago · verified today
Black Forest Labsfrontier lab
San Francisco (United States) · onsite · unknown
cudagpu-genericinferenceinference-enginesperformance-engineer
posted 24mo ago · verified today
Black Forest Labsfrontier lab
Freiburg (Germany) · onsite · unknown
cluster-datacentergo-langkubernetes-opsnvidiaobservability
posted 12mo ago · verified today
Baseteninference provider
San Francisco · hybrid · $170K–$230K/year · manager
cluster-datacentergo-langgpu-generickubernetes-opsnvidia
posted 8d ago · verified today
Fireworks AIinference provider
New York · San Mateo · hybrid · $200K–$230K/year · unknown
scheduling-orchestrationcpp-langnetwork-fabricpython-langstorage-checkpointing
posted 8d ago · verified today
Nscaleneocloud
Houston · New York +2 more · $130K–$200K/year est. · mid
reliability-sresrego-langobservabilitypython-lang
posted 8d ago · verified today
Nscaleneocloud
Houston · San Francisco +1 more · $170K–$265K · senior
reliability-sresrego-langkubernetes-opsobservability
posted 5mo ago · verified today
Cerebraschip vendor
Toronto, CAN · hybrid · senior
cluster-datacentergo-langkubernetes-opsscheduling-orchestrationobservability
posted 8d ago · verified today
Nscaleneocloud
UK · manager
eng-manageransiblecluster-datacentergpu-genericobservability
posted 9d ago · verified today
FluidStackneocloud
Travelling · U.S. Remote · remote · $258K–$300K/year · staff plus
network-engineernetwork-fabricreliability-sre
posted 10d ago · verified today
Harveyai startup
San Francisco · hybrid · $231K–$340K/year · staff plus
inferenceinference-enginessoftware-engineerml-platformobservability
posted 9d ago · verified today
Nscaleneocloud
London · UK · senior
kubernetes-opsgo-langobservabilitypython-langreliability-sre
posted 10d ago · verified today