AI infrastructure roles, filterable by the stack you actually work on.
Distributed training · inference serving · GPU fleets · network fabric — aggregated straight from company boards, never a copy of a copy.
1297 open roles · 85 companies · last verified today
145 roles
Cerebraschip vendor
Canada · United States · unknown
cpp-langpython-langreliability-sresoftware-engineercustom-asic
posted 1d ago · verified today
Matter Intelligenceai startup
San Francisco · onsite · unknown
observabilityreliability-sresoftware-engineerml-platformsre
posted 1d ago · verified today
DigitalOceanneocloud
DigitalOcean Seattle Office · Seattle · $167K–$209K/year est. · senior
go-langinferencekubernetes-opsgpu-genericinference-engines
posted 5d ago · verified today
Runpodneocloud
Remote - USA · remote · senior
cluster-datacentergo-langkubernetes-opslinux-kernelnetwork-fabric
posted 7d ago · verified today
Amazon (AWS)hyperscaler
Austin, Texas, USA · onsite · mid
reliability-sresretrainiumpython-lang
posted 7d ago · verified today
Reactorinference provider
San Francisco · onsite · unknown
gpu-generickubernetes-opsml-platformobservabilityterraform-iac
added 7d ago · verified today
Boson AIfrontier lab
Toronto · onsite · CA$125K–CA$250K/year · unknown
nvidiareliability-sresreansiblecluster-datacenter
posted 9w ago · verified today
DigitalOceanneocloud
Bangalore Metro · Bengaluru · senior
cudakubernetes-opsnvidiaamdrocm-hip
posted 11w ago · verified today
DigitalOceanneocloud
Austin Metro · Seattle · $107K–$134K/year est. · mid
reliability-sresregpu-generickubernetes-opsgo-lang
posted 7mo ago · verified today
Sakana AIfrontier lab
Tokyo · unknown
reliability-srenvidiainferencesregpu-generic
added 7d ago · verified today
Fish Audio (39 AI)ai startup
Location not specified · senior
cluster-datacentergpu-genericreliability-sresrekubernetes-ops
added 7d ago · verified today
Runwareinference provider
Remote · remote · senior
reliability-sresrego-langkubernetes-opsobservability
posted 8w ago · verified today
Runwareinference provider
United Kingdom · remote · senior
gpu-genericinferencenvidiareliability-sresre
posted 4mo ago · verified today
Arcee AIfrontier lab
San Francisco, CA · SF, CA · unknown
cluster-datacentergpu-generickubernetes-opsreliability-sresre
posted 3mo ago · verified today
Black Forest Labsfrontier lab
Freiburg (Germany) · onsite · unknown
cluster-datacentergo-langkubernetes-opsnvidiaobservability
posted 12mo ago · verified today
Nscaleneocloud
Houston · New York +2 more · $130K–$200K/year est. · mid
reliability-sresrego-langobservabilitypython-lang
posted 8d ago · verified today
Nscaleneocloud
Houston · San Francisco +1 more · $170K–$265K · senior
reliability-sresrego-langkubernetes-opsobservability
posted 5mo ago · verified today
Lila Sciencesai startup
Alewife, Cambridge, MA · Cambridge, MA USA · $192K–$272K · staff plus
inferenceinference-engineskubernetes-opsml-platformnvidia
posted 7w ago · verified today
Luma AIai startup
Redwood City, CA · hybrid · $235K–$353K/year · staff plus
gpu-genericreliability-sresrecluster-datacenterkubernetes-ops
posted 7w ago · verified today
Luma AIai startup
Redwood City, CA · hybrid · unknown
inferenceinference-enginessoftware-engineerkubernetes-opspython-lang
posted 7w ago · verified today
Luma AIai startup
Redwood City, CA · hybrid · $168K–$252K/year · senior
reliability-sresregpu-genericnvidiaamd
posted 7w ago · verified today
ElevenLabsai startup
United States · remote · unknown
gpu-genericnvidiacluster-datacentercudareliability-sre
posted 13d ago · verified today
Tenstorrentchip vendor
Tokyo · Tokyo, Japan · unknown
cluster-datacenterdatacenter-engineergpu-genericreliability-srenetwork-fabric
posted 12mo ago · verified today
Generalistai startup
San Francisco Bay Area (San Mateo) or Boston (Somerville) · onsite · $260K–$350K/year · unknown
gpu-genericnvidiasoftware-engineerinferencekubernetes-ops
posted 7mo ago · verified today
Lightning AIai startup
Singapore · 165K–205K SGD · senior
reliability-sresreansiblecluster-datacenterkubernetes-ops
posted 7w ago · verified today
Lightning AIai startup
London, England, United Kingdom · London, UK +7 more · remote · $160K–$200K · senior
reliability-sresreansiblecluster-datacenterkubernetes-ops
posted 3mo ago · verified today
Lightning AIai startup
London, England, United Kingdom · London, UK +7 more · remote · $180K–$220K · unknown
cluster-datacenterdatacenter-engineergpu-genericnvidiapython-lang
posted 4mo ago · verified today
Lightning AIai startup
New York, New York, United States · San Francisco, California +3 more · $115K–$140K · unknown
cudagpu-generickubernetes-opsnccl-libobservability
posted 3mo ago · verified today
Lightning AIai startup
London, England, United Kingdom · London, UK · £75K–£95K · unknown
cudagpu-generickubernetes-opsml-platformnccl-lib
posted 3mo ago · verified today
Lightning AIai startup
Philippines · Remote +1 more · unknown
reliability-sresrecluster-datacentergpu-generickubernetes-ops
posted 4mo ago · verified today