AI infrastructure roles, filterable by the stack you actually work on.
Distributed training · inference serving · GPU fleets · network fabric — aggregated straight from company boards, never a copy of a copy.
1297 open roles · 85 companies · last verified today
145 roles
Etchedchip vendor
San Jose · onsite · junior
cluster-datacentercpp-langnetwork-fabricpython-langreliability-sre
posted 3mo ago · verified today
Coherefrontier lab
Montreal · New York +2 more · remote · unknown
kubernetes-opssrecpp-langgo-langgpu-generic
posted 8mo ago · verified today
Reflection AIfrontier lab
London · New York, NY +1 more · onsite · unknown
kubernetes-opscluster-datacenternccl-libreliability-sregpu-generic
posted 6mo ago · verified today
RadixArkai startup
Palo Alto, CA · Palo Alto Office · $200K–$400K/year est. · mid
cudagpu-kernelsinference-enginescpp-langrocm-hip
posted 7w ago · verified today
falinference provider
Remote - APAC · remote · senior
inferencekubernetes-opsobservabilitypython-langreliability-sre
posted 11w ago · verified today
falinference provider
San Francisco · $180K–$250K/year · mid
ansiblecudanvidiaobservabilitypython-lang
posted 6mo ago · verified today
falinference provider
Remote - Global · unknown
python-langgpu-genericcudanvidiareliability-sre
posted 6mo ago · verified today
Parasailinference provider
San Mateo · senior
inferencereliability-sresreobservabilitykubernetes-ops
posted 7w ago · verified today
Inferactinference provider
Remote · remote · $200K–$400K/year · unknown
go-langkubernetes-opspython-langrust-langterraform-iac
posted 3w ago · verified today
Inferactinference provider
San Francisco · onsite · $200K–$400K/year · unknown
cluster-datacentergpu-genericreliability-sreansiblekubernetes-ops
posted 3w ago · verified today
Inferactinference provider
San Francisco · onsite · $200K–$400K/year · unknown
reliability-sresreinferenceobservabilityvllm-engine
posted 3w ago · verified today
NexGen Cloudneocloud
London · Nottingham · senior
kubernetes-opscluster-datacentergpu-genericnvidiascheduling-orchestration
posted 5mo ago · verified today
NexGen Cloudneocloud
Canada - Remote · Quebec, Canada · unknown
kubernetes-opscluster-datacenterdatacenter-engineergpu-genericreliability-sre
posted 14d ago · verified today
San Francisco Compute Companyneocloud
Remote · San Francisco, CA · hybrid · $220K–$300K/year · senior
cluster-datacenterdatacenter-engineergpu-genericnetwork-fabricinfiniband-ops
posted 1d ago · verified today
Radiantneocloud
London · hybrid · senior
ansibleobservabilitycluster-datacentergo-langgpu-generic
posted 3mo ago · verified today
Radiantneocloud
Gloucestershire · London · hybrid · senior
reliability-sresrecluster-datacentergpu-genericnvidia
posted 3mo ago · verified today
Radiantneocloud
Gloucestershire · hybrid · senior
ansibleobservabilitypython-langreliability-sresre
posted 5mo ago · verified today
Runpodneocloud
Remote - USA · remote · senior
observabilityreliability-sresrego-langpython-lang
posted 9w ago · verified today
TensorWaveneocloud
Las Vegas, Nevada · onsite · mid
amdgpu-generickubernetes-opsreliability-srescheduling-orchestration
posted 3w ago · verified today
TensorWaveneocloud
Remote · remote · staff plus
kubernetes-opscluster-datacenterscheduling-orchestrationreliability-sresoftware-engineer
posted 10w ago · verified today
TensorWaveneocloud
Remote · remote · senior
amdterraform-iacobservabilityreliability-sresre
posted 3mo ago · verified today
Liquid AIfrontier lab
San Francisco · hybrid · unknown
cluster-datacentergpu-genericreliability-sresoftware-engineersre
posted 7w ago · verified today
Anthropicfrontier lab
New York City, NY · Remote-Friendly (Travel-Required) +2 more · remote · $320K–$485K · staff plus
reliability-sresreml-platformrust-langinference
posted 12d ago · verified today
Thinking Machines Labfrontier lab
New York · San Francisco · onsite · unknown
reliability-sresrefine-tuninggpu-generickubernetes-ops
posted 15d ago · verified today
Thinking Machines Labfrontier lab
San Francisco · hybrid · unknown
cluster-datacenterreliability-sresregpu-generickubernetes-ops
posted 11w ago · verified today
Thinking Machines Labfrontier lab
New York · San Francisco · onsite · unknown
post-trainingreliability-sresrepython-langreinforcement-learning
posted 15d ago · verified today
Nscaleneocloud
Austin, TX · Houston +3 more · staff plus
ansiblecluster-datacentercollectivescpp-langdeepspeed-lib
posted 13d ago · verified today
xAIfrontier lab
Memphis, TN · Southaven, MS · senior
reliability-sresrecluster-datacenterobservabilitypython-lang
posted 14d ago · verified today
Nscaleneocloud
US · $160K–$230K · senior
observabilitygo-langkubernetes-opspython-langansible
posted 3w ago · verified today
Nscaleneocloud
EMEA · Germany +6 more · senior
ansiblecluster-datacenterdatacenter-engineerpython-langreliability-sre
posted 5mo ago · verified today