AI infrastructure roles, filterable by the stack you actually work on.
Distributed training · inference serving · GPU fleets · network fabric — aggregated straight from company boards, never a copy of a copy.
1296 open roles · 85 companies · last verified today
343 roles
Prime Intellectneocloud
San Francisco · unknown
post-trainingreinforcement-learningresearch-engineerevaluationml-platform
posted 10w ago · verified today
Prime Intellectneocloud
Remote · San Francisco · hybrid · unknown
cluster-datacentergpu-genericinfiniband-opskubernetes-opsnetwork-fabric
posted 10w ago · verified today
Prime Intellectneocloud
New York City, USA · hybrid · unknown
evaluationpost-trainingdistributed-inferenceray-distributedreinforcement-learning
posted 10w ago · verified today
Prime Intellectneocloud
San Francisco · unknown
python-langsoftware-engineercluster-datacentergpu-generickubernetes-ops
posted 10w ago · verified today
Radiantneocloud
London · hybrid · mid
datacenter-engineercluster-datacenternvidiagpu-genericobservability
posted 3w ago · verified today
Radiantneocloud
London · hybrid · senior
cluster-datacenterdatacenter-engineergpu-genericinfiniband-opsnetwork-fabric
posted 6w ago · verified today
Radiantneocloud
London · hybrid · senior
ansibleobservabilitycluster-datacentergo-langgpu-generic
posted 3mo ago · verified today
Radiantneocloud
Gloucestershire · London · hybrid · senior
reliability-sresrecluster-datacentergpu-genericnvidia
posted 3mo ago · verified today
Radiantneocloud
Gloucestershire · hybrid · senior
kubernetes-opsreliability-sreansibleobservabilitypython-lang
posted 4mo ago · verified today
Radiantneocloud
Gloucestershire · hybrid · senior
ansibleobservabilitypython-langreliability-sresre
posted 5mo ago · verified today
Radiantneocloud
Gloucestershire · hybrid · senior
ansiblenetwork-engineernetwork-fabricpython-langkubernetes-ops
posted 5mo ago · verified today
Runpodneocloud
Remote - USA · remote · senior
observabilityreliability-sresrego-langpython-lang
posted 9w ago · verified today
Runpodneocloud
Remote - USA · remote · manager
eng-manageransiblecluster-datacentergpu-genericinfiniband-ops
posted 9w ago · verified today
TensorWaveneocloud
Las Vegas, Nevada · onsite · mid
amdgpu-generickubernetes-opsreliability-srescheduling-orchestration
posted 3w ago · verified today
TensorWaveneocloud
Las Vegas, Nevada · Remote · remote · manager
network-fabricpython-langsoftware-engineerobservabilityeng-manager
posted 5w ago · verified today
TensorWaveneocloud
Miami, Florida · New Kensington, PA (Monroeville +2 more · onsite · manager
cluster-datacentereng-managerobservabilityreliability-sre
posted 11w ago · verified today
TensorWaveneocloud
Remote · remote · staff plus
network-engineernetwork-fabricroce-netansiblecluster-datacenter
posted 6w ago · verified today
TensorWaveneocloud
Remote · remote · staff plus
kubernetes-opscluster-datacenterscheduling-orchestrationreliability-sresoftware-engineer
posted 10w ago · verified today
TensorWaveneocloud
Remote · remote · senior
amdterraform-iacobservabilityreliability-sresre
posted 3mo ago · verified today
TensorWaveneocloud
Anniston, AL (Birmingham Area) · onsite · manager
cluster-datacenterdatacenter-engineerobservabilityreliability-sre
posted 3w ago · verified today
Ineffable Intelligencefrontier lab
London · onsite · unknown
cluster-datacenterkubernetes-opsscheduling-orchestrationobservabilitypython-lang
posted 5w ago · verified today
Poolsidefrontier lab
Remote (EMEA) · remote · unknown
gpu-genericinferencescheduling-orchestrationgo-langinference-engines
posted 9w ago · verified today
Together AIneocloud
Amsterdam · senior
go-langkubernetes-opsml-platformpython-langobservability
posted 10d ago · verified today
Baseteninference provider
San Francisco · hybrid · $225K–$235K/year · senior
cluster-datacentergpu-genericreliability-sreobservabilitynvidia
posted 11d ago · verified today
Mistral AIfrontier lab
Amsterdam · Lausanne +3 more · hybrid · unknown
post-trainingreinforcement-learningresearch-engineerevaluationfine-tuning
posted 12d ago · verified today
Thinking Machines Labfrontier lab
New York · San Francisco · onsite · unknown
reliability-sresrefine-tuninggpu-generickubernetes-ops
posted 15d ago · verified today
Thinking Machines Labfrontier lab
San Francisco · hybrid · unknown
distributed-inferenceinference-enginesresearch-engineergpu-genericinference
posted 6w ago · verified today
Thinking Machines Labfrontier lab
New York · San Francisco · onsite · unknown
post-trainingreliability-sresrepython-langreinforcement-learning
posted 15d ago · verified today
Cerebraschip vendor
Sunnyvale, CA · Toronto, CAN · hybrid · unknown
python-langsoftware-engineercluster-datacentercpp-langreliability-sre
posted 13d ago · verified today
Amazon (AWS)hyperscaler
Palo Alto, California, USA · Seattle, Washington, USA · onsite · mid
ml-platformsoftware-engineerinference-enginesscheduling-orchestrationinference
posted 13d ago · verified today