AI infrastructure roles, filterable by the stack you actually work on.
Distributed training · inference serving · GPU fleets · network fabric — aggregated straight from company boards, never a copy of a copy.
1296 open roles · 85 companies · last verified today
123 roles
Lightning AIai startup
London, England, United Kingdom · London, UK · £75K–£95K · unknown
cudagpu-generickubernetes-opsml-platformnccl-lib
posted 3mo ago · verified today
Lightning AIai startup
Philippines · Remote +1 more · unknown
reliability-sresrecluster-datacentergpu-generickubernetes-ops
posted 4mo ago · verified today
Coherefrontier lab
Canada · Montreal +4 more · remote · unknown
software-engineerstorage-checkpointingkubernetes-opsgo-langobject-storage
posted 4mo ago · verified today
Coherefrontier lab
London · Montreal +4 more · remote · senior
training-frameworksdistributed-inferencepre-trainingcudakubernetes-ops
posted 9mo ago · verified today
Reflection AIfrontier lab
London · New York, NY +1 more · onsite · manager
eng-managerkubernetes-opscluster-datacenterscheduling-orchestrationgpu-generic
posted 4w ago · verified today
Reflection AIfrontier lab
London · New York, NY +1 more · onsite · unknown
pre-trainingmodel-parallelismsoftware-engineertraining-frameworkscollectives
posted 5mo ago · verified today
Reflection AIfrontier lab
London · New York, NY +1 more · onsite · unknown
kubernetes-opscluster-datacenternccl-libreliability-sregpu-generic
posted 6mo ago · verified today
RadixArkai startup
Palo Alto, CA · Palo Alto Office · $200K–$400K/year est. · unknown
gpu-genericpost-traininggpu-kernelsmodel-parallelismtraining-frameworks
posted 7mo ago · verified today
falinference provider
Remote - USA · remote · $180K–$250K/year · senior
gpu-generickubernetes-opssoftware-engineercluster-datacentergpu-virtualization
posted 4w ago · verified today
falinference provider
Remote - Global · senior
ml-platformpython-langrust-langscheduling-orchestrationsoftware-engineer
posted 6mo ago · verified today
falinference provider
San Francisco · $180K–$250K/year · mid
ansiblecudanvidiaobservabilitypython-lang
posted 6mo ago · verified today
falinference provider
Remote - Global · unknown
python-langgpu-genericcudanvidiareliability-sre
posted 6mo ago · verified today
Inferactinference provider
San Francisco · onsite · $200K–$400K/year · unknown
cluster-datacentergpu-genericreliability-sreansiblekubernetes-ops
posted 3w ago · verified today
NexGen Cloudneocloud
UK - Remote · remote · senior
cudanvidiacluster-datacenternetwork-fabriccudnn-lib
posted 4mo ago · verified today
Prime Intellectneocloud
San Francisco · onsite · unknown
kubernetes-opsml-platformfine-tuninggpu-genericnvidia
posted 10w ago · verified today
Prime Intellectneocloud
Remote · San Francisco · hybrid · unknown
cluster-datacentergpu-genericinfiniband-opskubernetes-opsnetwork-fabric
posted 10w ago · verified today
Radiantneocloud
London · hybrid · senior
cluster-datacenterdatacenter-engineergpu-genericinfiniband-opsnetwork-fabric
posted 6w ago · verified today
Runpodneocloud
Remote - USA · remote · manager
eng-manageransiblecluster-datacentergpu-genericinfiniband-ops
posted 9w ago · verified today
TensorWaveneocloud
Las Vegas, Nevada · onsite · mid
amdgpu-generickubernetes-opsreliability-srescheduling-orchestration
posted 3w ago · verified today
Liquid AIfrontier lab
San Francisco · hybrid · unknown
cluster-datacentergpu-genericreliability-sresoftware-engineersre
posted 7w ago · verified today
Liquid AIfrontier lab
San Francisco · hybrid · unknown
deepspeed-libfsdppytorch-distsoftware-engineertraining-frameworks
posted 13mo ago · verified today
Nscaleneocloud
Houston · New York +2 more · $210K–$270K · staff plus
cluster-datacenterkubernetes-opsgpu-genericscheduling-orchestrationslurm-admin
posted 11d ago · verified today
Mistral AIfrontier lab
Amsterdam · Lausanne +3 more · hybrid · unknown
post-trainingreinforcement-learningresearch-engineerevaluationfine-tuning
posted 12d ago · verified today
Thinking Machines Labfrontier lab
New York · San Francisco · onsite · unknown
reliability-sresrefine-tuninggpu-generickubernetes-ops
posted 15d ago · verified today
Thinking Machines Labfrontier lab
New York · San Francisco · onsite · unknown
post-trainingreliability-sresrepython-langreinforcement-learning
posted 15d ago · verified today
Nscaleneocloud
Austin, TX · Houston +3 more · staff plus
ansiblecluster-datacentercollectivescpp-langdeepspeed-lib
posted 13d ago · verified today
Nscaleneocloud
UK · staff plus
object-storagesoftware-engineerstorage-checkpointingobservability
posted 5mo ago · verified today
Nscaleneocloud
Singapore · senior
gpu-genericsolutions-architectcluster-datacenterinfiniband-opsnetwork-fabric
posted 4mo ago · verified today
Nscaleneocloud
UK · senior
object-storagestorage-checkpointingobservabilitysoftware-engineerreliability-sre
posted 5mo ago · verified today
Nscaleneocloud
US · $150K–$300K · senior
object-storagesoftware-engineerobservabilitystorage-checkpointingcluster-datacenter
posted 3mo ago · verified today