AI infrastructure roles, filterable by the stack you actually work on.
Distributed training · inference serving · GPU fleets · network fabric — aggregated straight from company boards, never a copy of a copy.
1296 open roles · 85 companies · last verified today
130 roles
Reflection AIfrontier lab
London · New York, NY +1 more · onsite · unknown
pytorch-distreinforcement-learningsoftware-engineertraining-frameworkscollectives
posted 6mo ago · verified today
RadixArkai startup
Palo Alto, CA · Palo Alto Office · $200K–$400K/year est. · senior
cluster-datacenterscheduling-orchestrationgpu-generickubernetes-opsnetwork-fabric
posted 7mo ago · verified today
falinference provider
Remote - USA · remote · $180K–$250K/year · senior
gpu-generickubernetes-opssoftware-engineercluster-datacentergpu-virtualization
posted 4w ago · verified today
Inferactinference provider
Remote · remote · $200K–$400K/year · unknown
go-langkubernetes-opspython-langrust-langterraform-iac
posted 3w ago · verified today
Inferactinference provider
San Francisco · onsite · $200K–$400K/year · unknown
cluster-datacentergpu-genericreliability-sreansiblekubernetes-ops
posted 3w ago · verified today
Inferactinference provider
Singapore · onsite · 200K–400K SGD/year · unknown
cluster-datacentergo-langgpu-genericinferencekubernetes-ops
posted 12w ago · verified today
Inferactinference provider
San Francisco · onsite · $200K–$400K/year · unknown
kubernetes-opssoftware-engineercluster-datacentergo-langinference
posted 7mo ago · verified today
NexGen Cloudneocloud
London · London, England, United Kingdom, UK - Remote +1 more · remote · unknown
solutions-architectcudacudnn-libdeepspeed-libgpu-generic
posted 8w ago · verified today
NexGen Cloudneocloud
UK - Remote · remote · senior
cudanvidiacluster-datacenternetwork-fabriccudnn-lib
posted 4mo ago · verified today
San Francisco Compute Companyneocloud
Remote · San Francisco, CA · hybrid · $220K–$300K/year · senior
cluster-datacenterdatacenter-engineergpu-genericnetwork-fabricinfiniband-ops
posted 2d ago · verified today
Prime Intellectneocloud
Remote · San Francisco · hybrid · unknown
cluster-datacentergpu-genericinfiniband-opskubernetes-opsnetwork-fabric
posted 10w ago · verified today
Radiantneocloud
London · hybrid · senior
cluster-datacenterdatacenter-engineergpu-genericinfiniband-opsnetwork-fabric
posted 6w ago · verified today
TensorWaveneocloud
Las Vegas, Nevada · onsite · mid
amdgpu-generickubernetes-opsreliability-srescheduling-orchestration
posted 3w ago · verified today
TensorWaveneocloud
Las Vegas, Nevada · Remote · onsite · senior
gpu-generickubernetes-opsml-platformpython-langscheduling-orchestration
posted 10mo ago · verified today
Liquid AIfrontier lab
San Francisco · hybrid · unknown
cluster-datacentergpu-genericreliability-sresoftware-engineersre
posted 7w ago · verified today
Rekafrontier lab
US, UK, Singapore, Remote · remote · unknown
cpp-langcudafine-tuninggpu-genericgpu-kernels
posted 8mo ago · verified today
Nscaleneocloud
Houston · New York +2 more · $210K–$270K · staff plus
cluster-datacenterkubernetes-opsgpu-genericscheduling-orchestrationslurm-admin
posted 11d ago · verified today
Nscaleneocloud
London · UK · staff plus
solutions-architectgpu-generickubernetes-opsslurm-adminterraform-iac
posted 12d ago · verified today
Mistral AIfrontier lab
Amsterdam · Lausanne +3 more · hybrid · unknown
post-trainingreinforcement-learningresearch-engineerevaluationfine-tuning
posted 12d ago · verified today
Thinking Machines Labfrontier lab
San Francisco · hybrid · unknown
gpu-genericcluster-datacenterscheduling-orchestrationsoftware-engineerml-platform
posted 6w ago · verified today
Thinking Machines Labfrontier lab
San Francisco · hybrid · unknown
distributed-inferenceinference-enginesresearch-engineergpu-genericinference
posted 6w ago · verified today
Thinking Machines Labfrontier lab
San Francisco · hybrid · unknown
cluster-datacenterkubernetes-opsml-platformpython-langrust-lang
posted 6w ago · verified today
Thinking Machines Labfrontier lab
San Francisco · hybrid · unknown
cluster-datacenterreliability-sresregpu-generickubernetes-ops
posted 12w ago · verified today
Thinking Machines Labfrontier lab
New York · San Francisco · onsite · unknown
post-trainingreliability-sresrepython-langreinforcement-learning
posted 15d ago · verified today
Nscaleneocloud
Austin, TX · Houston +3 more · staff plus
ansiblecluster-datacentercollectivescpp-langdeepspeed-lib
posted 13d ago · verified today
Nscaleneocloud
US · $190K–$260K · staff plus
observabilitygo-langgpu-generickubernetes-opspython-lang
posted 3mo ago · verified today
Nscaleneocloud
Houston · New York +2 more · staff plus
nvidiasoftware-engineercluster-datacentercudago-lang
posted 4w ago · verified today
Nscaleneocloud
New York · $225K–$275K · staff plus
scheduling-orchestrationslurm-adminsoftware-engineergo-langpython-lang
posted 3mo ago · verified today
Nscaleneocloud
London · UK · staff plus
cluster-datacenterscheduling-orchestrationslurm-adminsoftware-engineergo-lang
posted 14d ago · verified today
Nscaleneocloud
Houston · New York +2 more · $140K–$193K · staff plus
solutions-architectkubernetes-opsgpu-genericslurm-admininfiniband-ops
posted 19d ago · verified today