AI infrastructure roles, filterable by the stack you actually work on.
Distributed training · inference serving · GPU fleets · network fabric — aggregated straight from company boards, never a copy of a copy.
1294 open roles · 85 companies · last verified today
1294 roles
TensorWaveneocloud
Las Vegas, Nevada · Remote · onsite · senior
gpu-generickubernetes-opsml-platformpython-langscheduling-orchestration
posted 10mo ago · verified today
TensorWaveneocloud
Miami, Florida · onsite · unknown
datacenter-engineercluster-datacentergpu-generic
posted 14d ago · verified today
TensorWaveneocloud
Tucson, Arizona · onsite · unknown
cluster-datacenterdatacenter-engineergpu-generic
posted 14d ago · verified today
Ineffable Intelligencefrontier lab
London · onsite · unknown
cluster-datacenterkubernetes-opsscheduling-orchestrationobservabilitypython-lang
posted 5w ago · verified today
Liquid AIfrontier lab
San Francisco · hybrid · unknown
cluster-datacentergpu-genericreliability-sresoftware-engineersre
posted 7w ago · verified today
Liquid AIfrontier lab
Boston · hybrid · unknown
cpp-langinferenceinference-enginespython-langevaluation
posted 3w ago · verified today
Liquid AIfrontier lab
San Francisco · hybrid · unknown
deepspeed-libfsdppytorch-distsoftware-engineertraining-frameworks
posted 13mo ago · verified today
Liquid AIfrontier lab
Boston · Remote +1 more · hybrid · unknown
cudagpu-genericgpu-kernelsperformance-engineercpp-lang
posted 13mo ago · verified today
Rekafrontier lab
US, UK, Singapore, Remote · remote · unknown
cpp-langcudafine-tuninggpu-genericgpu-kernels
posted 8mo ago · verified today
Poolsidefrontier lab
Remote (EMEA) · remote · unknown
gpu-genericinferencescheduling-orchestrationgo-langinference-engines
posted 9w ago · verified today
Periodic Labsfrontier lab
Menlo Park, CA · onsite · unknown
collectivescudacutlass-cutefsdpgpu-generic
posted 4mo ago · verified today
Amazon (AWS)hyperscaler
Tel Aviv-Yafo, Tel Aviv, ISR · onsite · junior
cpp-langdatacenter-engineersoftware-engineercluster-datacenterlinux-kernel
posted 11d ago · verified today
Together AIneocloud
Amsterdam · senior
go-langkubernetes-opsml-platformpython-langobservability
posted 10d ago · verified today
Baseteninference provider
San Francisco · hybrid · $225K–$235K/year · senior
cluster-datacentergpu-genericreliability-sreobservabilitynvidia
posted 12d ago · verified today
Baseteninference provider
San Francisco · hybrid · $265K–$285K/year · manager
cluster-datacenternvidiagpu-genericnetwork-fabricinfiniband-ops
posted 12d ago · verified today
Nscaleneocloud
Houston · New York +2 more · $210K–$270K · staff plus
cluster-datacenterkubernetes-opsgpu-genericscheduling-orchestrationslurm-admin
posted 12d ago · verified today
Nscaleneocloud
London · UK · staff plus
solutions-architectgpu-generickubernetes-opsslurm-adminterraform-iac
posted 12d ago · verified today
Mistral AIfrontier lab
Amsterdam · Lausanne +3 more · hybrid · unknown
post-trainingreinforcement-learningresearch-engineerevaluationfine-tuning
posted 12d ago · verified today
Anthropicfrontier lab
New York City, NY · Remote-Friendly (Travel-Required) +2 more · remote · $320K–$485K · staff plus
reliability-sresreml-platformrust-langinference
posted 12d ago · verified today
Thinking Machines Labfrontier lab
San Francisco · hybrid · unknown
reinforcement-learningresearch-engineerdistributed-inferenceinferenceinference-engines
posted 3w ago · verified today
Thinking Machines Labfrontier lab
San Francisco · hybrid · unknown
gpu-genericcluster-datacenterscheduling-orchestrationsoftware-engineerml-platform
posted 6w ago · verified today
Thinking Machines Labfrontier lab
San Francisco · hybrid · unknown
cudagpu-kernelstriton-langresearch-engineercutlass-cute
posted 6w ago · verified today
Thinking Machines Labfrontier lab
San Francisco · hybrid · unknown
research-engineertraining-frameworksdeepspeed-libgpu-genericmegatron-lm
posted 6w ago · verified today
Thinking Machines Labfrontier lab
New York · San Francisco · onsite · unknown
reliability-sresrefine-tuninggpu-generickubernetes-ops
posted 16d ago · verified today
Thinking Machines Labfrontier lab
San Francisco · hybrid · unknown
distributed-inferenceinference-enginesresearch-engineergpu-genericinference
posted 6w ago · verified today
Thinking Machines Labfrontier lab
San Francisco · hybrid · unknown
gpu-genericresearch-engineergpu-kernelsmodel-parallelismquantization
posted 6w ago · verified today
Thinking Machines Labfrontier lab
San Francisco · hybrid · unknown
cluster-datacenterkubernetes-opsml-platformpython-langrust-lang
posted 6w ago · verified today
Thinking Machines Labfrontier lab
San Francisco · hybrid · unknown
cluster-datacenterreliability-sresregpu-generickubernetes-ops
posted 12w ago · verified today
Thinking Machines Labfrontier lab
San Francisco · hybrid · unknown
gpu-genericnetwork-engineernetwork-fabriccollectives
posted 12w ago · verified today
Thinking Machines Labfrontier lab
New York · San Francisco · onsite · unknown
post-trainingreliability-sresrepython-langreinforcement-learning
posted 16d ago · verified today