AI infrastructure roles, filterable by the stack you actually work on.
Distributed training · inference serving · GPU fleets · network fabric — aggregated straight from company boards, never a copy of a copy.
1297 open roles · 85 companies · last verified today
619 roles
Poolsidefrontier lab
Remote (EMEA) · remote · unknown
gpu-genericinferencescheduling-orchestrationgo-langinference-engines
posted 9w ago · verified today
Periodic Labsfrontier lab
Menlo Park, CA · onsite · unknown
collectivescudacutlass-cutefsdpgpu-generic
posted 4mo ago · verified today
Baseteninference provider
San Francisco · hybrid · $225K–$235K/year · senior
cluster-datacentergpu-genericreliability-sreobservabilitynvidia
posted 12d ago · verified today
Baseteninference provider
San Francisco · hybrid · $265K–$285K/year · manager
cluster-datacenternvidiagpu-genericnetwork-fabricinfiniband-ops
posted 12d ago · verified today
Nscaleneocloud
Houston · New York +2 more · $210K–$270K · staff plus
cluster-datacenterkubernetes-opsgpu-genericscheduling-orchestrationslurm-admin
posted 12d ago · verified today
Nscaleneocloud
London · UK · staff plus
solutions-architectgpu-generickubernetes-opsslurm-adminterraform-iac
posted 12d ago · verified today
Thinking Machines Labfrontier lab
San Francisco · hybrid · unknown
gpu-genericcluster-datacenterscheduling-orchestrationsoftware-engineerml-platform
posted 6w ago · verified today
Thinking Machines Labfrontier lab
San Francisco · hybrid · unknown
cudagpu-kernelstriton-langresearch-engineercutlass-cute
posted 6w ago · verified today
Thinking Machines Labfrontier lab
San Francisco · hybrid · unknown
research-engineertraining-frameworksdeepspeed-libgpu-genericmegatron-lm
posted 6w ago · verified today
Thinking Machines Labfrontier lab
New York · San Francisco · onsite · unknown
reliability-sresrefine-tuninggpu-generickubernetes-ops
posted 16d ago · verified today
Thinking Machines Labfrontier lab
San Francisco · hybrid · unknown
distributed-inferenceinference-enginesresearch-engineergpu-genericinference
posted 6w ago · verified today
Thinking Machines Labfrontier lab
San Francisco · hybrid · unknown
gpu-genericresearch-engineergpu-kernelsmodel-parallelismquantization
posted 6w ago · verified today
Thinking Machines Labfrontier lab
San Francisco · hybrid · unknown
cluster-datacenterkubernetes-opsml-platformpython-langrust-lang
posted 6w ago · verified today
Thinking Machines Labfrontier lab
San Francisco · hybrid · unknown
cluster-datacenterreliability-sresregpu-generickubernetes-ops
posted 12w ago · verified today
Thinking Machines Labfrontier lab
San Francisco · hybrid · unknown
gpu-genericnetwork-engineernetwork-fabriccollectives
posted 12w ago · verified today
Thinking Machines Labfrontier lab
New York · San Francisco · onsite · unknown
post-trainingreliability-sresrepython-langreinforcement-learning
posted 16d ago · verified today
Anyscaleai startup
San Francisco · $215K–$265K/year · senior
cpp-langray-distributedsoftware-engineerscheduling-orchestrationgpu-generic
posted 13d ago · verified today
Amazon (AWS)hyperscaler
Palo Alto, California, USA · Seattle, Washington, USA · onsite · mid
ml-platformsoftware-engineerinference-enginesscheduling-orchestrationinference
posted 14d ago · verified today
Amazon (AWS)hyperscaler
Seattle, Washington, USA · onsite · mid
software-engineerml-platformscheduling-orchestrationgpu-genericinference
posted 14d ago · verified today
Nscaleneocloud
Austin, TX · Houston +3 more · staff plus
ansiblecluster-datacentercollectivescpp-langdeepspeed-lib
posted 13d ago · verified today
Cerebraschip vendor
Remote (US) · Sunnyvale, CA · senior
datacenter-engineercluster-datacentergpu-generictpu
posted 13d ago · verified today
Mistral AIfrontier lab
Palo Alto · hybrid · unknown
ml-platformscheduling-orchestrationkubernetes-opspython-langcuda
posted 13d ago · verified today
xAIfrontier lab
Memphis, TN · Southaven, MS · mid
cluster-datacenterdatacenter-engineerpython-langreliability-sregpu-generic
posted 13d ago · verified today
Nebiusneocloud
Béthune, Pas-de-Calais, France · Hauts-de-France · unknown
cluster-datacentergpu-genericdatacenter-engineer
posted 13d ago · verified today
CoreWeaveneocloud
Bellevue, WA · Livingston, NJ +1 more · $153K–$204K/year est. · senior
gpu-generickubernetes-opssoftware-engineercluster-datacentergo-lang
posted 14d ago · verified today
xAIfrontier lab
Memphis, TN · Southaven, MS · senior
reliability-sresrecluster-datacenterobservabilitypython-lang
posted 14d ago · verified today
Anthropicfrontier lab
London, UK · £325K–£390K · staff plus
ebpfobservabilitygpu-genericsoftware-engineertpu
posted 14d ago · verified today
Nscaleneocloud
Seattle · US · $220K–$320K · staff plus
cluster-datacenterpython-langsoftware-engineergpu-genericobservability
posted 5w ago · verified today
Nscaleneocloud
US · $190K–$260K · staff plus
observabilitygo-langgpu-generickubernetes-opspython-lang
posted 3mo ago · verified today
Nscaleneocloud
New York · $225K–$275K · staff plus
scheduling-orchestrationslurm-adminsoftware-engineergo-langpython-lang
posted 3mo ago · verified today