AI infrastructure roles, filterable by the stack you actually work on.
Distributed training · inference serving · GPU fleets · network fabric — aggregated straight from company boards, never a copy of a copy.
1294 open roles · 85 companies · last verified today
407 roles
Periodic Labsfrontier lab
Menlo Park, CA · onsite · unknown
collectivescudacutlass-cutefsdpgpu-generic
posted 4mo ago · verified today
Together AIneocloud
Amsterdam · senior
go-langkubernetes-opsml-platformpython-langobservability
posted 10d ago · verified today
Mistral AIfrontier lab
Amsterdam · Lausanne +3 more · hybrid · unknown
post-trainingreinforcement-learningresearch-engineerevaluationfine-tuning
posted 12d ago · verified today
Anthropicfrontier lab
New York City, NY · Remote-Friendly (Travel-Required) +2 more · remote · $320K–$485K · staff plus
reliability-sresreml-platformrust-langinference
posted 12d ago · verified today
Thinking Machines Labfrontier lab
San Francisco · hybrid · unknown
gpu-genericcluster-datacenterscheduling-orchestrationsoftware-engineerml-platform
posted 6w ago · verified today
Thinking Machines Labfrontier lab
New York · San Francisco · onsite · unknown
reliability-sresrefine-tuninggpu-generickubernetes-ops
posted 16d ago · verified today
Thinking Machines Labfrontier lab
San Francisco · hybrid · unknown
distributed-inferenceinference-enginesresearch-engineergpu-genericinference
posted 6w ago · verified today
Thinking Machines Labfrontier lab
San Francisco · hybrid · unknown
gpu-genericresearch-engineergpu-kernelsmodel-parallelismquantization
posted 6w ago · verified today
Thinking Machines Labfrontier lab
San Francisco · hybrid · unknown
cluster-datacenterkubernetes-opsml-platformpython-langrust-lang
posted 6w ago · verified today
Amazon (AWS)hyperscaler
Palo Alto, California, USA · Seattle, Washington, USA · onsite · mid
ml-platformsoftware-engineerinference-enginesscheduling-orchestrationinference
posted 14d ago · verified today
Amazon (AWS)hyperscaler
Seattle, Washington, USA · onsite · mid
software-engineerml-platformscheduling-orchestrationgpu-genericinference
posted 14d ago · verified today
Amazon (AWS)hyperscaler
Vancouver, British Columbia, CAN · onsite · senior
ml-platformsoftware-engineerevaluationinferencekubernetes-ops
posted 14d ago · verified today
Amazon (AWS)hyperscaler
New York, New York, USA · onsite · mid
software-engineertrainiumml-platformcpp-langevaluation
posted 14d ago · verified today
Amazon (AWS)hyperscaler
New York, New York, USA · onsite · unknown
software-engineertrainiumevaluationml-platformcpp-lang
posted 14d ago · verified today
Amazon (AWS)hyperscaler
Cupertino, California, USA · onsite · mid
trainiumcpp-langsoftware-engineermlir-llvmxla-compiler
posted 14d ago · verified today
Mistral AIfrontier lab
Palo Alto · hybrid · unknown
ml-platformscheduling-orchestrationkubernetes-opspython-langcuda
posted 13d ago · verified today
Amazon (AWS)hyperscaler
Cupertino, California, USA · Seattle, Washington, USA · onsite · manager
eng-managertrainiumml-platform
posted 15d ago · verified today
Nscaleneocloud
US · $190K–$260K · staff plus
observabilitygo-langgpu-generickubernetes-opspython-lang
posted 3mo ago · verified today
Nscaleneocloud
Houston · San Francisco +1 more · $220K–$265K · staff plus
kubernetes-opsgo-langobservabilitysoftware-engineerebpf
posted 20d ago · verified today
Nscaleneocloud
US · $160K–$230K · senior
observabilitygo-langkubernetes-opspython-langansible
posted 3w ago · verified today
Nscaleneocloud
US · $190K–$300K · staff plus
observabilitykubernetes-opsgo-langpython-langreliability-sre
posted 5mo ago · verified today
Amazon (AWS)hyperscaler
Arlington, Virginia, USA · Denver, Colorado, USA +1 more · onsite · senior
solutions-architectkubernetes-opsscheduling-orchestrationcluster-datacentergpu-generic
posted 16d ago · verified today
Anthropicfrontier lab
New York City, NY · San Francisco, CA +1 more · $405K–$625K · manager
eng-managerinferencescheduling-orchestrationinference-enginescluster-datacenter
posted 15d ago · verified today
Amazon (AWS)hyperscaler
Seattle, Washington, USA · onsite · manager
eng-managerml-platformperformance-engineercluster-datacentercuda
posted 17d ago · verified today
Crusoeneocloud
San Francisco, CA - US · Sunnyvale, CA - US · onsite · senior
go-langsoftware-engineerkubernetes-opsobservabilityterraform-iac
posted 16d ago · verified today
Amazon (AWS)hyperscaler
Seattle, Washington, USA · onsite · senior
ml-platformsoftware-engineerinferencejax-pallaspytorch-dist
posted 20d ago · verified today
Amazon (AWS)hyperscaler
Cupertino, California, USA · onsite · senior
training-frameworksdistributed-inferencecluster-datacentercollectivesdatacenter-engineer
posted 3w ago · verified today
Nebiusneocloud
Abu Dhabi, UAE · Middle East +1 more · remote · senior
solutions-architectgpu-generickubernetes-opsml-platforminference
posted 9w ago · verified today
Amazon (AWS)hyperscaler
Arlington, Virginia, USA · Denver, Colorado, USA +1 more · onsite · unknown
kubernetes-opsnvidiascheduling-orchestrationgpu-genericsolutions-architect
posted 3w ago · verified today
Nebiusneocloud
Canada · Remote - United States · CA$235K–CA$300K/year est. · unknown
solutions-architectgpu-generickubernetes-opsterraform-iaccuda
posted 3w ago · verified today