AI infrastructure roles, filterable by the stack you actually work on.

Distributed training · inference serving · GPU fleets · network fabric — aggregated straight from company boards, never a copy of a copy.

1294 open roles · 85 companies · last verified today

407 roles

ml-platform

Periodic Labsfrontier lab

Menlo Park, CA · onsite · unknown

collectivescudacutlass-cutefsdpgpu-generic

posted 4mo ago · verified today

Mistral AIfrontier lab

Amsterdam · Lausanne +3 more · hybrid · unknown

post-trainingreinforcement-learningresearch-engineerevaluationfine-tuning

posted 12d ago · verified today

Anthropicfrontier lab

New York City, NY · Remote-Friendly (Travel-Required) +2 more · remote · $320K–$485K · staff plus

reliability-sresreml-platformrust-langinference

posted 12d ago · verified today

Thinking Machines Labfrontier lab

New York · San Francisco · onsite · unknown

reliability-sresrefine-tuninggpu-generickubernetes-ops

posted 16d ago · verified today

Amazon (AWS)hyperscaler

Palo Alto, California, USA · Seattle, Washington, USA · onsite · mid

ml-platformsoftware-engineerinference-enginesscheduling-orchestrationinference

posted 14d ago · verified today

Amazon (AWS)hyperscaler

Vancouver, British Columbia, CAN · onsite · senior

ml-platformsoftware-engineerevaluationinferencekubernetes-ops

posted 14d ago · verified today

Mistral AIfrontier lab

Palo Alto · hybrid · unknown

ml-platformscheduling-orchestrationkubernetes-opspython-langcuda

posted 13d ago · verified today

Amazon (AWS)hyperscaler

Cupertino, California, USA · Seattle, Washington, USA · onsite · manager

eng-managertrainiumml-platform

posted 15d ago · verified today

Nscaleneocloud

US · $190K–$260K · staff plus

observabilitygo-langgpu-generickubernetes-opspython-lang

posted 3mo ago · verified today

Nscaleneocloud

Houston · San Francisco +1 more · $220K–$265K · staff plus

kubernetes-opsgo-langobservabilitysoftware-engineerebpf

posted 20d ago · verified today

Nscaleneocloud

US · $160K–$230K · senior

observabilitygo-langkubernetes-opspython-langansible

posted 3w ago · verified today

Nscaleneocloud

US · $190K–$300K · staff plus

observabilitykubernetes-opsgo-langpython-langreliability-sre

posted 5mo ago · verified today

Amazon (AWS)hyperscaler

Arlington, Virginia, USA · Denver, Colorado, USA +1 more · onsite · senior

solutions-architectkubernetes-opsscheduling-orchestrationcluster-datacentergpu-generic

posted 16d ago · verified today

Anthropicfrontier lab

New York City, NY · San Francisco, CA +1 more · $405K–$625K · manager

eng-managerinferencescheduling-orchestrationinference-enginescluster-datacenter

posted 15d ago · verified today

Amazon (AWS)hyperscaler

Seattle, Washington, USA · onsite · manager

eng-managerml-platformperformance-engineercluster-datacentercuda

posted 17d ago · verified today

Crusoeneocloud

San Francisco, CA - US · Sunnyvale, CA - US · onsite · senior

go-langsoftware-engineerkubernetes-opsobservabilityterraform-iac

posted 16d ago · verified today

Amazon (AWS)hyperscaler

Seattle, Washington, USA · onsite · senior

ml-platformsoftware-engineerinferencejax-pallaspytorch-dist

posted 20d ago · verified today

Nebiusneocloud

Abu Dhabi, UAE · Middle East +1 more · remote · senior

solutions-architectgpu-generickubernetes-opsml-platforminference

posted 9w ago · verified today

Amazon (AWS)hyperscaler

Arlington, Virginia, USA · Denver, Colorado, USA +1 more · onsite · unknown

kubernetes-opsnvidiascheduling-orchestrationgpu-genericsolutions-architect

posted 3w ago · verified today

Nebiusneocloud

Canada · Remote - United States · CA$235K–CA$300K/year est. · unknown

solutions-architectgpu-generickubernetes-opsterraform-iaccuda

posted 3w ago · verified today