AI infrastructure roles, filterable by the stack you actually work on.

Distributed training · inference serving · GPU fleets · network fabric — aggregated straight from company boards, never a copy of a copy.

1297 open roles · 85 companies · last verified today

197 roles

cuda

Periodic Labsfrontier lab

Menlo Park, CA · onsite · unknown

collectivescudacutlass-cutefsdpgpu-generic

posted 4mo ago · verified today

Mistral AIfrontier lab

Palo Alto · hybrid · unknown

ml-platformscheduling-orchestrationkubernetes-opspython-langcuda

posted 13d ago · verified today

Nscaleneocloud

Houston · New York +2 more · staff plus

nvidiasoftware-engineercluster-datacentercudago-lang

posted 4w ago · verified today

Nscaleneocloud

London · UK · senior

evaluationfine-tuninggpu-genericinferenceinference-engines

posted 5mo ago · verified today

Amazon (AWS)hyperscaler

Cupertino, California, USA · onsite · senior

gpu-kernelsperformance-engineertrainiumcudanvidia

posted 16d ago · verified today

Lambdaneocloud

Bellevue Office · San Francisco Office (Fremont St) +1 more · remote · $231K–$342K · senior

gpu-genericsolutions-architectcluster-datacenterinfiniband-opskubernetes-ops

posted 15d ago · verified today

Amazon (AWS)hyperscaler

Seattle, Washington, USA · onsite · manager

eng-managerml-platformperformance-engineercluster-datacentercuda

posted 17d ago · verified today

Amazon (AWS)hyperscaler

Boston, Massachusetts, USA · Seattle, Washington, USA +1 more · onsite · senior

gpu-kernelsinferenceinference-engineskv-cache-systemsperformance-engineer

posted 3w ago · verified today

Nebiusneocloud

Canada · Remote - United States · CA$235K–CA$300K/year est. · unknown

solutions-architectgpu-generickubernetes-opsterraform-iaccuda

posted 3w ago · verified today

Crusoeneocloud

Denver, CO - US · onsite · staff plus

inferenceinference-enginesperformance-engineersglang-enginevllm-engine

posted 3w ago · verified today

Amazon (AWS)hyperscaler

Seattle, Washington, USA · onsite · manager

cudaeng-managerml-platformperformance-engineercluster-datacenter

posted 3mo ago · verified today

Amazon (AWS)hyperscaler

San Francisco, California, USA · onsite · senior

cudagpu-kernelsnvidiapython-langresearch-engineer

posted 3mo ago · verified today

Amazon (AWS)hyperscaler

Cupertino, California, USA · Seattle, Washington, USA · onsite · manager

collectivescpp-langcudagpu-kernelsnccl-lib

posted 6w ago · verified today

Amazon (AWS)hyperscaler

Cupertino, California, USA · onsite · manager

eng-managertrainiumdeepspeed-libfsdpmodel-parallelism

posted 3mo ago · verified today

Amazon (AWS)hyperscaler

Tel Aviv-Yafo, Tel Aviv, ISR · onsite · senior

inferenceinference-enginesvllm-enginecudagpu-generic

posted 10w ago · verified today

Amazon (AWS)hyperscaler

Tel Aviv-Yafo, Tel Aviv, ISR · onsite · mid

cpp-langsoftware-engineercustom-asicdistributed-inferencegpu-kernels

posted 5w ago · verified today