AI infrastructure roles, filterable by the stack you actually work on.

Distributed training · inference serving · GPU fleets · network fabric — aggregated straight from company boards, never a copy of a copy.

1297 open roles · 85 companies · last verified today

116 roles

nccl-lib

Lightning AIai startup

New York, New York, United States · San Francisco, California +3 more · $115K–$140K · unknown

cudagpu-generickubernetes-opsnccl-libobservability

posted 3mo ago · verified today

Lightning AIai startup

London, England, United Kingdom · London, UK · £75K–£95K · unknown

cudagpu-generickubernetes-opsml-platformnccl-lib

posted 3mo ago · verified today

Lightning AIai startup

Philippines · Remote +1 more · unknown

reliability-sresrecluster-datacentergpu-generickubernetes-ops

posted 4mo ago · verified today

Coherefrontier lab

London · Montreal +4 more · remote · senior

training-frameworksdistributed-inferencepre-trainingcudakubernetes-ops

posted 9mo ago · verified today

Coherefrontier lab

Canada · United States · hybrid · senior

cluster-datacentergo-langgpu-generickubernetes-opsnccl-lib

posted 13d ago · verified today

Reflection AIfrontier lab

London · New York, NY +1 more · onsite · unknown

pre-trainingmodel-parallelismsoftware-engineertraining-frameworkscollectives

posted 5mo ago · verified today

Reflection AIfrontier lab

London · New York, NY +1 more · onsite · unknown

kubernetes-opscluster-datacenternccl-libreliability-sregpu-generic

posted 6mo ago · verified today

RadixArkai startup

Palo Alto, CA · Palo Alto Office · $200K–$400K/year est. · unknown

gpu-genericpost-traininggpu-kernelsmodel-parallelismtraining-frameworks

posted 7mo ago · verified today

RadixArkai startup

Palo Alto, CA · Palo Alto Office · $200K–$400K/year est. · unknown

amdnvidiagpu-kernelscollectivescpp-lang

posted 6w ago · verified today

RadixArkai startup

Palo Alto, CA · Palo Alto Office · $200K–$400K/year est. · senior

cpp-langcudagpu-kernelsperformance-engineercollectives

posted 7mo ago · verified today

falinference provider

Remote - USA · remote · $180K–$250K/year · senior

gpu-generickubernetes-opssoftware-engineercluster-datacentergpu-virtualization

posted 4w ago · verified today

Inferactinference provider

San Francisco · onsite · $200K–$400K/year · unknown

cluster-datacentergpu-genericreliability-sreansiblekubernetes-ops

posted 3w ago · verified today

FriendliAIinference provider

San Francisco · hybrid · senior

cluster-datacentergpu-genericinferencekubernetes-opsml-platform

posted 5w ago · verified today

NexGen Cloudneocloud

London · London, England, United Kingdom, UK - Remote +1 more · remote · unknown

solutions-architectcudacudnn-libdeepspeed-libgpu-generic

posted 8w ago · verified today

NexGen Cloudneocloud

UK - Remote · remote · senior

cudanvidiacluster-datacenternetwork-fabriccudnn-lib

posted 4mo ago · verified today

Radiantneocloud

Gloucestershire · London · hybrid · senior

reliability-sresrecluster-datacentergpu-genericnvidia

posted 3mo ago · verified today

Periodic Labsfrontier lab

Menlo Park, CA · onsite · unknown

collectivescudacutlass-cutefsdpgpu-generic

posted 4mo ago · verified today

Thinking Machines Labfrontier lab

New York · San Francisco · onsite · unknown

post-trainingreliability-sresrepython-langreinforcement-learning

posted 15d ago · verified today

Mistral AIfrontier lab

Palo Alto · hybrid · unknown

ml-platformscheduling-orchestrationkubernetes-opspython-langcuda

posted 13d ago · verified today

Nscaleneocloud

Houston · New York +2 more · staff plus

nvidiasoftware-engineercluster-datacentercudago-lang

posted 4w ago · verified today

Nscaleneocloud

Houston · San Francisco +1 more · $120K–$170K · senior

nvidiasrecluster-datacenterinfiniband-opsnccl-lib

posted 6mo ago · verified today

Nscaleneocloud

Houston · San Francisco +1 more · $100K–$140K · unknown

gpu-genericreliability-sresrecluster-datacenternvidia

posted 3mo ago · verified today

Amazon (AWS)hyperscaler

Boston, Massachusetts, USA · Seattle, Washington, USA +1 more · onsite · senior

gpu-kernelsinferenceinference-engineskv-cache-systemsperformance-engineer

posted 20d ago · verified today

Nebiusneocloud

Remote - United States · United States · remote · $180K–$220K/year est. · senior

gpu-genericnvidiasolutions-architectcluster-datacenterperformance-engineer

posted 5mo ago · verified today