AI infrastructure roles, filterable by the stack you actually work on.

Distributed training · inference serving · GPU fleets · network fabric — aggregated straight from company boards, never a copy of a copy.

1297 open roles · 85 companies · last verified today

180 roles

infiniband-ops

Lightning AIai startup

Lisle, Illinois, United States · ORD1 (Lisle) · $80K–$90K · unknown

cluster-datacenterdatacenter-engineerinfiniband-opsnetwork-fabricnvidia

posted 9w ago · verified today

Lightning AIai startup

Allen, Texas, United States · DFW1 (Allen) · $80K–$90K · mid

cluster-datacenterdatacenter-engineerinfiniband-opsnetwork-fabric

posted 9w ago · verified today

Lightning AIai startup

New York, New York, United States · San Francisco, California +3 more · $115K–$140K · unknown

cudagpu-generickubernetes-opsnccl-libobservability

posted 3mo ago · verified today

Lightning AIai startup

London, England, United Kingdom · London, UK · £75K–£95K · unknown

cudagpu-generickubernetes-opsml-platformnccl-lib

posted 3mo ago · verified today

Lightning AIai startup

Philippines · Remote +1 more · unknown

reliability-sresrecluster-datacentergpu-generickubernetes-ops

posted 4mo ago · verified today

Etchedchip vendor

San Jose · onsite · junior

cpp-langinferencepython-langdistributed-inferencecollectives

posted 9mo ago · verified today

RadixArkai startup

Palo Alto, CA · Palo Alto Office · $200K–$400K/year est. · senior

cpp-langcudagpu-kernelsperformance-engineercollectives

posted 7mo ago · verified today

RadixArkai startup

Palo Alto, CA · Palo Alto Office · $200K–$400K/year est. · senior

cluster-datacenterscheduling-orchestrationgpu-generickubernetes-opsnetwork-fabric

posted 7mo ago · verified today

falinference provider

Remote - USA · remote · $180K–$250K/year · senior

gpu-generickubernetes-opssoftware-engineercluster-datacentergpu-virtualization

posted 4w ago · verified today

falinference provider

San Francisco · $180K–$250K/year · mid

ansiblecudanvidiaobservabilitypython-lang

posted 6mo ago · verified today

falinference provider

Remote - Global · unknown

python-langgpu-genericcudanvidiareliability-sre

posted 6mo ago · verified today

Inferactinference provider

San Francisco · onsite · $200K–$400K/year · unknown

cluster-datacentergpu-genericreliability-sreansiblekubernetes-ops

posted 3w ago · verified today

Inferactinference provider

Singapore · onsite · 200K–400K SGD/year · unknown

software-engineerdistributed-inferenceinference-enginesinfiniband-opsnvlink-topology

posted 12w ago · verified today

Inferactinference provider

San Francisco · onsite · $200K–$400K/year · unknown

distributed-inferenceinference-enginesvllm-enginecpp-langgo-lang

posted 7mo ago · verified today

FriendliAIinference provider

San Francisco · hybrid · senior

cluster-datacentergpu-genericinferencekubernetes-opsml-platform

posted 5w ago · verified today

NexGen Cloudneocloud

London · London, England, United Kingdom, UK - Remote +1 more · remote · unknown

solutions-architectcudacudnn-libdeepspeed-libgpu-generic

posted 8w ago · verified today

NexGen Cloudneocloud

UK - Remote · remote · senior

cudanvidiacluster-datacenternetwork-fabriccudnn-lib

posted 4mo ago · verified today

San Francisco Compute Companyneocloud

Remote · San Francisco, CA · hybrid · $220K–$300K/year · senior

cluster-datacenterdatacenter-engineergpu-genericnetwork-fabricinfiniband-ops

posted 1d ago · verified today

Radiantneocloud

Gloucestershire · London · hybrid · senior

reliability-sresrecluster-datacentergpu-genericnvidia

posted 3mo ago · verified today

Radiantneocloud

Gloucestershire · hybrid · senior

ansibleobservabilitypython-langreliability-sresre

posted 5mo ago · verified today

Runpodneocloud

Remote - USA · remote · manager

eng-manageransiblecluster-datacentergpu-genericinfiniband-ops

posted 9w ago · verified today

Periodic Labsfrontier lab

Menlo Park, CA · onsite · unknown

collectivescudacutlass-cutefsdpgpu-generic

posted 4mo ago · verified today

Baseteninference provider

San Francisco · hybrid · $265K–$285K/year · manager

cluster-datacentergpu-genericinfiniband-opsnetwork-fabricroce-net

posted 11d ago · verified today

Nscaleneocloud

London · UK · staff plus

solutions-architectgpu-generickubernetes-opsslurm-adminterraform-iac

posted 12d ago · verified today

Thinking Machines Labfrontier lab

New York · San Francisco · onsite · unknown

post-trainingreliability-sresrepython-langreinforcement-learning

posted 15d ago · verified today

Nscaleneocloud

Seattle · US · $220K–$320K · staff plus

cluster-datacenterpython-langsoftware-engineergpu-genericobservability

posted 5w ago · verified today