AI infrastructure roles, filterable by the stack you actually work on.
Distributed training · inference serving · GPU fleets · network fabric — aggregated straight from company boards, never a copy of a copy.
1297 open roles · 85 companies · last verified today
24 roles
Amazon (AWS)hyperscaler
Austin, Texas, USA · New York, New York, USA +2 more · onsite · senior
solutions-architectkubernetes-opsmpinccl-libtraining-frameworks
posted 2d ago · verified today
Luma AIai startup
Redwood City, CA · hybrid · unknown
pytorch-disttraining-frameworkscollectivescudafsdp
posted 7w ago · verified today
Luma AIai startup
Redwood City, CA · hybrid · unknown
fsdpgpu-genericinference-enginespost-trainingpytorch-dist
posted 7w ago · verified today
Tenstorrentchip vendor
Austin · Austin, Texas, United States +4 more · $100K–$500K/year est. · unknown
software-engineercpp-langmpicollectivescluster-datacenter
posted 17mo ago · verified today
Lightning AIai startup
New York, New York · New York, New York, United States +5 more · remote · $170K–$210K · senior
infiniband-opsnetwork-engineernetwork-fabricnvidiaansible
posted 8w ago · verified today
RadixArkai startup
Palo Alto, CA · Palo Alto Office · $200K–$400K/year est. · unknown
amdnvidiagpu-kernelscollectivescpp-lang
posted 6w ago · verified today
RadixArkai startup
Palo Alto, CA · Palo Alto Office · $200K–$400K/year est. · senior
cpp-langcudagpu-kernelsperformance-engineercollectives
posted 7mo ago · verified today
NexGen Cloudneocloud
London · London, England, United Kingdom, UK - Remote +1 more · remote · unknown
solutions-architectcudacudnn-libdeepspeed-libgpu-generic
posted 8w ago · verified today
NexGen Cloudneocloud
UK - Remote · remote · senior
cudanvidiacluster-datacenternetwork-fabriccudnn-lib
posted 4mo ago · verified today
Nscaleneocloud
Houston · San Francisco +1 more · $120K–$170K · senior
nvidiasrecluster-datacenterinfiniband-opsnccl-lib
posted 6mo ago · verified today
Amazon (AWS)hyperscaler
Cupertino, California, USA · onsite · mid
cpp-langdistributed-inferencenetwork-fabricsoftware-engineercollectives
posted 7w ago · verified today
Amazon (AWS)hyperscaler
Austin, Texas, USA · Cupertino, California, USA · onsite · mid
cpp-langsoftware-engineercollectivespython-langrust-lang
posted 5mo ago · verified today
Amazon (AWS)hyperscaler
Cupertino, California, USA · Seattle, Washington, USA · onsite · senior
cpp-langdistributed-inferencenetwork-fabricsoftware-engineercollectives
posted 8w ago · verified today
Amazon (AWS)hyperscaler
Seattle, Washington, USA · onsite · junior
network-fabricefa-fabricmpinetwork-engineersoftware-engineer
posted 8w ago · verified today
Amazon (AWS)hyperscaler
Cupertino, California, USA · onsite · junior
collectivescpp-langnccl-libnetwork-fabricsoftware-engineer
posted 6w ago · verified today
SambaNovachip vendor
San Jose, CA · San Jose, California, United States · staff plus
custom-asicgpu-kernelsnetwork-fabricrdma-verbsroce-net
posted 5mo ago · verified today
CoreWeaveneocloud
Bellevue, WA · Manhattan, NY +2 more · $165K–$242K/year est. · senior
performance-engineerinfiniband-opslinux-kernelmpinetwork-fabric
posted 12mo ago · verified today
OpenAIfrontier lab
San Francisco · hybrid · $295K–$500K/year · unknown
performance-engineercollectivescpp-langcudagpu-generic
posted 11mo ago · verified today
OpenAIfrontier lab
San Francisco · $266K–$500K/year · unknown
cudainferenceinference-enginesnvidiasoftware-engineer
posted 19mo ago · verified today
Nebiusneocloud
Remote - Europe · Remote - United States · remote · $170K–$300K/year est. · senior
cluster-datacentercpp-langgpu-genericinfiniband-opsnetwork-fabric
posted 5mo ago · verified today
Nebiusneocloud
Amsterdam, Netherlands · Remote - Europe +1 more · remote · senior
performance-engineercollectivescpp-langgo-langgpu-generic
posted 4mo ago · verified today
Nebiusneocloud
Amsterdam, Netherlands · Berlin, Germany +2 more · remote · senior
cluster-datacentercpp-langgpu-genericinfiniband-opsnetwork-fabric
posted 25mo ago · verified today
Nebiusneocloud
Remote - United States · United States · remote · $170K–$300K · manager
gpu-genericperformance-engineercluster-datacentercpp-langinfiniband-ops
posted 4mo ago · verified today
Together AIneocloud
San Francisco · $200K–$290K/year est. · senior
inferenceinference-enginesnvidiasoftware-engineerdistributed-inference
posted 13mo ago · verified today