AI infrastructure roles, filterable by the stack you actually work on.
Distributed training · inference serving · GPU fleets · network fabric — aggregated straight from company boards, never a copy of a copy.
1297 open roles · 85 companies · last verified today
180 roles
Lightning AIai startup
Lisle, Illinois, United States · ORD1 (Lisle) · $80K–$90K · unknown
cluster-datacenterdatacenter-engineerinfiniband-opsnetwork-fabricnvidia
posted 9w ago · verified today
Lightning AIai startup
Allen, Texas, United States · DFW1 (Allen) · $80K–$90K · mid
cluster-datacenterdatacenter-engineerinfiniband-opsnetwork-fabric
posted 9w ago · verified today
Lightning AIai startup
New York, New York, United States · San Francisco, California +3 more · $115K–$140K · unknown
cudagpu-generickubernetes-opsnccl-libobservability
posted 3mo ago · verified today
Lightning AIai startup
London, England, United Kingdom · London, UK · £75K–£95K · unknown
cudagpu-generickubernetes-opsml-platformnccl-lib
posted 3mo ago · verified today
Lightning AIai startup
Philippines · Remote +1 more · unknown
reliability-sresrecluster-datacentergpu-generickubernetes-ops
posted 4mo ago · verified today
Etchedchip vendor
San Jose · onsite · junior
cpp-langinferencepython-langdistributed-inferencecollectives
posted 9mo ago · verified today
RadixArkai startup
Palo Alto, CA · Palo Alto Office · $200K–$400K/year est. · senior
cpp-langcudagpu-kernelsperformance-engineercollectives
posted 7mo ago · verified today
RadixArkai startup
Palo Alto, CA · Palo Alto Office · $200K–$400K/year est. · senior
cluster-datacenterscheduling-orchestrationgpu-generickubernetes-opsnetwork-fabric
posted 7mo ago · verified today
falinference provider
Remote - USA · remote · $180K–$250K/year · senior
gpu-generickubernetes-opssoftware-engineercluster-datacentergpu-virtualization
posted 4w ago · verified today
falinference provider
San Francisco · $180K–$250K/year · mid
ansiblecudanvidiaobservabilitypython-lang
posted 6mo ago · verified today
falinference provider
Remote - Global · unknown
python-langgpu-genericcudanvidiareliability-sre
posted 6mo ago · verified today
Inferactinference provider
San Francisco · onsite · $200K–$400K/year · unknown
cluster-datacentergpu-genericreliability-sreansiblekubernetes-ops
posted 3w ago · verified today
Inferactinference provider
Singapore · onsite · 200K–400K SGD/year · unknown
software-engineerdistributed-inferenceinference-enginesinfiniband-opsnvlink-topology
posted 12w ago · verified today
Inferactinference provider
San Francisco · onsite · $200K–$400K/year · unknown
distributed-inferenceinference-enginesvllm-enginecpp-langgo-lang
posted 7mo ago · verified today
FriendliAIinference provider
San Francisco · hybrid · senior
cluster-datacentergpu-genericinferencekubernetes-opsml-platform
posted 5w ago · verified today
FriendliAIinference provider
Seoul · onsite · senior
kubernetes-opssoftware-engineernetwork-fabricscheduling-orchestrationcluster-datacenter
posted 5w ago · verified today
NexGen Cloudneocloud
London · London, England, United Kingdom, UK - Remote +1 more · remote · unknown
solutions-architectcudacudnn-libdeepspeed-libgpu-generic
posted 8w ago · verified today
NexGen Cloudneocloud
UK - Remote · remote · senior
cudanvidiacluster-datacenternetwork-fabriccudnn-lib
posted 4mo ago · verified today
San Francisco Compute Companyneocloud
Remote · San Francisco, CA · hybrid · $220K–$300K/year · senior
cluster-datacenterdatacenter-engineergpu-genericnetwork-fabricinfiniband-ops
posted 1d ago · verified today
Prime Intellectneocloud
Remote · San Francisco · hybrid · unknown
cluster-datacentergpu-genericinfiniband-opskubernetes-opsnetwork-fabric
posted 10w ago · verified today
Radiantneocloud
London · hybrid · senior
cluster-datacenterdatacenter-engineergpu-genericinfiniband-opsnetwork-fabric
posted 6w ago · verified today
Radiantneocloud
Gloucestershire · London · hybrid · senior
reliability-sresrecluster-datacentergpu-genericnvidia
posted 3mo ago · verified today
Radiantneocloud
Gloucestershire · hybrid · senior
ansibleobservabilitypython-langreliability-sresre
posted 5mo ago · verified today
Runpodneocloud
Remote - USA · remote · manager
eng-manageransiblecluster-datacentergpu-genericinfiniband-ops
posted 9w ago · verified today
Periodic Labsfrontier lab
Menlo Park, CA · onsite · unknown
collectivescudacutlass-cutefsdpgpu-generic
posted 4mo ago · verified today
Baseteninference provider
San Francisco · hybrid · $265K–$285K/year · manager
cluster-datacentergpu-genericinfiniband-opsnetwork-fabricroce-net
posted 11d ago · verified today
Nscaleneocloud
London · UK · staff plus
solutions-architectgpu-generickubernetes-opsslurm-adminterraform-iac
posted 12d ago · verified today
Thinking Machines Labfrontier lab
New York · San Francisco · onsite · unknown
post-trainingreliability-sresrepython-langreinforcement-learning
posted 15d ago · verified today
Nscaleneocloud
Austin, TX · Houston +3 more · staff plus
ansiblecluster-datacentercollectivescpp-langdeepspeed-lib
posted 13d ago · verified today
Nscaleneocloud
Seattle · US · $220K–$320K · staff plus
cluster-datacenterpython-langsoftware-engineergpu-genericobservability
posted 5w ago · verified today