AI infrastructure roles, filterable by the stack you actually work on.
Distributed training · inference serving · GPU fleets · network fabric — aggregated straight from company boards, never a copy of a copy.
1297 open roles · 85 companies · last verified today
99 roles
Nscaleneocloud
Houston · New York +2 more · $160K–$290K · senior
infiniband-opsnetwork-fabricroce-netsolutions-architectcluster-datacenter
posted 1d ago · verified today
Amazon (AWS)hyperscaler
Austin, Texas, USA · New York, New York, USA +2 more · onsite · senior
solutions-architectkubernetes-opsmpinccl-libtraining-frameworks
posted 2d ago · verified today
SambaNovachip vendor
Tokyo, Japan · Tokyo Prefecture, Japan · 105K–130K JPY · senior
inferencepython-langsolutions-architectperformance-engineerevaluation
posted 2d ago · verified today
Amazon (AWS)hyperscaler
Austin, Texas, USA · Dallas, Texas, USA +2 more · onsite · staff plus
cluster-datacentercollectivesdeepspeed-libdistributed-inferenceefa-fabric
posted 7w ago · verified today
CoreWeaveneocloud
Bellevue, WA · San Francisco, CA +2 more · $182K–$242K/year est. · senior
kubernetes-opssolutions-architectnvidiainfiniband-opsnccl-lib
posted 5d ago · verified today
Mistral AIfrontier lab
Montréal · New York +1 more · remote · senior
solutions-architectgpu-genericscheduling-orchestrationcluster-datacenterinference
posted 6d ago · verified today
Amazon (AWS)hyperscaler
Austin, Texas, USA · Chicago, Illinois, USA +5 more · onsite · staff plus
solutions-architectcluster-datacentercollectivesefa-fabricfsdp
posted 7d ago · verified today
Reactorinference provider
San Francisco · onsite · mid
inferencepython-langsolutions-architectinference-enginesperformance-engineer
added 7d ago · verified today
DigitalOceanneocloud
Bay Area Metro · Seattle · $220K–$239K/year est. · staff plus
amdcudainferencekubernetes-opsnvidia
posted 4mo ago · verified today
DigitalOceanneocloud
Bay Area Metro · San Francisco · $220K–$239K/year est. · staff plus
amdcudadistributed-inferencego-langinference
posted 4mo ago · verified today
DigitalOceanneocloud
Bay Area Metro · New York · $150K–$186K/year est. · senior
solutions-architectcudagpu-genericinferencekubernetes-ops
posted 5mo ago · verified today
DigitalOceanneocloud
Bay Area Metro · San Francisco · $150K–$215K/year est. · senior
cudakubernetes-opssolutions-architectfine-tuninggpu-generic
posted 5mo ago · verified today
DigitalOceanneocloud
Bangalore Metro · Bengaluru · senior
cudakubernetes-opsnvidiaamdrocm-hip
posted 11w ago · verified today
DigitalOceanneocloud
Bangalore Metro · Bengaluru · senior
distributed-inferenceinferenceinference-engineskv-cache-systemsperformance-engineer
posted 6w ago · verified today
DeepInfrainference provider
Palo Alto, United States · onsite · unknown
inferenceinference-enginesnvidiapython-langsolutions-architect
posted 4w ago · verified today
Waferinference provider
San Francisco · onsite · $200K–$300K/year · unknown
gpu-kernelsinferenceinference-enginescluster-datacenterperformance-engineer
posted 8w ago · verified today
Novita AIinference provider
San Mateo · onsite · unknown
kubernetes-opspython-langsolutions-architectinferenceinference-engines
posted 11w ago · verified today
Crusoeneocloud
San Francisco, CA - US · onsite · staff plus
inferenceinference-enginesperformance-engineersglang-enginevllm-engine
posted 7d ago · verified today
Together AIneocloud
San Francisco · $270K–$300K/year est. · senior
distributed-inferencefine-tuninginferenceinference-engineskv-cache-systems
posted 4mo ago · verified today
Nebiusneocloud
Abu Dhabi, UAE · Dubai +1 more · senior
solutions-architectansiblecudagpu-generickubernetes-ops
posted 9d ago · verified today
Tenstorrentchip vendor
Austin, Texas, United States · Santa Clara +2 more · $100K–$500K/year est. · staff plus
inferencekubernetes-opsinference-enginesobservabilitysglang-engine
posted 5w ago · verified today
Lightning AIai startup
London, England, United Kingdom · New York, New York +5 more · $180K–$250K · unknown
python-langsolutions-architectdistributed-inferenceinferenceinference-engines
posted 3mo ago · verified today
FriendliAIinference provider
San Francisco · hybrid · mid
solutions-architectgpu-kernelsinference-engineskubernetes-opsobservability
posted 6mo ago · verified today
NexGen Cloudneocloud
London · London, England, United Kingdom, UK - Remote +1 more · remote · unknown
solutions-architectcudacudnn-libdeepspeed-libgpu-generic
posted 8w ago · verified today
NexGen Cloudneocloud
UK - Remote · remote · senior
cudanvidiacluster-datacenternetwork-fabriccudnn-lib
posted 4mo ago · verified today
Prime Intellectneocloud
Remote · San Francisco · hybrid · unknown
cluster-datacentergpu-genericinfiniband-opskubernetes-opsnetwork-fabric
posted 10w ago · verified today
Prime Intellectneocloud
New York City, USA · hybrid · unknown
evaluationpost-trainingdistributed-inferenceray-distributedreinforcement-learning
posted 10w ago · verified today
Radiantneocloud
London · hybrid · senior
cluster-datacenterdatacenter-engineergpu-genericinfiniband-opsnetwork-fabric
posted 6w ago · verified today
TensorWaveneocloud
Las Vegas, Nevada · Remote · remote · senior
ansiblecudakubernetes-opsnetwork-fabricpython-lang
posted 7w ago · verified today
Nscaleneocloud
London · UK · staff plus
solutions-architectgpu-generickubernetes-opsslurm-adminterraform-iac
posted 12d ago · verified today