AI infrastructure roles, filterable by the stack you actually work on.
Distributed training · inference serving · GPU fleets · network fabric — aggregated straight from company boards, never a copy of a copy.
1297 open roles · 85 companies · last verified today
94 roles
Nscaleneocloud
Houston · New York +2 more · $220K–$293K · staff plus
inferenceinference-engineskv-cache-systemspost-trainingpython-lang
posted 1d ago · verified today
SambaNovachip vendor
Tokyo, Japan · Tokyo Prefecture, Japan · 105K–130K JPY · senior
inferencepython-langsolutions-architectperformance-engineerevaluation
posted 2d ago · verified today
Amazon (AWS)hyperscaler
Austin, Texas, USA · Dallas, Texas, USA +2 more · onsite · staff plus
cluster-datacentercollectivesdeepspeed-libdistributed-inferenceefa-fabric
posted 7w ago · verified today
DoorDashenterprise
San Francisco · San Francisco, CA +2 more · senior
fine-tuninggpu-genericinferenceinference-enginesml-platform
posted 10w ago · verified today
Runwareinference provider
United Kingdom · remote · senior
inferenceinference-enginesgpu-genericgpu-kernelsperformance-engineer
posted 7d ago · verified today
DigitalOceanneocloud
Bay Area Metro · Seattle · $220K–$239K/year est. · staff plus
amdcudainferencekubernetes-opsnvidia
posted 4mo ago · verified today
DigitalOceanneocloud
Bay Area Metro · San Francisco · $220K–$239K/year est. · staff plus
amdcudadistributed-inferencego-langinference
posted 4mo ago · verified today
DigitalOceanneocloud
Bay Area Metro · New York · $150K–$186K/year est. · senior
solutions-architectcudagpu-genericinferencekubernetes-ops
posted 5mo ago · verified today
DigitalOceanneocloud
Bay Area Metro · San Francisco · $150K–$215K/year est. · senior
cudakubernetes-opssolutions-architectfine-tuninggpu-generic
posted 5mo ago · verified today
Interfazeinference provider
San Francisco, CA · senior
fine-tuninginferenceinference-enginespython-langquantization
added 7d ago · verified today
Together AIneocloud
San Francisco · $270K–$300K/year est. · senior
distributed-inferencefine-tuninginferenceinference-engineskv-cache-systems
posted 4mo ago · verified today
Lila Sciencesai startup
Alewife, Cambridge, MA · Cambridge, MA USA +2 more · $180K–$298K · staff plus
evaluationpython-langfine-tuningdeepspeed-libmegatron-lm
posted 11w ago · verified today
Lila Sciencesai startup
San Francisco, CA · San Francisco, CA USA · $268K–$384K · staff plus
ml-platformgpu-genericsoftware-engineertraining-frameworksinference-engines
posted 4mo ago · verified today
Lila Sciencesai startup
Cambridge, MA USA · One Charles Park, Cambridge, MA +1 more · $189K–$289K · unknown
post-trainingpython-langresearch-engineerreinforcement-learningtraining-frameworks
posted 11mo ago · verified today
Echo Neurotechnologiesai startup
San Francisco · SF Office · senior
ml-platformpython-langpytorch-distsoftware-engineertraining-frameworks
posted 8w ago · verified today
Hedraai startup
San Francisco · $175K–$275K/year · unknown
post-trainingpre-trainingresearch-engineerdeepspeed-libfsdp
posted 5mo ago · verified today
Tenstorrentchip vendor
Tokyo · Tokyo, Japan · unknown
performance-engineerpython-langfine-tuninginferencepre-training
posted 9mo ago · verified today
Lightning AIai startup
London, UK · New York, New York +6 more · remote · $165K–$310K · senior
python-langresearch-engineertraining-frameworksevaluationfine-tuning
posted 5w ago · verified today
Coherefrontier lab
London · Montreal +3 more · remote · unknown
post-trainingpython-langresearch-engineerjax-pallaspytorch-dist
posted 13mo ago · verified today
Coherefrontier lab
London · Montreal +4 more · hybrid · unknown
post-trainingfine-tuningjax-pallaskubernetes-opsml-platform
posted 15mo ago · verified today
Prime Intellectneocloud
San Francisco · onsite · unknown
kubernetes-opsml-platformfine-tuninggpu-genericnvidia
posted 10w ago · verified today
Prime Intellectneocloud
San Francisco · unknown
python-langsoftware-engineercluster-datacentergpu-generickubernetes-ops
posted 10w ago · verified today
TensorWaveneocloud
Las Vegas, Nevada · onsite · mid
amdgpu-generickubernetes-opsreliability-srescheduling-orchestration
posted 3w ago · verified today
Rekafrontier lab
US, UK, Singapore, Remote · remote · unknown
cpp-langcudafine-tuninggpu-genericgpu-kernels
posted 8mo ago · verified today
Mistral AIfrontier lab
Amsterdam · Lausanne +3 more · hybrid · unknown
post-trainingreinforcement-learningresearch-engineerevaluationfine-tuning
posted 12d ago · verified today
Thinking Machines Labfrontier lab
New York · San Francisco · onsite · unknown
reliability-sresrefine-tuninggpu-generickubernetes-ops
posted 15d ago · verified today
Nscaleneocloud
Austin, TX · Houston +3 more · staff plus
ansiblecluster-datacentercollectivescpp-langdeepspeed-lib
posted 13d ago · verified today
Mistral AIfrontier lab
Palo Alto · hybrid · unknown
ml-platformscheduling-orchestrationkubernetes-opspython-langcuda
posted 13d ago · verified today
Nscaleneocloud
London · UK · senior
evaluationfine-tuninggpu-genericinferenceinference-engines
posted 5mo ago · verified today
Amazon (AWS)hyperscaler
Boston, Massachusetts, USA · Seattle, Washington, USA +1 more · onsite · senior
gpu-kernelsinferenceinference-engineskv-cache-systemsperformance-engineer
posted 20d ago · verified today