AI infrastructure roles, filterable by the stack you actually work on.
Distributed training · inference serving · GPU fleets · network fabric — aggregated straight from company boards, never a copy of a copy.
983 open roles · 19 companies · last verified today
21 roles
CoreWeaveneocloud
San Francisco, CA · San Francisco, CA / Seattle, WA +1 more · senior
solutions-architectgpu-genericinfiniband-opskubernetes-opsnetwork-fabric
posted 1d ago · verified today
Amazon (AWS)hyperscaler
Bellevue, Washington, USA · mid
ml-platformgpu-genericreliability-srescheduling-orchestrationsoftware-engineer
posted 15d ago · verified today
Amazon (AWS)hyperscaler
San Francisco, California, USA · senior
cudagpu-genericgpu-kernelsnvidiapython-lang
posted 11w ago · verified today
Amazon (AWS)hyperscaler
Cupertino, California, USA · manager
deepspeed-libeng-managerfsdpjax-pallasmegatron-lm
posted 9w ago · verified today
Amazon (AWS)hyperscaler
Bellevue, Washington, USA · mid
ml-platformsoftware-engineerscheduling-orchestrationtraining-frameworksfsdp
posted 15d ago · verified today
Perplexityai startup
Palo Alto · San Francisco · $220K–$485K · mid
post-trainingpython-langmegatron-lmpytorch-distreinforcement-learning
posted 4mo ago · verified today
Cerebraschip vendor
US and Canada Offices · mid
evaluationfine-tuningpost-trainingpython-langpytorch-dist
posted 5mo ago · verified today
Baseteninference provider
New York · San Francisco · remote · $165K–$330K · senior
go-langkubernetes-opsml-platformnetwork-fabricobservability
posted 11mo ago · verified today
Baseteninference provider
New York · San Francisco · remote · $165K–$330K · senior
software-engineerfine-tuninggpu-genericml-platformpost-training
posted 6mo ago · verified today
SambaNovachip vendor
San Jose, CA · San Jose, California, United States · staff plus
inferenceinference-enginesperformance-engineercustom-asicquantization
posted 3d ago · verified today
SambaNovachip vendor
San Jose, CA · San Jose, California, United States · senior
performance-engineergpu-genericgpu-kernelsinferenceinference-engines
posted 3d ago · verified today
CoreWeaveneocloud
Bellevue, WA · Sunnyvale, CA +1 more · staff plus
post-trainingreinforcement-learningfine-tuninggpu-genericinference-engines
posted 3d ago · verified today
CoreWeaveneocloud
Bellevue, WA · Bellevue, WA / Sunnyvale, CA +1 more · senior
post-trainingpython-langreinforcement-learningresearch-engineerfine-tuning
posted 3d ago · verified today
CoreWeaveneocloud
Bellevue, WA · Sunnyvale, CA +1 more · senior
gpu-generickubernetes-opsml-platformpython-langnvidia
posted 3d ago · verified today
CoreWeaveneocloud
Bellevue, WA · Sunnyvale, CA +1 more · staff plus
inferenceinference-enginesnvidiaperformance-engineertraining-frameworks
posted 3d ago · verified today
CoreWeaveneocloud
Bellevue, WA · Sunnyvale, CA +1 more · senior
cudagpu-kernelsperformance-engineercpp-langinference
posted 3d ago · verified today
Nebiusneocloud
Canada · Canada; Remote - United States +1 more · remote · manager
solutions-architectcluster-datacenterdistributed-inferencefine-tuninginference
posted 3d ago · verified today
Nebiusneocloud
Palo Alto · Palo Alto, California, United States · senior
training-frameworksmodel-parallelismpost-trainingpython-langpytorch-dist
posted 3d ago · verified today
Nebiusneocloud
Amsterdam, Netherlands; Remote - Europe; Remote - United States · Czech Republic +3 more · remote · mid
performance-engineercudagpu-genericgpu-kernelsnccl-lib
posted 3d ago · verified today
Lambdaneocloud
Bellevue Office · San Francisco Office (Second St) +1 more · remote · $226K–$355K · senior
nvidiasolutions-architectdistributed-inferencegpu-genericinference
posted 16d ago · verified today
Together AIneocloud
San Francisco · mid
fine-tuninggpu-genericmodel-parallelismperformance-engineerpython-lang
posted 3d ago · verified today