AI infrastructure roles, filterable by the stack you actually work on.
Distributed training · inference serving · GPU fleets · network fabric — aggregated straight from company boards, never a copy of a copy.
1297 open roles · 85 companies · last verified today
40 roles
Modalinference provider
San Francisco · onsite · $300K–$350K/year · manager
eng-managergpu-genericinferenceinference-enginesml-platform
posted today · verified today
Nscaleneocloud
Houston · New York +2 more · $220K–$293K · staff plus
inferenceinference-engineskv-cache-systemspost-trainingpython-lang
posted 1d ago · verified today
Amazon (AWS)hyperscaler
Seattle, Washington, USA · onsite · senior
inference-enginessoftware-engineersglang-enginetrainiumvllm-engine
posted 4d ago · verified today
Amazon (AWS)hyperscaler
Seattle, Washington, USA · onsite · senior
trainiuminferenceinference-enginesmoe-systemsperformance-engineer
posted 6d ago · verified today
Amazon (AWS)hyperscaler
Seattle, Washington, USA · onsite · senior
trainiuminferenceinference-enginesmoe-systemsperformance-engineer
posted 7d ago · verified today
DigitalOceanneocloud
Bay Area Metro · Seattle · $220K–$239K/year est. · staff plus
amdcudainferencekubernetes-opsnvidia
posted 4mo ago · verified today
DigitalOceanneocloud
Bay Area Metro · San Francisco · $220K–$239K/year est. · staff plus
amdcudadistributed-inferencego-langinference
posted 4mo ago · verified today
DigitalOceanneocloud
Bangalore Metro · Bengaluru · senior
cudakubernetes-opsnvidiaamdrocm-hip
posted 11w ago · verified today
Sakana AIfrontier lab
Tokyo · unknown
reliability-srenvidiainferencesregpu-generic
added 7d ago · verified today
Inceptionfrontier lab
San Mateo, United States · onsite · unknown
cudagpu-genericinferenceinference-engineskubernetes-ops
posted 6mo ago · verified today
Together AIneocloud
San Francisco · $270K–$300K/year est. · senior
distributed-inferencefine-tuninginferenceinference-engineskv-cache-systems
posted 4mo ago · verified today
Lila Sciencesai startup
Alewife, Cambridge, MA · Cambridge, MA USA · $192K–$272K · staff plus
inferenceinference-engineskubernetes-opsml-platformnvidia
posted 7w ago · verified today
Coherefrontier lab
Montreal · New York +2 more · remote · senior
cpp-langcudagpu-genericgpu-kernelsinference
posted 10mo ago · verified today
RadixArkai startup
Palo Alto, CA · Palo Alto Office · $200K–$400K/year est. · unknown
distributed-inferenceinferenceinference-enginesperformance-engineernvidia
posted 4mo ago · verified today
RadixArkai startup
Palo Alto, CA · Palo Alto Office · $200K–$400K/year est. · mid
cudagpu-kernelsinference-enginescpp-langrocm-hip
posted 7w ago · verified today
Thinking Machines Labfrontier lab
San Francisco · hybrid · unknown
reinforcement-learningresearch-engineerdistributed-inferenceinferenceinference-engines
posted 3w ago · verified today
Nscaleneocloud
London · UK · senior
evaluationfine-tuninggpu-genericinferenceinference-engines
posted 5mo ago · verified today
Amazon (AWS)hyperscaler
Boston, Massachusetts, USA · Seattle, Washington, USA +1 more · onsite · senior
gpu-kernelsinferenceinference-engineskv-cache-systemsperformance-engineer
posted 20d ago · verified today
Nebiusneocloud
Remote - United States · United States · remote · $228K–$285K · manager
eng-managerfine-tuninginferenceinference-engineskubernetes-ops
posted 4w ago · verified today
Amazon (AWS)hyperscaler
Tel Aviv-Yafo, Tel Aviv, ISR · onsite · mid
cpp-langsoftware-engineercustom-asicdistributed-inferencegpu-kernels
posted 5w ago · verified today
Amazon (AWS)hyperscaler
Boston, Massachusetts, USA · Seattle, Washington, USA +1 more · onsite · senior
cudacutlass-cutedistributed-inferenceevaluationflash-attention
posted 4w ago · verified today
Perplexityai startup
London · unknown
cudagpu-kernelsinferenceinference-enginesgpu-generic
posted 5mo ago · verified today
Perplexityai startup
New York City · Palo Alto +1 more · $220K–$485K/year · mid
cudagpu-kernelsinferenceinference-enginescutlass-cute
posted 5mo ago · verified today
Baseteninference provider
Montreal · New York +2 more · hybrid · $180K–$360K/year · mid
inferenceinference-enginessoftware-engineercudadistributed-inference
posted 11mo ago · verified today
Baseteninference provider
Montreal · New York +2 more · hybrid · $180K–$360K/year · unknown
cpp-langcudainferenceinference-engineskv-cache-systems
posted 30mo ago · verified today
SambaNovachip vendor
Austin, Texas, United States · Austin, TX +2 more · $200K–$275K · senior
inferenceinference-enginessoftware-engineerspeculative-decodingvllm-engine
posted 3mo ago · verified today
SambaNovachip vendor
San Jose, CA · San Jose, California, United States · $220K–$300K · staff plus
fine-tuninginferencesoftware-engineerinference-enginespre-training
posted 4w ago · verified today
Modalinference provider
New York · San Francisco · $150K–$350K/year · unknown
inferenceinference-engineskv-cache-systemsquantizationresearch-engineer
posted 10w ago · verified today
Fireworks AIinference provider
London · senior
fine-tuninginferenceinference-enginespost-trainingpython-lang
posted 14d ago · verified today
xAIfrontier lab
Palo Alto, CA · $180K–$440K/year est. · unknown
cpp-langgpu-genericgpu-kernelsinferenceinference-engines
posted 23mo ago · verified today