AI infrastructure roles, filterable by the stack you actually work on.
Distributed training · inference serving · GPU fleets · network fabric — aggregated straight from company boards, never a copy of a copy.
1296 open roles · 85 companies · last verified today
85 roles
Prime Intellectneocloud
Remote · San Francisco · unknown
inference-enginesreinforcement-learningresearch-engineersglang-enginevllm-engine
posted 10w ago · verified today
Prime Intellectneocloud
Remote · San Francisco · unknown
cudadeepspeed-libfsdpgpu-kernelsmodel-parallelism
posted 10w ago · verified today
TensorWaveneocloud
Las Vegas, Nevada · onsite · mid
amdgpu-generickubernetes-opsreliability-srescheduling-orchestration
posted 3w ago · verified today
Rekafrontier lab
US, UK, Singapore, Remote · remote · unknown
cpp-langcudafine-tuninggpu-genericgpu-kernels
posted 8mo ago · verified today
Poolsidefrontier lab
Remote (EMEA) · remote · unknown
gpu-genericinferencescheduling-orchestrationgo-langinference-engines
posted 9w ago · verified today
Periodic Labsfrontier lab
Menlo Park, CA · onsite · unknown
collectivescudacutlass-cutefsdpgpu-generic
posted 4mo ago · verified today
Mistral AIfrontier lab
Amsterdam · Lausanne +3 more · hybrid · unknown
post-trainingreinforcement-learningresearch-engineerevaluationfine-tuning
posted 12d ago · verified today
Thinking Machines Labfrontier lab
San Francisco · hybrid · unknown
reinforcement-learningresearch-engineerdistributed-inferenceinferenceinference-engines
posted 3w ago · verified today
Thinking Machines Labfrontier lab
New York · San Francisco · onsite · unknown
post-trainingreliability-sresrepython-langreinforcement-learning
posted 15d ago · verified today
Nscaleneocloud
London · UK · senior
evaluationfine-tuninggpu-genericinferenceinference-engines
posted 5mo ago · verified today
Amazon (AWS)hyperscaler
Cupertino, California, USA · onsite · senior
fsdppost-trainingpre-trainingreinforcement-learningsoftware-engineer
posted 7mo ago · verified today
Amazon (AWS)hyperscaler
Boston, Massachusetts, USA · Seattle, Washington, USA +1 more · onsite · senior
gpu-kernelsinferenceinference-engineskv-cache-systemsperformance-engineer
posted 20d ago · verified today
Amazon (AWS)hyperscaler
Cupertino, California, USA · onsite · senior
training-frameworksdistributed-inferencecluster-datacentercollectivesdatacenter-engineer
posted 3w ago · verified today
Baseteninference provider
New York · San Francisco · hybrid · $200K–$400K/year · mid
inferenceinference-enginessolutions-architectevaluationperformance-engineer
posted 4w ago · verified today
Amazon (AWS)hyperscaler
Seattle, Washington, USA · onsite · senior
software-engineertraining-frameworkstrainiumpytorch-distfsdp
posted 11mo ago · verified today
Amazon (AWS)hyperscaler
Seattle, Washington, USA · onsite · mid
software-engineerml-platformpost-trainingreinforcement-learningtraining-frameworks
posted 3mo ago · verified today
Amazon (AWS)hyperscaler
Boston, Massachusetts, USA · Seattle, Washington, USA +1 more · onsite · senior
cudacutlass-cutedistributed-inferenceevaluationflash-attention
posted 4w ago · verified today
Cerebraschip vendor
Canada · United States · unknown
python-langevaluationpost-trainingreinforcement-learningfine-tuning
posted 6mo ago · verified today
Cerebraschip vendor
Canada · United States · senior
amdcpp-langinferenceinference-enginespython-lang
posted 9mo ago · verified today
Baseteninference provider
San Francisco · hybrid · $200K–$275K/year · unknown
post-traininggpu-genericmodel-parallelismpytorch-distresearch-engineer
posted 5mo ago · verified today
Baseteninference provider
New York · San Francisco · hybrid · $165K–$330K/year · unknown
ml-platformsoftware-engineerfine-tuningkubernetes-opspost-training
posted 7mo ago · verified today
Baseteninference provider
New York · San Francisco · hybrid · $165K–$330K/year · unknown
go-langkubernetes-opsscheduling-orchestrationsoftware-engineerml-platform
posted 12mo ago · verified today
SambaNovachip vendor
San Jose, CA · San Jose, California, United States · $220K–$300K · staff plus
fine-tuninginferencesoftware-engineerinference-enginespre-training
posted 4w ago · verified today
Scale AIai startup
New York, NY · San Francisco, CA · $290K–$363K · manager
cudaeng-managerpost-trainingtraining-frameworksflash-attention
posted 11mo ago · verified today
Scale AIai startup
New York, NY · San Francisco, CA · $180K–$225K · mid
software-engineertraining-data-inframl-platformscheduling-orchestrationevaluation
posted 13mo ago · verified today
Modalinference provider
Stockholm · mid
inferencesolutions-architectdistributed-inferencefine-tuninggpu-generic
posted 6mo ago · verified today
Modalinference provider
New York · San Francisco · $180K–$250K/year · unknown
solutions-architectsglang-enginevllm-engineperformance-engineerinference
posted 6mo ago · verified today
Fireworks AIinference provider
London · senior
fine-tuninginferenceinference-enginespost-trainingpython-lang
posted 14d ago · verified today
Fireworks AIinference provider
London · senior
fine-tuninginferenceinference-enginespython-langsglang-engine
posted 8w ago · verified today
xAIfrontier lab
Palo Alto, CA · $180K–$440K/year est. · unknown
cpp-langgpu-genericgpu-kernelsinferenceinference-engines
posted 23mo ago · verified today