AI infrastructure roles, filterable by the stack you actually work on.
Distributed training · inference serving · GPU fleets · network fabric — aggregated straight from company boards, never a copy of a copy.
1296 open roles · 85 companies · last verified today
83 roles
1Xai startup
San Carlos, CA · onsite · $250K–$350K/year · unknown
gpu-genericpre-trainingpython-langresearch-engineertraining-data-infra
posted 3mo ago · verified today
RadixArkai startup
Palo Alto, CA · Palo Alto Office · $200K–$400K/year est. · mid
gpu-genericinferenceinference-enginesreliability-sreamd
posted 4mo ago · verified today
RadixArkai startup
Palo Alto, CA · Palo Alto Office · $200K–$400K/year est. · unknown
amdnvidiagpu-kernelscollectivescpp-lang
posted 6w ago · verified today
RadixArkai startup
Palo Alto, CA · Palo Alto Office · $200K–$400K/year est. · mid
cudagpu-kernelsinference-enginescpp-langrocm-hip
posted 7w ago · verified today
Inferactinference provider
San Francisco · onsite · $200K–$400K/year · unknown
amdgpu-kernelsinferenceperformance-engineerrocm-hip
posted 11w ago · verified today
Inferactinference provider
Singapore · onsite · 200K–400K SGD/year · unknown
tpuinferenceinference-enginesjax-pallasperformance-engineer
posted 11w ago · verified today
Inferactinference provider
Singapore · onsite · 200K–400K SGD/year · unknown
amdgpu-kernelsinferenceinference-enginesperformance-engineer
posted 11w ago · verified today
Inferactinference provider
Singapore · onsite · $200K–$400K/year · unknown
gpu-kernelscudaperformance-engineercpp-langinference
posted 12w ago · verified today
Inferactinference provider
San Francisco · onsite · $200K–$400K/year · unknown
tpuinference-enginesjax-pallasperformance-engineerxla-compiler
posted 3mo ago · verified today
FriendliAIinference provider
Seoul · onsite · mid
amdcpp-langcudacutlass-cutegpu-kernels
posted 4mo ago · verified today
FriendliAIinference provider
San Francisco · hybrid · mid
amdcpp-langcudagpu-genericgpu-kernels
posted 6mo ago · verified today
Liquid AIfrontier lab
Boston · hybrid · unknown
cpp-langinferenceinference-enginespython-langevaluation
posted 3w ago · verified today
Thinking Machines Labfrontier lab
San Francisco · hybrid · unknown
reinforcement-learningresearch-engineerdistributed-inferenceinferenceinference-engines
posted 3w ago · verified today
Thinking Machines Labfrontier lab
San Francisco · hybrid · unknown
cudagpu-kernelstriton-langresearch-engineercutlass-cute
posted 6w ago · verified today
Thinking Machines Labfrontier lab
San Francisco · hybrid · unknown
gpu-genericresearch-engineergpu-kernelsmodel-parallelismquantization
posted 6w ago · verified today
Amazon (AWS)hyperscaler
Vancouver, British Columbia, CAN · onsite · senior
ml-platformsoftware-engineerevaluationinferencekubernetes-ops
posted 13d ago · verified today
Nscaleneocloud
London · UK · senior
evaluationfine-tuninggpu-genericinferenceinference-engines
posted 5mo ago · verified today
Amazon (AWS)hyperscaler
Boston, Massachusetts, USA · Seattle, Washington, USA +1 more · onsite · senior
gpu-kernelsinferenceinference-engineskv-cache-systemsperformance-engineer
posted 20d ago · verified today
Amazon (AWS)hyperscaler
New York, New York, USA · Sunnyvale, California, USA · onsite · senior
kubernetes-opsml-platforminferenceinference-enginesquantization
posted 3w ago · verified today
Nebiusneocloud
New York City, New York, United States · Remote - United States +1 more · remote · $200K–$245K · senior
gpu-genericnvidiasolutions-architectinferenceinference-engines
posted 7mo ago · verified today
Nebiusneocloud
Remote - United States · United States · remote · $228K–$285K · manager
eng-managerfine-tuninginferenceinference-engineskubernetes-ops
posted 4w ago · verified today
Amazon (AWS)hyperscaler
Boston, Massachusetts, USA · Seattle, Washington, USA +1 more · onsite · senior
cudacutlass-cutedistributed-inferenceevaluationflash-attention
posted 4w ago · verified today
Amazon (AWS)hyperscaler
Herndon, Virginia, USA · New York, New York, USA +1 more · onsite · senior
solutions-architectgpu-genericinference-enginesmodel-parallelismtraining-frameworks
posted 3mo ago · verified today
Perplexityai startup
London · unknown
cudagpu-kernelsinferenceinference-enginesgpu-generic
posted 5mo ago · verified today
Perplexityai startup
New York City · Palo Alto +1 more · $220K–$485K/year · mid
cudagpu-kernelsinferenceinference-enginescutlass-cute
posted 5mo ago · verified today
Cerebraschip vendor
Sunnyvale, CA · Toronto, CAN · hybrid · staff plus
amdinference-enginespython-langsoftware-engineervllm-engine
posted 7w ago · verified today
Cerebraschip vendor
Canada · United States · senior
amdcpp-langinferenceinference-enginespython-lang
posted 9mo ago · verified today
Baseteninference provider
Montreal · New York +2 more · hybrid · $180K–$360K/year · mid
inferenceinference-enginessoftware-engineercudadistributed-inference
posted 11mo ago · verified today
Baseteninference provider
Montreal · New York +2 more · hybrid · $180K–$360K/year · unknown
cudagpu-kernelscpp-langgpu-genericinference
posted 14mo ago · verified today
Baseteninference provider
Montreal · New York +2 more · hybrid · $180K–$360K/year · unknown
cpp-langcudainferenceinference-engineskv-cache-systems
posted 30mo ago · verified today