AI infrastructure roles, filterable by the stack you actually work on.

Distributed training · inference serving · GPU fleets · network fabric — aggregated straight from company boards, never a copy of a copy.

1297 open roles · 85 companies · last verified today

42 roles

model-parallelism

Nscaleneocloud

Houston · New York +2 more · $220K–$293K · staff plus

inferenceinference-engineskv-cache-systemspost-trainingpython-lang

posted 1d ago · verified today

Inferactinference provider

San Francisco · onsite · junior

amdcudagpu-kernelsinferenceinference-engines

posted 1d ago · verified today

Lila Sciencesai startup

Alewife, Cambridge, MA · Cambridge, MA USA · $192K–$272K · staff plus

inferenceinference-engineskubernetes-opsml-platformnvidia

posted 7w ago · verified today

Cognitionai startup

San Francisco · onsite · unknown

cpp-langgpu-genericmodel-parallelismpython-langcluster-datacenter

posted 5w ago · verified today

Reflection AIfrontier lab

London · New York, NY +1 more · onsite · unknown

pre-trainingmodel-parallelismsoftware-engineertraining-frameworkscollectives

posted 5mo ago · verified today

RadixArkai startup

Palo Alto, CA · Palo Alto Office · $200K–$400K/year est. · unknown

gpu-genericpost-traininggpu-kernelsmodel-parallelismtraining-frameworks

posted 7mo ago · verified today

Prime Intellectneocloud

Remote · San Francisco · unknown

cudadeepspeed-libfsdpgpu-kernelsmodel-parallelism

posted 10w ago · verified today

Amazon (AWS)hyperscaler

Vancouver, British Columbia, CAN · onsite · senior

ml-platformsoftware-engineerevaluationinferencekubernetes-ops

posted 13d ago · verified today

Nscaleneocloud

London · UK · senior

evaluationfine-tuninggpu-genericinferenceinference-engines

posted 5mo ago · verified today

Amazon (AWS)hyperscaler

Boston, Massachusetts, USA · Seattle, Washington, USA +1 more · onsite · senior

gpu-kernelsinferenceinference-engineskv-cache-systemsperformance-engineer

posted 20d ago · verified today

OpenAIfrontier lab

San Francisco · hybrid · $266K–$445K/year · unknown

software-engineerinference-enginesinferencekv-cache-systemsscheduling-orchestration

posted 3w ago · verified today

Amazon (AWS)hyperscaler

Cupertino, California, USA · Seattle, Washington, USA · onsite · mid

trainiummodel-parallelismtraining-frameworkscollectivesperformance-engineer

posted 4w ago · verified today

Amazon (AWS)hyperscaler

Cupertino, California, USA · Seattle, Washington, USA · onsite · senior

software-engineertrainiummodel-parallelismpre-trainingtraining-frameworks

posted 5w ago · verified today

Amazon (AWS)hyperscaler

Cupertino, California, USA · onsite · manager

eng-managertrainiumdeepspeed-libfsdpmodel-parallelism

posted 3mo ago · verified today

Amazon (AWS)hyperscaler

Tel Aviv-Yafo, Tel Aviv, ISR · onsite · senior

inferenceinference-enginesvllm-enginecudagpu-generic

posted 10w ago · verified today

Amazon (AWS)hyperscaler

Cupertino, California, USA · onsite · senior

gpu-kernelsvllm-enginemodel-parallelismdistributed-inferenceml-platform

posted 6w ago · verified today

Amazon (AWS)hyperscaler

Cupertino, California, USA · onsite · mid

gpu-kernelsinferenceinference-enginesdistributed-inferenceml-platform

posted 6w ago · verified today

Amazon (AWS)hyperscaler

Tel Aviv-Yafo, Tel Aviv, ISR · onsite · mid

cpp-langsoftware-engineercustom-asicdistributed-inferencegpu-kernels

posted 5w ago · verified today

Amazon (AWS)hyperscaler

Santa Clara, California, USA · onsite · unknown

cudagpu-kernelstriton-langgpu-genericmodel-parallelism

posted 4w ago · verified today

Amazon (AWS)hyperscaler

Houston, Texas, USA · onsite · senior

cudagpu-kernelsmodel-parallelismpytorch-disttraining-frameworks

posted 4w ago · verified today

Amazon (AWS)hyperscaler

Boston, Massachusetts, USA · Seattle, Washington, USA +1 more · onsite · senior

cudacutlass-cutedistributed-inferenceevaluationflash-attention

posted 4w ago · verified today