AI infrastructure roles, filterable by the stack you actually work on.

Distributed training · inference serving · GPU fleets · network fabric — aggregated straight from company boards, never a copy of a copy.

1296 open roles · 85 companies · last verified today

128 roles

pre-training

SambaNovachip vendor

Stockholm, Sweden · staff plus

cpp-langsoftware-engineerinferencelinux-kernelnetwork-fabric

posted 19d ago · verified today

Nebiusneocloud

Abu Dhabi, UAE · Middle East +1 more · remote · senior

solutions-architectgpu-generickubernetes-opsml-platforminference

posted 9w ago · verified today

CoreWeaveneocloud

Bellevue, WA · San Francisco, CA +1 more · $198K–$264K/year est. · staff plus

amdcluster-datacenterdeepspeed-libdistributed-inferencefsdp

posted 4w ago · verified today

Amazon (AWS)hyperscaler

San Francisco, California, USA · onsite · senior

cudagpu-kernelsnvidiapython-langresearch-engineer

posted 3mo ago · verified today

Amazon (AWS)hyperscaler

Cupertino, California, USA · Seattle, Washington, USA · onsite · mid

trainiumperformance-engineersoftware-engineercollectivesgpu-generic

posted 7mo ago · verified today

Amazon (AWS)hyperscaler

Cupertino, California, USA · Seattle, Washington, USA · onsite · mid

trainiummodel-parallelismtraining-frameworkscollectivesperformance-engineer

posted 4w ago · verified today

Amazon (AWS)hyperscaler

Cupertino, California, USA · Seattle, Washington, USA · onsite · senior

software-engineertrainiummodel-parallelismpre-trainingtraining-frameworks

posted 5w ago · verified today

Amazon (AWS)hyperscaler

Cupertino, California, USA · onsite · manager

eng-managertrainiumdeepspeed-libfsdpmodel-parallelism

posted 3mo ago · verified today

Amazon (AWS)hyperscaler

Seattle, Washington, USA · onsite · staff plus

solutions-architectcudagpu-generickubernetes-opscluster-datacenter

posted 10w ago · verified today

Amazon (AWS)hyperscaler

Santa Clara, California, USA · onsite · unknown

cudagpu-kernelstriton-langgpu-genericmodel-parallelism

posted 4w ago · verified today

Amazon (AWS)hyperscaler

Houston, Texas, USA · onsite · senior

cudagpu-kernelsmodel-parallelismpytorch-disttraining-frameworks

posted 4w ago · verified today

Amazon (AWS)hyperscaler

Bellevue, Washington, USA · onsite · mid

software-engineerml-platformtrainiumscheduling-orchestrationfsdp

posted 6w ago · verified today

Amazon (AWS)hyperscaler

Herndon, Virginia, USA · New York, New York, USA +1 more · onsite · senior

solutions-architectgpu-genericinference-enginesmodel-parallelismtraining-frameworks

posted 3mo ago · verified today

Perplexityai startup

Palo Alto · San Francisco · $220K–$405K/year · unknown

cluster-datacentercpp-langgpu-genericinferencekubernetes-ops

posted 5mo ago · verified today

Cerebraschip vendor

Canada · United States · senior

amdcpp-langinferenceinference-enginespython-lang

posted 9mo ago · verified today

Cerebraschip vendor

Canada · United States · senior

cpp-langsoftware-engineerml-platforminferencepython-lang

posted 10mo ago · verified today

SambaNovachip vendor

San Jose, CA · San Jose, California, United States · staff plus

custom-asicgpu-kernelsnetwork-fabricrdma-verbsroce-net

posted 5mo ago · verified today

SambaNovachip vendor

San Jose, CA · San Jose, California, United States · $220K–$300K · staff plus

fine-tuninginferencesoftware-engineerinference-enginespre-training

posted 4w ago · verified today

Scale AIai startup

New York, NY · San Francisco, CA +1 more · $190K–$237K · unknown

cudaflash-attentioninference-enginesml-platformresearch-engineer

posted 18mo ago · verified today

xAIfrontier lab

Palo Alto, CA · $180K–$440K/year est. · unknown

gpu-genericpost-trainingpre-trainingpython-langresearch-engineer

posted 5mo ago · verified today

CoreWeaveneocloud

Bellevue, WA · Sunnyvale, CA · $206K–$333K/year est. · staff plus

cudaperformance-engineergpu-genericinferenceinference-engines

posted 9mo ago · verified today