AI infrastructure roles, filterable by the stack you actually work on.

Distributed training · inference serving · GPU fleets · network fabric — aggregated straight from company boards, never a copy of a copy.

1296 open roles · 85 companies · last verified today

116 roles

nccl-lib

Amazon (AWS)hyperscaler

Austin, Texas, USA · Cupertino, California, USA · onsite · mid

cpp-langsoftware-engineercollectivespython-langrust-lang

posted 5mo ago · verified today

Amazon (AWS)hyperscaler

Cupertino, California, USA · Seattle, Washington, USA · onsite · manager

collectivescpp-langcudagpu-kernelsnccl-lib

posted 6w ago · verified today

Amazon (AWS)hyperscaler

Boston, Massachusetts, USA · Seattle, Washington, USA +1 more · onsite · senior

cudacutlass-cutedistributed-inferenceevaluationflash-attention

posted 4w ago · verified today

FluidStackneocloud

Austin, TX · New York, NY +2 more · onsite · $173K–$224K/year · unknown

reliability-sresreinfiniband-opskubernetes-opsnccl-lib

posted 8w ago · verified today

Perplexityai startup

New York City · Palo Alto +1 more · $220K–$485K/year · mid

cudagpu-kernelsinferenceinference-enginescutlass-cute

posted 5mo ago · verified today

Baseteninference provider

Montreal · New York +2 more · hybrid · $165K–$330K/year · unknown

cpp-langnetwork-fabricnvidiasoftware-engineercollectives

posted 6mo ago · verified today

Baseteninference provider

New York · San Francisco · hybrid · $165K–$330K/year · unknown

ml-platformsoftware-engineerfine-tuningkubernetes-opspost-training

posted 7mo ago · verified today

Baseteninference provider

New York · San Francisco · hybrid · $165K–$330K/year · unknown

go-langkubernetes-opsscheduling-orchestrationsoftware-engineerml-platform

posted 12mo ago · verified today

SambaNovachip vendor

San Jose, CA · San Jose, California, United States · staff plus

custom-asicgpu-kernelsnetwork-fabricrdma-verbsroce-net

posted 5mo ago · verified today

xAIfrontier lab

Austin, TX · Bastrop, TX +2 more · senior

cluster-datacentergpu-genericsolutions-architectinfiniband-opsnetwork-fabric

posted 8w ago · verified today

CoreWeaveneocloud

Bellevue, WA · Sunnyvale, CA · $188K–$275K/year est. · staff plus

inferenceinference-enginesdistributed-inferencegpu-generickubernetes-ops

posted 4mo ago · verified today

CoreWeaveneocloud

Bellevue, WA · San Francisco, CA +1 more · $188K–$275K/year est. · senior

gpu-genericnvidiacluster-datacenterinfiniband-opsnetwork-fabric

posted 3mo ago · verified today

CoreWeaveneocloud

Bellevue, WA · Sunnyvale, CA · $182K–$242K/year est. · senior

kubernetes-opsperformance-engineerpython-langsoftware-engineergo-lang

posted 4mo ago · verified today

CoreWeaveneocloud

Bellevue, WA · Sunnyvale, CA · $152K–$204K/year est. · senior

cudadistributed-inferencego-langgpu-genericgpu-kernels

posted 11mo ago · verified today

CoreWeaveneocloud

Bellevue, WA · Sunnyvale, CA · $139K–$204K/year est. · senior

software-engineerinferenceinference-engineskubernetes-opsgpu-generic

posted 7mo ago · verified today

CoreWeaveneocloud

Bellevue, WA · Sunnyvale, CA · $182K–$242K/year est. · senior

go-langgpu-genericnetwork-fabricobservabilitypython-lang

posted 4w ago · verified today

CoreWeaveneocloud

Bellevue, WA · Sunnyvale, CA · $206K–$333K/year est. · staff plus

cudaperformance-engineergpu-genericinferenceinference-engines

posted 9mo ago · verified today

OpenAIfrontier lab

London, UK · New York City +2 more · hybrid · $230K–$405K/year · unknown

cluster-datacentercollectivesgpu-genericinferencekubernetes-ops

posted 4w ago · verified today

OpenAIfrontier lab

San Francisco · Seattle · hybrid · $293K–$385K/year · senior

collectivescpp-langcudadistributed-inferencegpu-generic

posted 5mo ago · verified today

OpenAIfrontier lab

San Francisco · hybrid · $295K–$500K/year · unknown

performance-engineercollectivescpp-langcudagpu-generic

posted 11mo ago · verified today

OpenAIfrontier lab

San Francisco · $266K–$500K/year · unknown

cudainferenceinference-enginesnvidiasoftware-engineer

posted 19mo ago · verified today