AI infrastructure roles, filterable by the stack you actually work on.

Distributed training · inference serving · GPU fleets · network fabric — aggregated straight from company boards, never a copy of a copy.

1296 open roles · 85 companies · last verified today

77 roles

ray-distributed

Coherefrontier lab

London · Montreal +4 more · hybrid · unknown

post-trainingfine-tuningjax-pallaskubernetes-opsml-platform

posted 15mo ago · verified today

Coherefrontier lab

London · Montreal +4 more · remote · mid

jax-pallasml-platformpython-langpytorch-distsoftware-engineer

posted 19mo ago · verified today

RadixArkai startup

Palo Alto, CA · Palo Alto Office · $200K–$400K/year est. · senior

cluster-datacenterscheduling-orchestrationgpu-generickubernetes-opsnetwork-fabric

posted 7mo ago · verified today

Inferactinference provider

Remote · remote · $200K–$400K/year · unknown

go-langkubernetes-opspython-langrust-langterraform-iac

posted 3w ago · verified today

Inferactinference provider

San Francisco · onsite · manager

eng-managerinferenceinference-enginesvllm-enginedistributed-inference

posted 6w ago · verified today

Inferactinference provider

Singapore · onsite · 200K–400K SGD/year · unknown

cluster-datacentergo-langgpu-genericinferencekubernetes-ops

posted 12w ago · verified today

Inferactinference provider

San Francisco · onsite · $200K–$400K/year · unknown

kubernetes-opssoftware-engineercluster-datacentergo-langinference

posted 7mo ago · verified today

Prime Intellectneocloud

Remote · San Francisco · unknown

cudadeepspeed-libfsdpgpu-kernelsmodel-parallelism

posted 10w ago · verified today

Periodic Labsfrontier lab

Menlo Park, CA · onsite · unknown

collectivescudacutlass-cutefsdpgpu-generic

posted 4mo ago · verified today

Mistral AIfrontier lab

Amsterdam · Lausanne +3 more · hybrid · unknown

post-trainingreinforcement-learningresearch-engineerevaluationfine-tuning

posted 12d ago · verified today

Thinking Machines Labfrontier lab

New York · San Francisco · onsite · unknown

post-trainingreliability-sresrepython-langreinforcement-learning

posted 15d ago · verified today

Anyscaleai startup

San Francisco · $215K–$265K/year · senior

cpp-langray-distributedsoftware-engineerscheduling-orchestrationgpu-generic

posted 13d ago · verified today

SambaNovachip vendor

Austin, Texas, United States · Austin, TX +2 more · $210K–$280K · staff plus

inferenceml-platformobservabilityreliability-sresre

posted 9mo ago · verified today

Baseteninference provider

New York · San Francisco · hybrid · $200K–$400K/year · mid

inferenceinference-enginessolutions-architectevaluationperformance-engineer

posted 4w ago · verified today

Amazon (AWS)hyperscaler

Bellevue, Washington, USA · onsite · mid

software-engineerml-platformtrainiumscheduling-orchestrationfsdp

posted 6w ago · verified today

Baseteninference provider

San Francisco · hybrid · $200K–$275K/year · unknown

post-traininggpu-genericmodel-parallelismpytorch-distresearch-engineer

posted 5mo ago · verified today

Anyscaleai startup

San Francisco · hybrid · $200K–$240K/year · mid

kubernetes-opsray-distributedsoftware-engineergo-langml-platform

posted 4w ago · verified today

Anyscaleai startup

Palo Alto · San Francisco · hybrid · $170K–$245K/year · unknown

distributed-inferenceinferenceinference-enginessoftware-engineervllm-engine

posted 3mo ago · verified today

xAIfrontier lab

Palo Alto, CA · $180K–$440K/year est. · unknown

gpu-genericpost-trainingpre-trainingpython-langresearch-engineer

posted 5mo ago · verified today

CoreWeaveneocloud

Bellevue, WA · Sunnyvale, CA · $185K–$275K/year est. · staff plus

go-langkubernetes-opsscheduling-orchestrationsoftware-engineerslurm-admin

posted 6mo ago · verified today

CoreWeaveneocloud

Bellevue, WA · Sunnyvale, CA · $188K–$275K/year est. · staff plus

scheduling-orchestrationgo-langkubernetes-opsml-platformsoftware-engineer

posted 7mo ago · verified today