ai-infra-jobs

AI infrastructure roles, filterable by the stack you actually work on.

Distributed training · inference serving · GPU fleets · network fabric — aggregated straight from company boards, never a copy of a copy.

983 open roles · 19 companies · last verified today

Baseteninference provider

New York · San Francisco · remote · $200K–$400K · mid

inferenceinference-enginessolutions-architectdistributed-inferencepost-training

posted 1d ago · verified today

Amazon (AWS)hyperscaler

Herndon, Virginia, USA · New York, New York, USA · senior

solutions-architectinference-enginesnvidiatrainiumdistributed-inference

posted 9w ago · verified today

CoreWeaveneocloud

Austin, TX · US - Remote · senior

python-langsolutions-architectgpu-genericinferenceinference-engines

posted 1d ago · verified today

FluidStackneocloud

Austin, TX · New York, NY +2 more · $225K–$351K · manager

reliability-srecluster-datacentereng-managergpu-genericsre

posted 4w ago · verified today

FluidStackneocloud

Austin, TX · New York, NY +2 more · remote · $164K–$206K · mid

network-engineernetwork-fabriccluster-datacenterdatacenter-engineerinfiniband-ops

posted 9d ago · verified today

FluidStackneocloud

Austin, TX · New York, NY +2 more · $173K–$224K · mid

reliability-sresrecluster-datacentergpu-genericinfiniband-ops

posted 4w ago · verified today

Perplexityai startup

Palo Alto · San Francisco · $220K–$405K · mid

kubernetes-opsscheduling-orchestrationslurm-admincpp-langgpu-generic

posted 4mo ago · verified today

Baseteninference provider

San Francisco · remote · $200K–$275K · senior

post-trainingresearch-engineermodel-parallelismfine-tuninggpu-generic

posted 4mo ago · verified today

Baseteninference provider

Montreal · New York +2 more · remote · $165K–$330K · senior

cluster-datacenterkubernetes-opslinux-kernelscheduling-orchestrationml-platform

posted 3mo ago · verified today

xAIfrontier lab

Palo Alto, CA · Palo Alto, California · mid

gpu-genericml-platformpython-langcpp-langnvidia

posted 3d ago · verified today

CoreWeaveneocloud

Bellevue, WA · Bellevue, WA / Sunnyvale, CA +1 more · staff plus

go-langkubernetes-opsscheduling-orchestrationsoftware-engineerml-platform

posted 3d ago · verified today

CoreWeaveneocloud

Bellevue, WA · Livingston, NJ +4 more · manager

eng-managergpu-generickubernetes-opsnvidiascheduling-orchestration

posted 3d ago · verified today

CoreWeaveneocloud

Bellevue, WA · Sunnyvale, CA +1 more · senior

cudagpu-kernelsperformance-engineercpp-langinference

posted 3d ago · verified today

CoreWeaveneocloud

Bellevue, WA · Sunnyvale, CA +1 more · staff plus

go-langgpu-generickubernetes-opsml-platformscheduling-orchestration

posted 3d ago · verified today

CoreWeaveneocloud

Bellevue, WA · Bellevue, WA, Sunnyvale, CA, New York, NY, Livingston, NJ · manager

ml-platformscheduling-orchestrationgpu-generickubernetes-opsslurm-admin

posted 3d ago · verified today

CoreWeaveneocloud

Bellevue, WA · Sunnyvale, CA +1 more · staff plus

cluster-datacenterscheduling-orchestrationgo-langkubernetes-opsslurm-admin

posted 3d ago · verified today

CoreWeaveneocloud

Bellevue, WA · Bellevue, WA / Sunnyvale, CA +1 more · staff plus

go-langkubernetes-opsscheduling-orchestrationslurm-admingpu-generic

posted 3d ago · verified today

CoreWeaveneocloud

Washington, D.C. · Washington, DC · senior

solutions-architectpython-langfine-tuninggpu-genericinference

posted 3d ago · verified today