ai-infra-jobs

AI infrastructure roles, filterable by the stack you actually work on.

Distributed training · inference serving · GPU fleets · network fabric — aggregated straight from company boards, never a copy of a copy.

995 open roles · 19 companies · last verified today

Lambdaneocloud

Bellevue Office · San Francisco Office (Fremont St) · remote · $240K–$356K · senior

ansiblecluster-datacenterinfiniband-opsnccl-libnetwork-fabric

posted 6d ago · verified today

Lambdaneocloud

Remote, USA · San Jose Office (Zanker) · remote · $125K–$195K · senior

reliability-srecluster-datacentergpu-genericsrenetwork-fabric

posted 11w ago · verified today

Lambdaneocloud

San Francisco Office (Fremont St) · San Jose Office (First St) · remote · $296K–$346K · senior

cluster-datacenterscheduling-orchestrationreliability-srego-langkubernetes-ops

posted 3w ago · verified today

Lambdaneocloud

Bellevue Office · San Francisco Office (Fremont St) +1 more · remote · $266K–$395K · senior

go-langkubernetes-opspython-langscheduling-orchestrationsoftware-engineer

posted 6d ago · verified today

Lambdaneocloud

Bellevue Office · San Francisco Office (Fremont St) +1 more · remote · $314K–$465K · staff plus

kubernetes-opsgo-langgpu-genericml-platformnvidia

posted 15d ago · verified today

Lambdaneocloud

Remote, USA · remote · $122K–$162K · mid

cluster-datacentergpu-genericinfiniband-opsreliability-sresre

posted 20d ago · verified today

Lambdaneocloud

Bellevue Office · San Francisco Office (Fremont St) +1 more · remote · $266K–$395K · senior

software-engineerstorage-checkpointingcpp-langgo-langlinux-kernel

posted 5w ago · verified today

Lambdaneocloud

Bellevue Office · San Francisco Office (Second St) +1 more · remote · $226K–$355K · senior

nvidiasolutions-architectdistributed-inferencegpu-genericinference

posted 17d ago · verified today

Lambdaneocloud

Bellevue Office · San Francisco Office (Fremont St) +1 more · remote · $240K–$356K · senior

go-langkubernetes-opspython-langreliability-sresre

posted 6d ago · verified today

Together AIneocloud

San Francisco · San Francisco · mid

kubernetes-opsml-platformpython-langscheduling-orchestrationgpu-generic

posted 4d ago · verified today

Together AIneocloud

Remote · San Francisco, Singapore, Amsterdam · mid

cpp-langcudadistributed-inferencegpu-genericinference

posted 4d ago · verified today

Together AIneocloud

San Francisco · senior

solutions-architectpython-langgpu-genericansibleinference

posted 4d ago · verified today