AI infrastructure roles, filterable by the stack you actually work on.
Distributed training · inference serving · GPU fleets · network fabric — aggregated straight from company boards, never a copy of a copy.
1296 open roles · 85 companies · last verified today
122 roles
CoreWeaveneocloud
Manhattan, NY · New York, NY · $153K–$204K/year est. · senior
object-storagesoftware-engineerstorage-checkpointinggo-langobservability
posted 19d ago · verified today
Amazon (AWS)hyperscaler
Cupertino, California, USA · onsite · senior
training-frameworksdistributed-inferencecluster-datacentercollectivesdatacenter-engineer
posted 3w ago · verified today
Amazon (AWS)hyperscaler
New York, New York, USA · Sunnyvale, California, USA · onsite · senior
kubernetes-opsml-platforminferenceinference-enginesquantization
posted 3w ago · verified today
Nebiusneocloud
New York City, New York, United States · Remote - United States +1 more · remote · $200K–$245K · senior
gpu-genericnvidiasolutions-architectinferenceinference-engines
posted 7mo ago · verified today
Lambdaneocloud
San Francisco Office (Second St) · San Jose Office (First St) · remote · $251K–$335K · unknown
gpu-generickubernetes-opscluster-datacentersolutions-architectnetwork-fabric
posted 3w ago · verified today
Amazon (AWS)hyperscaler
Boston, Massachusetts, USA · onsite · unknown
software-engineerparallel-fsstorage-checkpointing
posted 6w ago · verified today
Amazon (AWS)hyperscaler
Tel Aviv-Yafo, Tel Aviv, ISR · onsite · senior
solutions-architectscheduling-orchestrationcluster-datacenterparallel-fsgpu-generic
posted 3mo ago · verified today
Amazon (AWS)hyperscaler
Bellevue, Washington, USA · onsite · mid
ml-platformscheduling-orchestrationsoftware-engineertraining-frameworksstorage-checkpointing
posted 6w ago · verified today
Amazon (AWS)hyperscaler
Bellevue, Washington, USA · onsite · mid
software-engineerml-platformtrainiumscheduling-orchestrationfsdp
posted 6w ago · verified today
FluidStackneocloud
Austin, TX · New York, NY +2 more · onsite · $224K–$344K/year · manager
cluster-datacentereng-managergpu-virtualizationkubernetes-opsreliability-sre
posted 8w ago · verified today
Baseteninference provider
San Francisco · hybrid · $165K–$330K/year · manager
eng-managercpp-langgo-langgpu-genericinference
posted 3mo ago · verified today
Baseteninference provider
Montreal · New York +2 more · hybrid · $165K–$330K/year · unknown
cpp-langnetwork-fabricnvidiasoftware-engineercollectives
posted 6mo ago · verified today
Baseteninference provider
New York · San Francisco · hybrid · $165K–$330K/year · unknown
ml-platformsoftware-engineerfine-tuningkubernetes-opspost-training
posted 7mo ago · verified today
Baseteninference provider
New York · San Francisco · hybrid · $165K–$330K/year · unknown
go-langkubernetes-opsscheduling-orchestrationsoftware-engineerml-platform
posted 12mo ago · verified today
Fireworks AIinference provider
New York · San Mateo · hybrid · $210K–$320K/year · unknown
software-engineertraining-frameworksfsdpkubernetes-opsml-platform
posted 6w ago · verified today
CoreWeaveneocloud
Bellevue, WA · California +5 more · $207K–$303K/year est. · staff plus
object-storagesoftware-engineerstorage-checkpointingobservabilityparallel-fs
posted 6w ago · verified today
CoreWeaveneocloud
Bellevue, WA · Livingston, NJ +5 more · $207K–$303K/year est. · staff plus
cluster-datacentercpp-langgo-langinfiniband-opskubernetes-ops
posted 5w ago · verified today
CoreWeaveneocloud
Bellevue, WA · Livingston, NJ +3 more · $188K–$275K/year est. · staff plus
object-storagesoftware-engineergo-langrust-langkubernetes-ops
posted 10mo ago · verified today
CoreWeaveneocloud
Bellevue, WA · Livingston, NJ +4 more · $143K–$210K/year est. · senior
go-langkubernetes-opsobject-storageobservabilityparallel-fs
posted 5mo ago · verified today
OpenAIfrontier lab
San Francisco · hybrid · $380K–$500K/year · manager
eng-managertraining-data-infrapython-langsoftware-engineerstorage-checkpointing
posted 3mo ago · verified today
OpenAIfrontier lab
London, UK · New York City +2 more · hybrid · $230K–$405K/year · unknown
cluster-datacentercollectivesgpu-genericinferencekubernetes-ops
posted 4w ago · verified today
OpenAIfrontier lab
San Francisco · hybrid · $295K–$500K/year · unknown
performance-engineercollectivescpp-langcudagpu-generic
posted 11mo ago · verified today
OpenAIfrontier lab
San Francisco · hybrid · $230K–$385K/year · unknown
object-storagerust-langsoftware-engineerkubernetes-opsstorage-checkpointing
posted 21mo ago · verified today
OpenAIfrontier lab
New York City · San Francisco · $230K–$490K/year · unknown
cluster-datacentergpu-genericsoftware-engineerkubernetes-opsscheduling-orchestration
posted 19mo ago · verified today
OpenAIfrontier lab
San Francisco · $295K–$500K/year · unknown
python-langsoftware-engineertraining-frameworksgpu-genericperformance-engineer
posted 10mo ago · verified today
Anthropicfrontier lab
San Francisco, CA · $405K–$485K · staff plus
ml-platformsoftware-engineerpython-langscheduling-orchestrationstorage-checkpointing
posted 5w ago · verified today
Mistral AIfrontier lab
Palo Alto · San Francisco · hybrid · unknown
kubernetes-opsml-platformpython-langsoftware-engineertraining-data-infra
posted 9w ago · verified today
Mistral AIfrontier lab
London · Paris +2 more · hybrid · unknown
kubernetes-opsml-platformpython-langscheduling-orchestrationtraining-data-infra
posted 9w ago · verified today
Mistral AIfrontier lab
Amsterdam · Berlin +3 more · remote · manager
go-langkubernetes-opsml-platformeng-managerscheduling-orchestration
posted 12d ago · verified today
Mistral AIfrontier lab
London · Paris +2 more · unknown
research-engineerdeepspeed-libfsdppython-langslurm-admin
posted 10w ago · verified today