AI infrastructure roles, filterable by the stack you actually work on.
Distributed training · inference serving · GPU fleets · network fabric — aggregated straight from company boards, never a copy of a copy.
1294 open roles · 85 companies · last verified today
1294 roles
Amazon (AWS)hyperscaler
Arlington, Virginia, USA · Denver, Colorado, USA +1 more · onsite · unknown
kubernetes-opsnvidiascheduling-orchestrationgpu-genericsolutions-architect
posted 3w ago · verified today
Nebiusneocloud
Canada · Remote - United States · CA$235K–CA$300K/year est. · unknown
solutions-architectgpu-generickubernetes-opsterraform-iaccuda
posted 3w ago · verified today
Perplexityai startup
Palo Alto · San Francisco · $220K–$405K/year · senior
software-engineerml-platformpython-langrust-langinference
posted 3w ago · verified today
Nebiusneocloud
Remote - Europe · remote · senior
datacenter-engineercluster-datacenternetwork-fabric
posted 3w ago · verified today
Amazon (AWS)hyperscaler
New York, New York, USA · Sunnyvale, California, USA · onsite · senior
kubernetes-opsml-platforminferenceinference-enginesquantization
posted 3w ago · verified today
FluidStackneocloud
Austin, TX · New York, NY +2 more · hybrid · $162K–$202K/year · unknown
network-engineernetwork-fabriccluster-datacenter
posted 14d ago · verified today
Nebiusneocloud
South Korea · staff plus
solutions-architectcudagpu-genericansiblekubernetes-ops
posted 3w ago · verified today
Nebiusneocloud
Japan · staff plus
solutions-architectcudagpu-generickubernetes-opsml-platform
posted 3w ago · verified today
Together AIneocloud
India · Remote · senior
software-engineergo-langkubernetes-opsml-platformobservability
posted 3w ago · verified today
Crusoeneocloud
Denver, CO - US · onsite · staff plus
inferenceinference-enginesperformance-engineersglang-enginevllm-engine
posted 3w ago · verified today
OpenAIfrontier lab
San Francisco · hybrid · $266K–$445K/year · unknown
software-engineercpp-langpython-langobservabilityperformance-engineer
posted 3w ago · verified today
OpenAIfrontier lab
San Francisco · hybrid · $266K–$445K/year · unknown
software-engineergpu-kernelscpp-langrust-langlinux-kernel
posted 3w ago · verified today
Amazon (AWS)hyperscaler
Tel Aviv-Yafo, Tel Aviv, ISR · onsite · senior
cpp-langgpu-kernelsperformance-engineerpython-langtrainium
posted 3w ago · verified today
Anthropicfrontier lab
New York City, NY · San Francisco, CA +1 more · $320K–$485K · staff plus
inferenceinference-enginesscheduling-orchestrationsoftware-engineerdistributed-inference
posted 3w ago · verified today
Crusoeneocloud
San Francisco, CA - US · Sunnyvale, CA - US · onsite · manager
eng-managerkubernetes-opsslurm-admincluster-datacenterscheduling-orchestration
posted 3w ago · verified today
OpenAIfrontier lab
San Francisco · hybrid · $266K–$445K/year · unknown
software-engineerinference-enginesinferencekv-cache-systemsscheduling-orchestration
posted 3w ago · verified today
Lambdaneocloud
Bellevue Office · San Francisco Office (Fremont St) +1 more · remote · $297K–$440K · manager
eng-managercluster-datacentergpu-genericreliability-srescheduling-orchestration
posted 3w ago · verified today
Anthropicfrontier lab
Remote-Friendly (Travel Required) · Remote-Friendly US (Travel Required) +1 more · remote · $320K–$405K · manager
cluster-datacenterdatacenter-engineerreliability-sregpu-generic
posted 3w ago · verified today
Amazon (AWS)hyperscaler
Seattle, Washington, USA · onsite · unknown
software-engineertrainiumcluster-datacentercpp-langscheduling-orchestration
posted 3w ago · verified today
Nebiusneocloud
New York City, New York, United States · Remote - United States +1 more · remote · $200K–$245K · senior
gpu-genericnvidiasolutions-architectinferenceinference-engines
posted 7mo ago · verified today
xAIfrontier lab
Memphis, TN · Southaven, MS · senior
cluster-datacenterdatacenter-engineer
posted 3w ago · verified today
Nebiusneocloud
Finland · France +6 more · remote · manager
solutions-architecteng-managerkubernetes-ops
posted 3w ago · verified today
CoreWeaveneocloud
Bellevue, WA · Las Vegas, NV - DC +4 more · $120K–$145K/year est. · manager
cluster-datacenterdatacenter-engineereng-manager
posted 3w ago · verified today
Amazon (AWS)hyperscaler
New York, New York, USA · onsite · mid
software-engineertrainiumml-platformevaluation
posted 3w ago · verified today
Lambdaneocloud
Elk Grove Village, IL - Data Center · remote · $137K–$183K · manager
cluster-datacenterdatacenter-engineereng-managerinfiniband-opsnetwork-fabric
posted 3w ago · verified today
Amazon (AWS)hyperscaler
Arlington, Virginia, USA · Denver, Colorado, USA +1 more · onsite · senior
gpu-generickubernetes-opsscheduling-orchestrationsolutions-architectml-platform
posted 4w ago · verified today
Amazon (AWS)hyperscaler
Arlington, Virginia, USA · Denver, Colorado, USA +1 more · onsite · senior
cluster-datacentergpu-generickubernetes-opsnvidiascheduling-orchestration
posted 4w ago · verified today
SambaNovachip vendor
Bengaluru, India · Bengaluru, Karnataka, India · 10300K–12500K INR · staff plus
inference-enginessoftware-engineergo-langinferencekubernetes-ops
posted 4w ago · verified today
SambaNovachip vendor
Bengaluru, India · Bengaluru, Karnataka, India · 10300K–12500K INR · manager
cluster-datacenterobservabilitysoftware-engineerml-platformreliability-sre
posted 3w ago · verified today
SambaNovachip vendor
Bengaluru, India · Bengaluru, Karnataka, India · 8900K–10800K INR · staff plus
cluster-datacentercpp-langgo-langobservabilitypython-lang
posted 3w ago · verified today