AI infrastructure roles, filterable by the stack you actually work on.

Distributed training · inference serving · GPU fleets · network fabric — aggregated straight from company boards, never a copy of a copy.

1296 open roles · 85 companies · last verified today

130 roles

slurm-admin

RadixArkai startup

Palo Alto, CA · Palo Alto Office · $200K–$400K/year est. · senior

cluster-datacenterscheduling-orchestrationgpu-generickubernetes-opsnetwork-fabric

posted 7mo ago · verified today

falinference provider

Remote - USA · remote · $180K–$250K/year · senior

gpu-generickubernetes-opssoftware-engineercluster-datacentergpu-virtualization

posted 4w ago · verified today

Inferactinference provider

Remote · remote · $200K–$400K/year · unknown

go-langkubernetes-opspython-langrust-langterraform-iac

posted 3w ago · verified today

Inferactinference provider

San Francisco · onsite · $200K–$400K/year · unknown

cluster-datacentergpu-genericreliability-sreansiblekubernetes-ops

posted 3w ago · verified today

Inferactinference provider

Singapore · onsite · 200K–400K SGD/year · unknown

cluster-datacentergo-langgpu-genericinferencekubernetes-ops

posted 12w ago · verified today

Inferactinference provider

San Francisco · onsite · $200K–$400K/year · unknown

kubernetes-opssoftware-engineercluster-datacentergo-langinference

posted 7mo ago · verified today

NexGen Cloudneocloud

London · London, England, United Kingdom, UK - Remote +1 more · remote · unknown

solutions-architectcudacudnn-libdeepspeed-libgpu-generic

posted 8w ago · verified today

NexGen Cloudneocloud

UK - Remote · remote · senior

cudanvidiacluster-datacenternetwork-fabriccudnn-lib

posted 4mo ago · verified today

San Francisco Compute Companyneocloud

Remote · San Francisco, CA · hybrid · $220K–$300K/year · senior

cluster-datacenterdatacenter-engineergpu-genericnetwork-fabricinfiniband-ops

posted 2d ago · verified today

TensorWaveneocloud

Las Vegas, Nevada · onsite · mid

amdgpu-generickubernetes-opsreliability-srescheduling-orchestration

posted 3w ago · verified today

TensorWaveneocloud

Las Vegas, Nevada · Remote · onsite · senior

gpu-generickubernetes-opsml-platformpython-langscheduling-orchestration

posted 10mo ago · verified today

Rekafrontier lab

US, UK, Singapore, Remote · remote · unknown

cpp-langcudafine-tuninggpu-genericgpu-kernels

posted 8mo ago · verified today

Nscaleneocloud

Houston · New York +2 more · $210K–$270K · staff plus

cluster-datacenterkubernetes-opsgpu-genericscheduling-orchestrationslurm-admin

posted 11d ago · verified today

Nscaleneocloud

London · UK · staff plus

solutions-architectgpu-generickubernetes-opsslurm-adminterraform-iac

posted 12d ago · verified today

Mistral AIfrontier lab

Amsterdam · Lausanne +3 more · hybrid · unknown

post-trainingreinforcement-learningresearch-engineerevaluationfine-tuning

posted 12d ago · verified today

Thinking Machines Labfrontier lab

New York · San Francisco · onsite · unknown

post-trainingreliability-sresrepython-langreinforcement-learning

posted 15d ago · verified today

Nscaleneocloud

US · $190K–$260K · staff plus

observabilitygo-langgpu-generickubernetes-opspython-lang

posted 3mo ago · verified today

Nscaleneocloud

Houston · New York +2 more · staff plus

nvidiasoftware-engineercluster-datacentercudago-lang

posted 4w ago · verified today

Nscaleneocloud

New York · $225K–$275K · staff plus

scheduling-orchestrationslurm-adminsoftware-engineergo-langpython-lang

posted 3mo ago · verified today

Nscaleneocloud

Houston · New York +2 more · $140K–$193K · staff plus

solutions-architectkubernetes-opsgpu-genericslurm-admininfiniband-ops

posted 19d ago · verified today