AI infrastructure roles, filterable by the stack you actually work on.

Distributed training · inference serving · GPU fleets · network fabric — aggregated straight from company boards, never a copy of a copy.

1297 open roles · 85 companies · last verified today

449 roles

reliability-sre

DigitalOceanneocloud

Atlanta · *United States · manager

cluster-datacenterdatacenter-engineergpu-genericreliability-sre

posted 20d ago · verified today

DigitalOceanneocloud

Seattle Metro · Wenatchee · manager

cluster-datacenterdatacenter-engineergpu-genericreliability-sre

posted 20d ago · verified today

DigitalOceanneocloud

Austin Metro · Seattle · $107K–$134K/year est. · mid

reliability-sresregpu-generickubernetes-opsgo-lang

posted 7mo ago · verified today

Scalewayneocloud

Paris · hybrid · manager

eng-managercluster-datacenterkubernetes-opsnvidiareliability-sre

posted 3w ago · verified today

DeepInfrainference provider

Bulgaria - Remote · remote · mid

cpp-langcudagpu-genericinferenceinference-engines

posted 5mo ago · verified today

DeepInfrainference provider

Palo Alto, United States · onsite · mid

cudainferencepython-langsoftware-engineercpp-lang

posted 5mo ago · verified today

DeepInfrainference provider

Palo Alto, United States · onsite · $140K–$150K/year est. · junior

inference-enginessoftware-engineercpp-langcudaml-platform

posted 8mo ago · verified today

DeepInfrainference provider

Bulgaria - Remote · remote · junior

cpp-langcudainferenceinference-enginespython-lang

posted 11mo ago · verified today

Fish Audio (39 AI)ai startup

Location not specified · senior

cluster-datacentergpu-genericreliability-sresrekubernetes-ops

added 7d ago · verified today

Waferinference provider

San Francisco · onsite · $200K–$300K/year · unknown

gpu-kernelsinferenceinference-enginescluster-datacenterperformance-engineer

posted 8w ago · verified today

Runwareinference provider

Remote · remote · senior

reliability-sresrego-langkubernetes-opsobservability

posted 8w ago · verified today

Runwareinference provider

United Kingdom · remote · senior

gpu-genericinferencenvidiareliability-sresre

posted 4mo ago · verified today

Relaceinference provider

San Francisco · onsite · mid

software-engineercluster-datacenterscheduling-orchestrationinferenceml-platform

posted 10mo ago · verified today

Exaai startup

Singapore · onsite · 90K–300K SGD/year · unknown

cluster-datacenterkubernetes-opsscheduling-orchestrationsoftware-engineerdistributed-inference

posted 6mo ago · verified today

Inceptionfrontier lab

San Mateo, United States · onsite · senior

inferenceinference-enginespython-langsoftware-engineergpu-generic

posted 6mo ago · verified today

Inceptionfrontier lab

San Mateo, United States · onsite · unknown

reinforcement-learningml-platformpost-trainingray-distributedsoftware-engineer

posted 6mo ago · verified today

Inceptionfrontier lab

San Mateo, United States · onsite · unknown

cudagpu-genericinferenceinference-engineskubernetes-ops

posted 6mo ago · verified today

Arcee AIfrontier lab

San Francisco, CA · SF, CA · unknown

cluster-datacentergpu-generickubernetes-opsreliability-sresre

posted 3mo ago · verified today

Black Forest Labsfrontier lab

San Francisco (United States) · onsite · unknown

cudagpu-genericinferenceinference-enginesperformance-engineer

posted 24mo ago · verified today

Baseteninference provider

San Francisco · hybrid · $170K–$230K/year · manager

cluster-datacentergo-langgpu-generickubernetes-opsnvidia

posted 8d ago · verified today

Fireworks AIinference provider

New York · San Mateo · hybrid · $200K–$230K/year · unknown

scheduling-orchestrationcpp-langnetwork-fabricpython-langstorage-checkpointing

posted 8d ago · verified today

Nscaleneocloud

Houston · New York +2 more · $130K–$200K/year est. · mid

reliability-sresrego-langobservabilitypython-lang

posted 8d ago · verified today

Nscaleneocloud

Houston · San Francisco +1 more · $170K–$265K · senior

reliability-sresrego-langkubernetes-opsobservability

posted 5mo ago · verified today

FluidStackneocloud

Travelling · U.S. Remote · remote · $258K–$300K/year · staff plus

network-engineernetwork-fabricreliability-sre

posted 10d ago · verified today

Harveyai startup

San Francisco · hybrid · $231K–$340K/year · staff plus

inferenceinference-enginessoftware-engineerml-platformobservability

posted 9d ago · verified today