AI infrastructure roles, filterable by the stack you actually work on.
Distributed training · inference serving · GPU fleets · network fabric — aggregated straight from company boards, never a copy of a copy.
1297 open roles · 85 companies · last verified today
408 roles
falinference provider
Remote - USA · remote · $180K–$250K/year · senior
gpu-generickubernetes-opssoftware-engineercluster-datacentergpu-virtualization
posted 4w ago · verified today
falinference provider
Remote - APAC · remote · senior
inferencekubernetes-opsobservabilitypython-langreliability-sre
posted 12w ago · verified today
falinference provider
Remote - Global · senior
ml-platformpython-langrust-langscheduling-orchestrationsoftware-engineer
posted 6mo ago · verified today
falinference provider
San Francisco · $180K–$250K/year · unknown
ml-platformpython-langrust-langscheduling-orchestrationsoftware-engineer
posted 14mo ago · verified today
Parasailinference provider
San Mateo · senior
inferencereliability-sresreobservabilitykubernetes-ops
posted 7w ago · verified today
Parasailinference provider
San Mateo · senior
go-langsoftware-engineerkubernetes-opsml-platformscheduling-orchestration
posted 12w ago · verified today
Inferactinference provider
Remote · remote · $200K–$400K/year · unknown
go-langkubernetes-opspython-langrust-langterraform-iac
posted 3w ago · verified today
Inferactinference provider
San Francisco · onsite · $200K–$400K/year · unknown
cluster-datacentergpu-genericreliability-sreansiblekubernetes-ops
posted 3w ago · verified today
Inferactinference provider
San Francisco · onsite · $200K–$400K/year · unknown
reliability-sresreinferenceobservabilityvllm-engine
posted 3w ago · verified today
Inferactinference provider
Singapore · onsite · 200K–400K SGD/year · unknown
software-engineerdistributed-inferenceinference-enginesinfiniband-opsnvlink-topology
posted 12w ago · verified today
Inferactinference provider
Singapore · onsite · 200K–400K SGD/year · unknown
cluster-datacentergo-langgpu-genericinferencekubernetes-ops
posted 12w ago · verified today
Inferactinference provider
San Francisco · onsite · $200K–$400K/year · unknown
kubernetes-opssoftware-engineercluster-datacentergo-langinference
posted 7mo ago · verified today
FriendliAIinference provider
San Francisco · hybrid · senior
cluster-datacentergpu-genericinferencekubernetes-opsml-platform
posted 5w ago · verified today
FriendliAIinference provider
Seoul · onsite · senior
kubernetes-opssoftware-engineernetwork-fabricscheduling-orchestrationcluster-datacenter
posted 5w ago · verified today
NexGen Cloudneocloud
London · London, England, United Kingdom, UK - Remote +1 more · remote · unknown
solutions-architectcudacudnn-libdeepspeed-libgpu-generic
posted 8w ago · verified today
NexGen Cloudneocloud
Canada - Remote · Quebec, Canada · unknown
kubernetes-opscluster-datacenterdatacenter-engineergpu-genericreliability-sre
posted 14d ago · verified today
Prime Intellectneocloud
San Francisco · onsite · unknown
kubernetes-opsml-platformfine-tuninggpu-genericnvidia
posted 10w ago · verified today
Prime Intellectneocloud
San Francisco · unknown
post-trainingreinforcement-learningresearch-engineerevaluationml-platform
posted 10w ago · verified today
Prime Intellectneocloud
Remote · San Francisco · hybrid · unknown
cluster-datacentergpu-genericinfiniband-opskubernetes-opsnetwork-fabric
posted 10w ago · verified today
Prime Intellectneocloud
New York City, USA · hybrid · unknown
evaluationpost-trainingdistributed-inferenceray-distributedreinforcement-learning
posted 10w ago · verified today
Prime Intellectneocloud
San Francisco · unknown
python-langsoftware-engineercluster-datacentergpu-generickubernetes-ops
posted 10w ago · verified today
Prime Intellectneocloud
Remote · San Francisco · unknown
inference-enginesreinforcement-learningresearch-engineersglang-enginevllm-engine
posted 10w ago · verified today
Radiantneocloud
London · hybrid · senior
ansibleobservabilitycluster-datacentergo-langgpu-generic
posted 3mo ago · verified today
Radiantneocloud
Gloucestershire · hybrid · senior
kubernetes-opsreliability-sreansibleobservabilitypython-lang
posted 4mo ago · verified today
TensorWaveneocloud
Las Vegas, Nevada · onsite · mid
amdgpu-generickubernetes-opsreliability-srescheduling-orchestration
posted 3w ago · verified today
TensorWaveneocloud
Remote · remote · staff plus
kubernetes-opscluster-datacenterscheduling-orchestrationreliability-sresoftware-engineer
posted 10w ago · verified today
TensorWaveneocloud
Remote · remote · senior
amdterraform-iacobservabilityreliability-sresre
posted 3mo ago · verified today
TensorWaveneocloud
Las Vegas, Nevada · Remote · onsite · senior
gpu-generickubernetes-opsml-platformpython-langscheduling-orchestration
posted 10mo ago · verified today
Liquid AIfrontier lab
San Francisco · hybrid · unknown
cluster-datacentergpu-genericreliability-sresoftware-engineersre
posted 7w ago · verified today
Poolsidefrontier lab
Remote (EMEA) · remote · unknown
gpu-genericinferencescheduling-orchestrationgo-langinference-engines
posted 9w ago · verified today