AI infrastructure roles, filterable by the stack you actually work on.

Distributed training · inference serving · GPU fleets · network fabric — aggregated straight from company boards, never a copy of a copy.

1296 open roles · 85 companies · last verified today

452 roles

reliability-sre

Lightning AIai startup

London, England, United Kingdom · London, UK · £75K–£95K · unknown

cudagpu-generickubernetes-opsml-platformnccl-lib

posted 3mo ago · verified today

Lightning AIai startup

Philippines · Remote +1 more · unknown

reliability-sresrecluster-datacentergpu-generickubernetes-ops

posted 4mo ago · verified today

Etchedchip vendor

San Jose · onsite · manager

cluster-datacenterdatacenter-engineerreliability-sregpu-generic

posted 4mo ago · verified today

Etchedchip vendor

San Jose · onsite · $130K–$210K/year · unknown

cluster-datacenterdatacenter-engineernetwork-engineernetwork-fabricansible

posted 11w ago · verified today

Etchedchip vendor

San Jose · onsite · junior

cluster-datacentercpp-langnetwork-fabricpython-langreliability-sre

posted 4mo ago · verified today

Etchedchip vendor

San Jose · onsite · manager

cluster-datacentereng-managerinferenceinference-enginesnetwork-fabric

posted 5mo ago · verified today

Etchedchip vendor

San Jose · onsite · $150K–$275K/year · mid

performance-engineergpu-genericinferencesoftware-engineerml-platform

posted 8mo ago · verified today

Coherefrontier lab

Canada · United States · onsite · manager

eng-managercluster-datacenterkubernetes-opsgpu-genericobservability

posted 13d ago · verified today

Coherefrontier lab

Montreal · New York +2 more · remote · unknown

kubernetes-opssrecpp-langgo-langgpu-generic

posted 8mo ago · verified today

Coherefrontier lab

Montreal · New York +2 more · hybrid · staff plus

inferenceinference-enginessoftware-engineergpu-generickubernetes-ops

posted 8mo ago · verified today

Coherefrontier lab

Canada · United States · hybrid · senior

cluster-datacentergo-langgpu-generickubernetes-opsnccl-lib

posted 13d ago · verified today

Reflection AIfrontier lab

London · New York, NY +1 more · onsite · unknown

kubernetes-opscluster-datacenternccl-libreliability-sregpu-generic

posted 6mo ago · verified today

RadixArkai startup

Palo Alto, CA · Palo Alto Office · $200K–$400K/year est. · unknown

gpu-genericpost-traininggpu-kernelsmodel-parallelismtraining-frameworks

posted 7mo ago · verified today

RadixArkai startup

Palo Alto, CA · Palo Alto Office · $200K–$400K/year est. · mid

gpu-genericinferenceinference-enginesreliability-sreamd

posted 4mo ago · verified today

RadixArkai startup

Palo Alto, CA · Palo Alto Office · $200K–$400K/year est. · senior

inference-enginesgpu-genericinferenceperformance-engineerdistributed-inference

posted 7mo ago · verified today

RadixArkai startup

Palo Alto, CA · Palo Alto Office · $200K–$400K/year est. · senior

cluster-datacenterscheduling-orchestrationgpu-generickubernetes-opsnetwork-fabric

posted 7mo ago · verified today

RadixArkai startup

Palo Alto, CA · Radixark · $200K–$400K/year est. · mid

ml-platformgo-langkubernetes-opsobservabilitypython-lang

posted 8mo ago · verified today

falinference provider

Remote - USA · remote · $180K–$250K/year · senior

gpu-generickubernetes-opssoftware-engineercluster-datacentergpu-virtualization

posted 4w ago · verified today

falinference provider

Remote - APAC · remote · senior

inferencekubernetes-opsobservabilitypython-langreliability-sre

posted 11w ago · verified today

falinference provider

San Francisco · $180K–$250K/year · mid

ansiblecudanvidiaobservabilitypython-lang

posted 6mo ago · verified today

falinference provider

Remote - Global · unknown

python-langgpu-genericcudanvidiareliability-sre

posted 6mo ago · verified today

Parasailinference provider

San Mateo · senior

inferencereliability-sresreobservabilitykubernetes-ops

posted 7w ago · verified today

Inferactinference provider

Remote · remote · $200K–$400K/year · unknown

go-langkubernetes-opspython-langrust-langterraform-iac

posted 3w ago · verified today

Inferactinference provider

San Francisco · onsite · $200K–$400K/year · unknown

cluster-datacentergpu-genericreliability-sreansiblekubernetes-ops

posted 3w ago · verified today

Inferactinference provider

San Francisco · onsite · $200K–$400K/year · unknown

reliability-sresreinferenceobservabilityvllm-engine

posted 3w ago · verified today

Inferactinference provider

Singapore · onsite · 200K–400K SGD/year · unknown

software-engineerdistributed-inferenceinference-enginesinfiniband-opsnvlink-topology

posted 12w ago · verified today

Inferactinference provider

Singapore · onsite · 200K–400K SGD/year · unknown

cluster-datacentergo-langgpu-genericinferencekubernetes-ops

posted 12w ago · verified today