AI infrastructure roles, filterable by the stack you actually work on.
Distributed training · inference serving · GPU fleets · network fabric — aggregated straight from company boards, never a copy of a copy.
1295 open roles · 85 companies · last verified today
1295 roles
Coherefrontier lab
London · Montreal +4 more · remote · unknown
cudapython-langgpu-genericgpu-kernelsmlir-llvm
posted 22mo ago · verified today
Reflection AIfrontier lab
London · New York, NY +1 more · onsite · manager
eng-managerkubernetes-opscluster-datacenterscheduling-orchestrationgpu-generic
posted 4w ago · verified today
Reflection AIfrontier lab
London · New York, NY +1 more · onsite · unknown
software-engineerinference-enginesreinforcement-learningdistributed-inferenceevaluation
posted 5mo ago · verified today
Reflection AIfrontier lab
London · New York, NY +1 more · onsite · unknown
pre-trainingmodel-parallelismsoftware-engineertraining-frameworkscollectives
posted 5mo ago · verified today
Reflection AIfrontier lab
London · New York, NY +1 more · onsite · unknown
kubernetes-opscluster-datacenternccl-libreliability-sregpu-generic
posted 6mo ago · verified today
Reflection AIfrontier lab
London · New York, NY +1 more · onsite · unknown
pytorch-distreinforcement-learningsoftware-engineertraining-frameworkscollectives
posted 6mo ago · verified today
RadixArkai startup
Palo Alto, CA · Palo Alto Office · $160K–$300K/year est. · manager
distributed-inferenceinferenceinference-enginessglang-enginetraining-frameworks
posted 4mo ago · verified today
RadixArkai startup
Palo Alto, CA · Palo Alto Office · $200K–$400K/year est. · unknown
gpu-genericpost-traininggpu-kernelsmodel-parallelismtraining-frameworks
posted 7mo ago · verified today
RadixArkai startup
Palo Alto, CA · Palo Alto Office · $200K–$400K/year est. · mid
gpu-genericinferenceinference-enginesreliability-sreamd
posted 5mo ago · verified today
RadixArkai startup
Palo Alto, CA · Palo Alto Office · $200K–$400K/year est. · unknown
distributed-inferenceinferenceinference-enginesperformance-engineernvidia
posted 4mo ago · verified today
RadixArkai startup
Palo Alto, CA · Palo Alto Office · $200K–$400K/year est. · mid
distributed-inferenceinferencejax-pallastpuxla-compiler
posted 8mo ago · verified today
RadixArkai startup
Palo Alto, CA · Palo Alto Office · $200K–$400K/year est. · senior
gpu-genericresearch-engineersoftware-engineertraining-frameworksinference
posted 7mo ago · verified today
RadixArkai startup
Palo Alto, CA · Palo Alto Office · $200K–$400K/year est. · unknown
amdnvidiagpu-kernelscollectivescpp-lang
posted 6w ago · verified today
RadixArkai startup
Palo Alto, CA · Palo Alto Office · $200K–$400K/year est. · senior
cpp-langcudagpu-kernelsperformance-engineercollectives
posted 7mo ago · verified today
RadixArkai startup
Palo Alto, CA · Palo Alto Office · $200K–$400K/year est. · senior
inference-enginesgpu-genericinferenceperformance-engineerdistributed-inference
posted 7mo ago · verified today
RadixArkai startup
Palo Alto, CA · Palo Alto Office · $200K–$400K/year est. · mid
cudagpu-kernelsinference-enginescpp-langrocm-hip
posted 7w ago · verified today
RadixArkai startup
Palo Alto, CA · Palo Alto Office · $200K–$400K/year est. · senior
cluster-datacenterscheduling-orchestrationgpu-generickubernetes-opsnetwork-fabric
posted 7mo ago · verified today
RadixArkai startup
Palo Alto, CA · Radixark · $200K–$400K/year est. · mid
ml-platformgo-langkubernetes-opsobservabilitypython-lang
posted 8mo ago · verified today
RadixArkai startup
Palo Alto, CA · Palo Alto Office · junior
cpp-langgpu-kernelsinferenceinference-enginespython-lang
posted 8mo ago · verified today
falinference provider
Remote - USA · remote · $180K–$250K/year · senior
gpu-generickubernetes-opssoftware-engineercluster-datacentergpu-virtualization
posted 4w ago · verified today
falinference provider
Remote - APAC · remote · senior
inferencekubernetes-opsobservabilitypython-langreliability-sre
posted 12w ago · verified today
falinference provider
Remote - Global · senior
ml-platformpython-langrust-langscheduling-orchestrationsoftware-engineer
posted 6mo ago · verified today
falinference provider
San Francisco · $180K–$250K/year · mid
ansiblecudanvidiaobservabilitypython-lang
posted 6mo ago · verified today
falinference provider
Remote - Global · unknown
python-langgpu-genericcudanvidiareliability-sre
posted 6mo ago · verified today
falinference provider
San Francisco · $180K–$250K/year · unknown
ml-platformpython-langrust-langscheduling-orchestrationsoftware-engineer
posted 14mo ago · verified today
Parasailinference provider
San Mateo · senior
inferencereliability-sresreobservabilitykubernetes-ops
posted 7w ago · verified today
Parasailinference provider
San Mateo · senior
go-langsoftware-engineerkubernetes-opsml-platformscheduling-orchestration
posted 12w ago · verified today
Inferactinference provider
Remote · remote · unknown
inferenceinference-enginesvllm-enginepython-langsoftware-engineer
posted 12d ago · verified today
Inferactinference provider
Remote · remote · $200K–$400K/year · unknown
go-langkubernetes-opspython-langrust-langterraform-iac
posted 3w ago · verified today
Inferactinference provider
San Francisco · onsite · $200K–$400K/year · unknown
cluster-datacentergpu-genericreliability-sreansiblekubernetes-ops
posted 3w ago · verified today