AI infrastructure roles, filterable by the stack you actually work on.

Distributed training · inference serving · GPU fleets · network fabric — aggregated straight from company boards, never a copy of a copy.

1296 open roles · 85 companies · last verified today

185 roles

training-frameworks

RadixArkai startup

Palo Alto, CA · Palo Alto Office · $200K–$400K/year est. · unknown

gpu-genericpost-traininggpu-kernelsmodel-parallelismtraining-frameworks

posted 7mo ago · verified today

RadixArkai startup

Palo Alto, CA · Palo Alto Office · $200K–$400K/year est. · mid

distributed-inferenceinferencejax-pallastpuxla-compiler

posted 8mo ago · verified today

RadixArkai startup

Palo Alto, CA · Palo Alto Office · $200K–$400K/year est. · senior

gpu-genericresearch-engineersoftware-engineertraining-frameworksinference

posted 7mo ago · verified today

RadixArkai startup

Palo Alto, CA · Palo Alto Office · $200K–$400K/year est. · unknown

amdnvidiagpu-kernelscollectivescpp-lang

posted 6w ago · verified today

RadixArkai startup

Palo Alto, CA · Palo Alto Office · $200K–$400K/year est. · mid

cudagpu-kernelsinference-enginescpp-langrocm-hip

posted 7w ago · verified today

RadixArkai startup

Palo Alto, CA · Palo Alto Office · $200K–$400K/year est. · senior

cluster-datacenterscheduling-orchestrationgpu-generickubernetes-opsnetwork-fabric

posted 7mo ago · verified today

RadixArkai startup

Palo Alto, CA · Palo Alto Office · junior

cpp-langinferenceinference-enginespython-langsglang-engine

posted 8mo ago · verified today

NexGen Cloudneocloud

London · London, England, United Kingdom, UK - Remote +1 more · remote · unknown

solutions-architectcudacudnn-libdeepspeed-libgpu-generic

posted 8w ago · verified today

NexGen Cloudneocloud

UK - Remote · remote · senior

cudanvidiacluster-datacenternetwork-fabriccudnn-lib

posted 4mo ago · verified today

Prime Intellectneocloud

Remote · San Francisco · unknown

cudadeepspeed-libfsdpgpu-kernelsmodel-parallelism

posted 10w ago · verified today

TensorWaveneocloud

Las Vegas, Nevada · onsite · mid

amdgpu-generickubernetes-opsreliability-srescheduling-orchestration

posted 3w ago · verified today

Liquid AIfrontier lab

Boston · Remote +1 more · hybrid · unknown

cudagpu-genericgpu-kernelsperformance-engineercpp-lang

posted 13mo ago · verified today

Rekafrontier lab

US, UK, Singapore, Remote · remote · unknown

cpp-langcudafine-tuninggpu-genericgpu-kernels

posted 8mo ago · verified today

Periodic Labsfrontier lab

Menlo Park, CA · onsite · unknown

collectivescudacutlass-cutefsdpgpu-generic

posted 4mo ago · verified today

Mistral AIfrontier lab

Amsterdam · Lausanne +3 more · hybrid · unknown

post-trainingreinforcement-learningresearch-engineerevaluationfine-tuning

posted 12d ago · verified today

Amazon (AWS)hyperscaler

Vancouver, British Columbia, CAN · onsite · senior

ml-platformsoftware-engineerevaluationinferencekubernetes-ops

posted 13d ago · verified today

Nscaleneocloud

London · UK · senior

evaluationfine-tuninggpu-genericinferenceinference-engines

posted 5mo ago · verified today