AI infrastructure roles, filterable by the stack you actually work on.
Distributed training · inference serving · GPU fleets · network fabric — aggregated straight from company boards, never a copy of a copy.
1296 open roles · 85 companies · last verified today
185 roles
RadixArkai startup
Palo Alto, CA · Palo Alto Office · $200K–$400K/year est. · unknown
gpu-genericpost-traininggpu-kernelsmodel-parallelismtraining-frameworks
posted 7mo ago · verified today
RadixArkai startup
Palo Alto, CA · Palo Alto Office · $200K–$400K/year est. · mid
distributed-inferenceinferencejax-pallastpuxla-compiler
posted 8mo ago · verified today
RadixArkai startup
Palo Alto, CA · Palo Alto Office · $200K–$400K/year est. · senior
gpu-genericresearch-engineersoftware-engineertraining-frameworksinference
posted 7mo ago · verified today
RadixArkai startup
Palo Alto, CA · Palo Alto Office · $200K–$400K/year est. · unknown
amdnvidiagpu-kernelscollectivescpp-lang
posted 6w ago · verified today
RadixArkai startup
Palo Alto, CA · Palo Alto Office · $200K–$400K/year est. · mid
cudagpu-kernelsinference-enginescpp-langrocm-hip
posted 7w ago · verified today
RadixArkai startup
Palo Alto, CA · Palo Alto Office · $200K–$400K/year est. · senior
cluster-datacenterscheduling-orchestrationgpu-generickubernetes-opsnetwork-fabric
posted 7mo ago · verified today
RadixArkai startup
Palo Alto, CA · Palo Alto Office · junior
cpp-langinferenceinference-enginespython-langsglang-engine
posted 8mo ago · verified today
NexGen Cloudneocloud
London · London, England, United Kingdom, UK - Remote +1 more · remote · unknown
solutions-architectcudacudnn-libdeepspeed-libgpu-generic
posted 8w ago · verified today
NexGen Cloudneocloud
UK - Remote · remote · senior
cudanvidiacluster-datacenternetwork-fabriccudnn-lib
posted 4mo ago · verified today
Prime Intellectneocloud
Remote · San Francisco · unknown
cudapytorch-distreinforcement-learningresearch-engineertriton-lang
posted 10w ago · verified today
Prime Intellectneocloud
San Francisco · onsite · unknown
kubernetes-opsml-platformfine-tuninggpu-genericnvidia
posted 10w ago · verified today
Prime Intellectneocloud
San Francisco · unknown
post-trainingreinforcement-learningresearch-engineerevaluationml-platform
posted 10w ago · verified today
Prime Intellectneocloud
New York City, USA · hybrid · unknown
evaluationpost-trainingdistributed-inferenceray-distributedreinforcement-learning
posted 10w ago · verified today
Prime Intellectneocloud
San Francisco · unknown
python-langsoftware-engineercluster-datacentergpu-generickubernetes-ops
posted 10w ago · verified today
Prime Intellectneocloud
Remote · San Francisco · unknown
inference-enginesreinforcement-learningresearch-engineersglang-enginevllm-engine
posted 10w ago · verified today
Prime Intellectneocloud
Remote · San Francisco · unknown
cudadeepspeed-libfsdpgpu-kernelsmodel-parallelism
posted 10w ago · verified today
TensorWaveneocloud
Las Vegas, Nevada · onsite · mid
amdgpu-generickubernetes-opsreliability-srescheduling-orchestration
posted 3w ago · verified today
Liquid AIfrontier lab
San Francisco · hybrid · unknown
deepspeed-libfsdppytorch-distsoftware-engineertraining-frameworks
posted 13mo ago · verified today
Liquid AIfrontier lab
Boston · Remote +1 more · hybrid · unknown
cudagpu-genericgpu-kernelsperformance-engineercpp-lang
posted 13mo ago · verified today
Rekafrontier lab
US, UK, Singapore, Remote · remote · unknown
cpp-langcudafine-tuninggpu-genericgpu-kernels
posted 8mo ago · verified today
Periodic Labsfrontier lab
Menlo Park, CA · onsite · unknown
collectivescudacutlass-cutefsdpgpu-generic
posted 4mo ago · verified today
Mistral AIfrontier lab
Amsterdam · Lausanne +3 more · hybrid · unknown
post-trainingreinforcement-learningresearch-engineerevaluationfine-tuning
posted 12d ago · verified today
Thinking Machines Labfrontier lab
San Francisco · hybrid · unknown
reinforcement-learningresearch-engineerdistributed-inferenceinferenceinference-engines
posted 3w ago · verified today
Thinking Machines Labfrontier lab
San Francisco · hybrid · unknown
cudagpu-kernelstriton-langresearch-engineercutlass-cute
posted 6w ago · verified today
Thinking Machines Labfrontier lab
San Francisco · hybrid · unknown
research-engineertraining-frameworksdeepspeed-libgpu-genericmegatron-lm
posted 6w ago · verified today
Thinking Machines Labfrontier lab
San Francisco · hybrid · unknown
gpu-genericresearch-engineergpu-kernelsmodel-parallelismquantization
posted 6w ago · verified today
Amazon (AWS)hyperscaler
Vancouver, British Columbia, CAN · onsite · senior
ml-platformsoftware-engineerevaluationinferencekubernetes-ops
posted 13d ago · verified today
Nscaleneocloud
Austin, TX · Houston +3 more · staff plus
ansiblecluster-datacentercollectivescpp-langdeepspeed-lib
posted 13d ago · verified today
Nscaleneocloud
London · UK · senior
evaluationfine-tuninggpu-genericinferenceinference-engines
posted 5mo ago · verified today
Amazon (AWS)hyperscaler
Cupertino, California, USA · onsite · senior
gpu-kernelstrainiumperformance-engineersoftware-engineerinference-engines
posted 15d ago · verified today