AI infrastructure roles, filterable by the stack you actually work on.

Distributed training · inference serving · GPU fleets · network fabric — aggregated straight from company boards, never a copy of a copy.

1296 open roles · 85 companies · last verified today

128 roles

pre-training

Generalistai startup

San Francisco Bay Area (San Mateo) or Boston (Somerville) · onsite · $260K–$350K/year · unknown

gpu-genericnvidiasoftware-engineerinferencekubernetes-ops

posted 7mo ago · verified today

1Xai startup

San Carlos, CA · onsite · $250K–$350K/year · unknown

gpu-genericpre-trainingpython-langresearch-engineertraining-data-infra

posted 3mo ago · verified today

Lightning AIai startup

London, UK · New York, New York +6 more · remote · $165K–$310K · senior

python-langresearch-engineertraining-frameworksevaluationfine-tuning

posted 5w ago · verified today

Lightning AIai startup

New York, New York, United States · San Francisco, California +3 more · $115K–$140K · unknown

cudagpu-generickubernetes-opsnccl-libobservability

posted 3mo ago · verified today

Lightning AIai startup

London, England, United Kingdom · London, UK · £75K–£95K · unknown

cudagpu-generickubernetes-opsml-platformnccl-lib

posted 3mo ago · verified today

Coherefrontier lab

London · Montreal +4 more · remote · senior

training-frameworksdistributed-inferencepre-trainingcudakubernetes-ops

posted 9mo ago · verified today

Coherefrontier lab

Canada · Toronto · onsite · CA$295K–CA$535K/year · unknown

pre-trainingpython-langtraining-frameworksresearch-engineergpu-generic

posted 10mo ago · verified today

Coherefrontier lab

London · Montreal +4 more · remote · mid

jax-pallasml-platformpython-langpytorch-distsoftware-engineer

posted 19mo ago · verified today

Coherefrontier lab

London · Montreal +3 more · hybrid · unknown

cudagpu-kernelsperformance-engineertriton-langpre-training

posted 19mo ago · verified today

Coherefrontier lab

London · Montreal +4 more · remote · unknown

cudapython-langgpu-genericgpu-kernelsmlir-llvm

posted 22mo ago · verified today

Reflection AIfrontier lab

London · New York, NY +1 more · onsite · unknown

pre-trainingmodel-parallelismsoftware-engineertraining-frameworkscollectives

posted 5mo ago · verified today

RadixArkai startup

Palo Alto, CA · Palo Alto Office · $200K–$400K/year est. · unknown

distributed-inferenceinferenceinference-enginesperformance-engineernvidia

posted 4mo ago · verified today

RadixArkai startup

Palo Alto, CA · Palo Alto Office · $200K–$400K/year est. · senior

gpu-genericresearch-engineersoftware-engineertraining-frameworksinference

posted 7mo ago · verified today

RadixArkai startup

Palo Alto, CA · Palo Alto Office · $200K–$400K/year est. · mid

cudagpu-kernelsinference-enginescpp-langrocm-hip

posted 7w ago · verified today

RadixArkai startup

Palo Alto, CA · Palo Alto Office · $200K–$400K/year est. · senior

cluster-datacenterscheduling-orchestrationgpu-generickubernetes-opsnetwork-fabric

posted 7mo ago · verified today

Prime Intellectneocloud

Remote · San Francisco · unknown

cudadeepspeed-libfsdpgpu-kernelsmodel-parallelism

posted 10w ago · verified today

TensorWaveneocloud

Las Vegas, Nevada · onsite · mid

amdgpu-generickubernetes-opsreliability-srescheduling-orchestration

posted 3w ago · verified today

TensorWaveneocloud

Las Vegas, Nevada · Remote · remote · senior

ansiblecudakubernetes-opsnetwork-fabricpython-lang

posted 7w ago · verified today

Liquid AIfrontier lab

Boston · Remote +1 more · hybrid · unknown

cudagpu-genericgpu-kernelsperformance-engineercpp-lang

posted 13mo ago · verified today

Rekafrontier lab

US, UK, Singapore, Remote · remote · unknown

cpp-langcudafine-tuninggpu-genericgpu-kernels

posted 8mo ago · verified today

Amazon (AWS)hyperscaler

Vancouver, British Columbia, CAN · onsite · senior

ml-platformsoftware-engineerevaluationinferencekubernetes-ops

posted 13d ago · verified today

Nscaleneocloud

London · UK · senior

evaluationfine-tuninggpu-genericinferenceinference-engines

posted 5mo ago · verified today

Amazon (AWS)hyperscaler

Arlington, Virginia, USA · Denver, Colorado, USA +1 more · onsite · senior

solutions-architectkubernetes-opsscheduling-orchestrationcluster-datacentergpu-generic

posted 15d ago · verified today

Amazon (AWS)hyperscaler

Seattle, Washington, USA · onsite · manager

eng-managerml-platformperformance-engineercluster-datacentercuda

posted 16d ago · verified today