AI infrastructure roles, filterable by the stack you actually work on.
Distributed training · inference serving · GPU fleets · network fabric — aggregated straight from company boards, never a copy of a copy.
1296 open roles · 85 companies · last verified today
611 roles
Interfazeinference provider
San Francisco, CA · senior
fine-tuninginferenceinference-enginespython-langquantization
added 7d ago · verified today
Runwareinference provider
Remote · remote · senior
reliability-sresrego-langkubernetes-opsobservability
posted 8w ago · verified today
Runwareinference provider
United Kingdom · remote · senior
gpu-genericinferencenvidiareliability-sresre
posted 4mo ago · verified today
Exaai startup
San Francisco, California · onsite · $180K–$350K/year · unknown
software-engineergpu-generickubernetes-opsml-platformray-distributed
posted 12mo ago · verified today
Inceptionfrontier lab
San Mateo, United States · onsite · senior
inferenceinference-enginespython-langsoftware-engineergpu-generic
posted 6mo ago · verified today
Inceptionfrontier lab
San Mateo, United States · onsite · unknown
python-langsoftware-engineertraining-frameworksgpu-genericpytorch-dist
posted 6mo ago · verified today
Inceptionfrontier lab
San Mateo, United States · onsite · unknown
cudagpu-genericinferenceinference-engineskubernetes-ops
posted 6mo ago · verified today
Arcee AIfrontier lab
San Francisco, CA · SF, CA · unknown
cluster-datacentergpu-generickubernetes-opsreliability-sresre
posted 3mo ago · verified today
Black Forest Labsfrontier lab
San Francisco (United States) · onsite · unknown
cudagpu-genericinferenceinference-enginesperformance-engineer
posted 24mo ago · verified today
Baseteninference provider
San Francisco · hybrid · $170K–$230K/year · senior
cluster-datacentergpu-generickubernetes-opsnvidiascheduling-orchestration
posted 8d ago · verified today
Crusoeneocloud
San Francisco, CA - US · onsite · staff plus
inferenceinference-enginesperformance-engineersglang-enginevllm-engine
posted 8d ago · verified today
Xaira Therapeuticsai startup
Seattle, WA · Seattle, Washington, United States +1 more · $185K–$308K/year est. · senior
gpu-genericgpu-kernelsinferenceresearch-engineerinference-engines
posted 10mo ago · verified today
Nscaleneocloud
Houston · New York +2 more · $130K–$200K/year est. · mid
reliability-sresrego-langobservabilitypython-lang
posted 8d ago · verified today
Nscaleneocloud
Houston · San Francisco +1 more · $170K–$265K · senior
reliability-sresrego-langkubernetes-opsobservability
posted 5mo ago · verified today
Cerebraschip vendor
Toronto, CAN · hybrid · senior
cluster-datacentergo-langkubernetes-opsscheduling-orchestrationobservability
posted 8d ago · verified today
Nebiusneocloud
Abu Dhabi, UAE · Dubai +1 more · senior
solutions-architectansiblecudagpu-generickubernetes-ops
posted 9d ago · verified today
Nscaleneocloud
UK · manager
eng-manageransiblecluster-datacentergpu-genericobservability
posted 9d ago · verified today
Nscaleneocloud
London · UK · staff plus
kubernetes-opssoftware-engineergo-langobservabilitynetwork-fabric
posted 10d ago · verified today
Lila Sciencesai startup
Alewife, Cambridge, MA · Cambridge, MA USA +4 more · $224K–$294K · staff plus
gpu-generickubernetes-opsml-platformpython-langray-distributed
posted 6w ago · verified today
Lila Sciencesai startup
San Francisco, CA · San Francisco, CA USA · $268K–$384K · staff plus
ml-platformgpu-genericsoftware-engineertraining-frameworksinference-engines
posted 4mo ago · verified today
Prior Labsai startup
Berlin · Freiburg +1 more · onsite · unknown
cluster-datacentergpu-genericpython-langpytorch-distscheduling-orchestration
posted 6w ago · verified today
Genesis Molecular AIai startup
New York, NY · San Mateo, CA · hybrid · senior
research-engineergpu-genericpre-trainingpython-langcuda
posted 13mo ago · verified today
Luma AIai startup
Redwood City, CA · hybrid · manager
inferenceinference-enginesml-platformscheduling-orchestrationcluster-datacenter
posted 7w ago · verified today
Luma AIai startup
Redwood City, CA · hybrid · $235K–$353K/year · staff plus
gpu-genericreliability-sresrecluster-datacenterkubernetes-ops
posted 7w ago · verified today
Luma AIai startup
Redwood City, CA · hybrid · unknown
inferenceinference-enginessoftware-engineerkubernetes-opspython-lang
posted 7w ago · verified today
Luma AIai startup
Redwood City, CA · hybrid · $168K–$252K/year · senior
reliability-sresregpu-genericnvidiaamd
posted 7w ago · verified today
Luma AIai startup
Redwood City, CA · hybrid · unknown
pytorch-disttraining-frameworkscollectivescudafsdp
posted 7w ago · verified today
Luma AIai startup
Redwood City, CA · hybrid · unknown
fsdpgpu-genericinference-enginespost-trainingpytorch-dist
posted 7w ago · verified today
Luma AIai startup
Redwood City, CA · hybrid · unknown
cudagpu-kernelstriton-langgpu-genericnvidia
posted 7w ago · verified today
ElevenLabsai startup
United States · remote · unknown
gpu-genericnvidiacluster-datacentercudareliability-sre
posted 13d ago · verified today