AI infrastructure roles, filterable by the stack you actually work on.
Distributed training · inference serving · GPU fleets · network fabric — aggregated straight from company boards, never a copy of a copy.
1296 open roles · 85 companies · last verified today
449 roles
Lila Sciencesai startup
Alewife, Cambridge, MA · Cambridge, MA USA +4 more · $224K–$294K · staff plus
gpu-generickubernetes-opsml-platformpython-langray-distributed
posted 6w ago · verified today
Lila Sciencesai startup
Alewife, Cambridge, MA · Cambridge, MA USA · $192K–$272K · staff plus
inferenceinference-engineskubernetes-opsml-platformnvidia
posted 7w ago · verified today
Lila Sciencesai startup
San Francisco, CA · San Francisco, CA USA · $268K–$384K · staff plus
ml-platformgpu-genericsoftware-engineertraining-frameworksinference-engines
posted 4mo ago · verified today
Prior Labsai startup
Berlin · Freiburg +1 more · onsite · unknown
cluster-datacentergpu-genericpython-langpytorch-distscheduling-orchestration
posted 6w ago · verified today
Sunoai startup
Boston · onsite · $229K–$364K/year · senior
software-engineerkubernetes-opsml-platforminference-enginesobservability
posted 29mo ago · verified today
Luma AIai startup
Redwood City, CA · hybrid · manager
inferenceinference-enginesml-platformscheduling-orchestrationcluster-datacenter
posted 7w ago · verified today
Luma AIai startup
Redwood City, CA · hybrid · $235K–$353K/year · staff plus
gpu-genericreliability-sresrecluster-datacenterkubernetes-ops
posted 7w ago · verified today
Luma AIai startup
Redwood City, CA · hybrid · unknown
inferenceinference-enginessoftware-engineerkubernetes-opspython-lang
posted 7w ago · verified today
Luma AIai startup
Redwood City, CA · hybrid · $168K–$252K/year · senior
reliability-sresregpu-genericnvidiaamd
posted 7w ago · verified today
Hedraai startup
San Francisco · $175K–$275K/year · senior
kubernetes-opsml-platformpython-langsoftware-engineerinference
posted 3mo ago · verified today
ElevenLabsai startup
United States · remote · unknown
gpu-genericnvidiacluster-datacentercudareliability-sre
posted 13d ago · verified today
Harveyai startup
New York · $161K–$242K/year · senior
cluster-datacenterkubernetes-opsreliability-sreml-platformobservability
posted 6w ago · verified today
Harveyai startup
San Francisco · $236K–$290K/year · staff plus
inferenceinference-enginesml-platformobservabilitysoftware-engineer
posted 8w ago · verified today
Harveyai startup
San Francisco · hybrid · $260K–$340K/year · manager
eng-managerml-platformobservabilityreliability-sreinference
posted 8w ago · verified today
Cognitionai startup
San Francisco · onsite · unknown
cpp-langgpu-genericmodel-parallelismpython-langcluster-datacenter
posted 5w ago · verified today
Cursorai startup
New York · San Francisco · onsite · unknown
ml-platformsoftware-engineerscheduling-orchestrationgpu-generickubernetes-ops
posted 16d ago · verified today
Cursorai startup
New York · San Francisco · onsite · unknown
inferencesoftware-engineerinference-enginesml-platformreliability-sre
posted 5mo ago · verified today
Cursorai startup
New York · San Francisco · onsite · unknown
go-langkubernetes-opsnvidiapython-langsoftware-engineer
posted 7mo ago · verified today
Maven Roboticsai startup
San Francisco Bay Area · San Francisco Bay Area, California USA · unknown
ml-platformgpu-generickubernetes-opsray-distributedscheduling-orchestration
posted 12w ago · verified today
Tenstorrentchip vendor
Tokyo · Tokyo, Japan · unknown
cluster-datacenterdatacenter-engineergpu-genericreliability-srenetwork-fabric
posted 12mo ago · verified today
Generalistai startup
San Francisco Bay Area (San Mateo) or Boston (Somerville) · onsite · $260K–$350K/year · unknown
gpu-genericnvidiasoftware-engineerinferencekubernetes-ops
posted 7mo ago · verified today
Lightning AIai startup
New York, New York · New York, New York, United States +5 more · remote · $170K–$210K · senior
infiniband-opsnetwork-engineernetwork-fabricnvidiaansible
posted 8w ago · verified today
Lightning AIai startup
London, England, United Kingdom · London, UK +7 more · remote · $180K–$220K · senior
cluster-datacenterpython-langsoftware-engineergpu-genericobservability
posted 16d ago · verified today
Lightning AIai startup
New York, New York · New York, New York, United States +4 more · $180K–$250K · senior
kubernetes-opsml-platformgo-langpython-langscheduling-orchestration
posted 7w ago · verified today
Lightning AIai startup
Singapore · 165K–205K SGD · senior
reliability-sresreansiblecluster-datacenterkubernetes-ops
posted 7w ago · verified today
Lightning AIai startup
London, England, United Kingdom · London, UK +7 more · remote · $160K–$200K · senior
reliability-sresreansiblecluster-datacenterkubernetes-ops
posted 3mo ago · verified today
Lightning AIai startup
London, England, United Kingdom · London, UK +7 more · remote · $180K–$220K · unknown
cluster-datacenterdatacenter-engineergpu-genericnvidiapython-lang
posted 4mo ago · verified today
Lightning AIai startup
Puyallup, Washington, United States · SEA1 (Puyallup) · $80K–$90K · mid
cluster-datacenterdatacenter-engineerinfiniband-opsnetwork-fabricreliability-sre
posted 7w ago · verified today
Lightning AIai startup
Lisle, Illinois, United States · ORD1 (Lisle) · $80K–$90K · unknown
cluster-datacenterdatacenter-engineerinfiniband-opsnetwork-fabricnvidia
posted 9w ago · verified today
Lightning AIai startup
New York, New York, United States · San Francisco, California +3 more · $115K–$140K · unknown
cudagpu-generickubernetes-opsnccl-libobservability
posted 3mo ago · verified today