AI infrastructure roles, filterable by the stack you actually work on.
Distributed training · inference serving · GPU fleets · network fabric — aggregated straight from company boards, never a copy of a copy.
1296 open roles · 85 companies · last verified today
685 roles
Fish Audio (39 AI)ai startup
Location not specified · senior
cluster-datacentergpu-genericreliability-sresrekubernetes-ops
added 7d ago · verified today
Waferinference provider
San Francisco · onsite · $200K–$300K/year · unknown
gpu-kernelsinferenceinference-enginescluster-datacenterperformance-engineer
posted 8w ago · verified today
Relaceinference provider
San Francisco · onsite · mid
software-engineercluster-datacenterscheduling-orchestrationinferenceml-platform
posted 10mo ago · verified today
Exaai startup
Singapore · onsite · 90K–300K SGD/year · unknown
cluster-datacenterkubernetes-opsscheduling-orchestrationsoftware-engineerdistributed-inference
posted 6mo ago · verified today
Exaai startup
San Francisco, California · onsite · $180K–$350K/year · unknown
software-engineergpu-generickubernetes-opsml-platformray-distributed
posted 12mo ago · verified today
Arcee AIfrontier lab
San Francisco, CA · SF, CA · unknown
cluster-datacentergpu-generickubernetes-opsreliability-sresre
posted 3mo ago · verified today
Black Forest Labsfrontier lab
Freiburg (Germany) · onsite · unknown
cluster-datacentergo-langkubernetes-opsnvidiaobservability
posted 12mo ago · verified today
Baseteninference provider
San Francisco · hybrid · $170K–$230K/year · senior
cluster-datacentergpu-generickubernetes-opsnvidiascheduling-orchestration
posted 8d ago · verified today
Fireworks AIinference provider
New York · San Mateo · hybrid · $200K–$230K/year · unknown
scheduling-orchestrationcpp-langnetwork-fabricpython-langstorage-checkpointing
posted 8d ago · verified today
Crusoeneocloud
Dallas, TX - US · onsite · staff plus
datacenter-engineercluster-datacenter
posted 8d ago · verified today
Cerebraschip vendor
Toronto, CAN · hybrid · senior
cluster-datacentergo-langkubernetes-opsscheduling-orchestrationobservability
posted 8d ago · verified today
Nscaleneocloud
UK · manager
eng-manageransiblecluster-datacentergpu-genericobservability
posted 9d ago · verified today
Lila Sciencesai startup
Alewife, Cambridge, MA · Cambridge, MA USA +4 more · $224K–$294K · staff plus
gpu-generickubernetes-opsml-platformpython-langray-distributed
posted 6w ago · verified today
Lila Sciencesai startup
Alewife, Cambridge, MA · Cambridge, MA USA · $192K–$272K · staff plus
inferenceinference-engineskubernetes-opsml-platformnvidia
posted 7w ago · verified today
Prior Labsai startup
Berlin · Freiburg +1 more · onsite · unknown
cluster-datacentergpu-genericpython-langpytorch-distscheduling-orchestration
posted 6w ago · verified today
Genesis Molecular AIai startup
San Mateo, CA · senior
kubernetes-opsml-platformpython-langray-distributedscheduling-orchestration
posted 7mo ago · verified today
Luma AIai startup
Redwood City, CA · hybrid · manager
inferenceinference-enginesml-platformscheduling-orchestrationcluster-datacenter
posted 7w ago · verified today
Luma AIai startup
Redwood City, CA · hybrid · $235K–$353K/year · staff plus
gpu-genericreliability-sresrecluster-datacenterkubernetes-ops
posted 7w ago · verified today
Luma AIai startup
Redwood City, CA · hybrid · unknown
inferenceinference-enginessoftware-engineerkubernetes-opspython-lang
posted 7w ago · verified today
Luma AIai startup
Redwood City, CA · hybrid · $168K–$252K/year · senior
reliability-sresregpu-genericnvidiaamd
posted 7w ago · verified today
Luma AIai startup
Redwood City, CA · hybrid · unknown
pytorch-disttraining-frameworkscollectivescudafsdp
posted 7w ago · verified today
ElevenLabsai startup
United States · remote · unknown
gpu-genericnvidiacluster-datacentercudareliability-sre
posted 13d ago · verified today
Harveyai startup
New York · $161K–$242K/year · senior
cluster-datacenterkubernetes-opsreliability-sreml-platformobservability
posted 6w ago · verified today
Cognitionai startup
San Francisco · onsite · unknown
cpp-langgpu-genericmodel-parallelismpython-langcluster-datacenter
posted 5w ago · verified today
Cursorai startup
New York · San Francisco · onsite · unknown
go-langkubernetes-opsnvidiapython-langsoftware-engineer
posted 7mo ago · verified today
Maven Roboticsai startup
San Francisco Bay Area · San Francisco Bay Area, California USA · unknown
ml-platformgpu-generickubernetes-opsray-distributedscheduling-orchestration
posted 12w ago · verified today
Nimble Roboticsai startup
San Francisco, CA · SFHQ · senior
cudagpu-genericgpu-kernelssoftware-engineercpp-lang
posted 4w ago · verified today
Figureai startup
HQ · San Jose, CA · $200K–$400K/year est. · senior
performance-engineercollectivescpp-langcudafsdp
posted 4w ago · verified today
Tenstorrentchip vendor
Austin · Austin, Texas, United States +2 more · unknown
cluster-datacenterdatacenter-engineer
posted 6mo ago · verified today
Tenstorrentchip vendor
Tokyo · Tokyo, Japan · unknown
cluster-datacenterdatacenter-engineergpu-genericreliability-srenetwork-fabric
posted 12mo ago · verified today