AI infrastructure roles, filterable by the stack you actually work on.

Distributed training · inference serving · GPU fleets · network fabric — aggregated straight from company boards, never a copy of a copy.

1294 open roles · 85 companies · last verified today

453 roles

reliability-sre

Thinking Machines Labfrontier lab

New York · San Francisco · onsite · unknown

reliability-sresrefine-tuninggpu-generickubernetes-ops

posted 16d ago · verified today

Thinking Machines Labfrontier lab

New York · San Francisco · onsite · unknown

post-trainingreliability-sresrepython-langreinforcement-learning

posted 16d ago · verified today

Cerebraschip vendor

Sunnyvale, CA · Toronto, CAN · hybrid · unknown

python-langsoftware-engineercluster-datacentercpp-langreliability-sre

posted 13d ago · verified today

Amazon (AWS)hyperscaler

Vancouver, British Columbia, CAN · onsite · senior

ml-platformsoftware-engineerevaluationinferencekubernetes-ops

posted 14d ago · verified today

Mistral AIfrontier lab

Palo Alto · hybrid · unknown

ml-platformscheduling-orchestrationkubernetes-opspython-langcuda

posted 13d ago · verified today

CoreWeaveneocloud

Bellevue, WA · Livingston, NJ +1 more · $153K–$204K/year est. · senior

gpu-generickubernetes-opssoftware-engineercluster-datacentergo-lang

posted 14d ago · verified today

CoreWeaveneocloud

Bellevue, WA · Manhattan, NY +2 more · $157K–$210K/year est. · senior

cluster-datacenterreliability-sreobservability

posted 14d ago · verified today

xAIfrontier lab

Memphis, TN · Southaven, MS · senior

reliability-sresrecluster-datacenterobservabilitypython-lang

posted 14d ago · verified today

Anthropicfrontier lab

London, UK · £325K–£390K · staff plus

ebpfobservabilitygpu-genericsoftware-engineertpu

posted 14d ago · verified today

CoreWeaveneocloud

Quincy, WA · Quincy, WA - DC · $65K–$83K/year est. · unknown

datacenter-engineercluster-datacenterpython-langnetwork-fabricreliability-sre

posted 14d ago · verified today

Nscaleneocloud

Seattle · US · $220K–$320K · staff plus

cluster-datacenterpython-langsoftware-engineergpu-genericobservability

posted 5w ago · verified today

Nscaleneocloud

US · $190K–$260K · staff plus

observabilitygo-langgpu-generickubernetes-opspython-lang

posted 3mo ago · verified today

Nscaleneocloud

Houston · New York +2 more · staff plus

nvidiasoftware-engineercluster-datacentercudago-lang

posted 4w ago · verified today

Nscaleneocloud

New York · $225K–$275K · staff plus

scheduling-orchestrationslurm-adminsoftware-engineergo-langpython-lang

posted 3mo ago · verified today

Nscaleneocloud

US · $160K–$230K · senior

observabilitygo-langkubernetes-opspython-langansible

posted 3w ago · verified today

Nscaleneocloud

US · $150K–$210K · senior

ansibleinfiniband-opsnetwork-engineernetwork-fabricpython-lang

posted 5mo ago · verified today

Nscaleneocloud

Houston · San Francisco +1 more · $120K–$170K · senior

nvidiasrecluster-datacenterinfiniband-opsnccl-lib

posted 6mo ago · verified today

Nscaleneocloud

US · $190K–$300K · staff plus

observabilitykubernetes-opsgo-langpython-langreliability-sre

posted 5mo ago · verified today

Nscaleneocloud

AMER · Seattle · $270K–$330K · staff plus

infiniband-opsnetwork-engineernetwork-fabricansiblepython-lang

posted 4w ago · verified today