AI infrastructure roles, filterable by the stack you actually work on.

Distributed training · inference serving · GPU fleets · network fabric — aggregated straight from company boards, never a copy of a copy.

1297 open roles · 85 companies · last verified today

TensorWaveneocloud

Las Vegas, Nevada · onsite · mid

amdgpu-generickubernetes-opsreliability-srescheduling-orchestration

posted 3w ago · verified today

TensorWaveneocloud

Las Vegas, Nevada · Remote · onsite · senior

gpu-generickubernetes-opsml-platformpython-langscheduling-orchestration

posted 10mo ago · verified today

Periodic Labsfrontier lab

Menlo Park, CA · onsite · unknown

collectivescudacutlass-cutefsdpgpu-generic

posted 4mo ago · verified today

Amazon (AWS)hyperscaler

Vancouver, British Columbia, CAN · onsite · senior

ml-platformsoftware-engineerevaluationinferencekubernetes-ops

posted 14d ago · verified today

Mistral AIfrontier lab

Palo Alto · hybrid · unknown

ml-platformscheduling-orchestrationkubernetes-opspython-langcuda

posted 13d ago · verified today

Nscaleneocloud

London · UK · senior

evaluationfine-tuninggpu-genericinferenceinference-engines

posted 5mo ago · verified today

Anthropicfrontier lab

New York City, NY · San Francisco, CA +1 more · $405K–$625K · manager

eng-managerinferencescheduling-orchestrationinference-enginescluster-datacenter

posted 15d ago · verified today

SambaNovachip vendor

Stockholm, Sweden · staff plus

cpp-langsoftware-engineerinferencelinux-kernelnetwork-fabric

posted 19d ago · verified today

Amazon (AWS)hyperscaler

Austin, Texas, USA · Cupertino, California, USA +2 more · onsite · junior

trainiumcpp-langpython-langsoftware-engineerdistributed-inference

posted 3w ago · verified today

Amazon (AWS)hyperscaler

Boston, Massachusetts, USA · Seattle, Washington, USA +1 more · onsite · senior

gpu-kernelsinferenceinference-engineskv-cache-systemsperformance-engineer

posted 3w ago · verified today

Nebiusneocloud

Abu Dhabi, UAE · Middle East +1 more · remote · senior

solutions-architectgpu-generickubernetes-opsml-platforminference

posted 9w ago · verified today

Amazon (AWS)hyperscaler

New York, New York, USA · Sunnyvale, California, USA · onsite · senior

kubernetes-opsml-platforminferenceinference-enginesquantization

posted 3w ago · verified today

Crusoeneocloud

Denver, CO - US · onsite · staff plus

inferenceinference-enginesperformance-engineersglang-enginevllm-engine

posted 3w ago · verified today

Anthropicfrontier lab

New York City, NY · San Francisco, CA +1 more · $320K–$485K · staff plus

inferenceinference-enginesscheduling-orchestrationsoftware-engineerdistributed-inference

posted 3w ago · verified today

OpenAIfrontier lab

San Francisco · hybrid · $266K–$445K/year · unknown

software-engineerinference-enginesinferencekv-cache-systemsscheduling-orchestration

posted 3w ago · verified today

SambaNovachip vendor

Bengaluru, India · Bengaluru, Karnataka, India · 10300K–12500K INR · staff plus

inference-enginessoftware-engineergo-langinferencekubernetes-ops

posted 4w ago · verified today

CoreWeaveneocloud

Bellevue, WA · San Francisco, CA +1 more · $198K–$264K/year est. · staff plus

amdcluster-datacenterdeepspeed-libdistributed-inferencefsdp

posted 4w ago · verified today

Amazon (AWS)hyperscaler

Seattle, Washington, USA · onsite · manager

eng-managernvidiacluster-datacentercpp-langml-platform

posted 5w ago · verified today

Amazon (AWS)hyperscaler

New York, New York, USA · Seattle, Washington, USA · onsite · senior

software-engineerinferenceinference-enginesdistributed-inferencetensorrt-stack

posted 6w ago · verified today