AI infrastructure roles, filterable by the stack you actually work on.

Distributed training · inference serving · GPU fleets · network fabric — aggregated straight from company boards, never a copy of a copy.

1297 open roles · 85 companies · last verified today

382 roles

inference

Amazon (AWS)hyperscaler

Cupertino, California, USA · Seattle, Washington, USA · onsite · mid

trainiumperformance-engineersoftware-engineercollectivesgpu-generic

posted 7mo ago · verified today

Amazon (AWS)hyperscaler

Cupertino, California, USA · Seattle, Washington, USA · onsite · mid

trainiummodel-parallelismtraining-frameworkscollectivesperformance-engineer

posted 5w ago · verified today

Amazon (AWS)hyperscaler

Tel Aviv-Yafo, Tel Aviv, ISR · onsite · senior

inferenceinference-enginesvllm-enginecudagpu-generic

posted 10w ago · verified today

Amazon (AWS)hyperscaler

Cupertino, California, USA · onsite · mid

gpu-kernelsinferenceinference-enginesdistributed-inferenceml-platform

posted 6w ago · verified today

Amazon (AWS)hyperscaler

Tel Aviv-Yafo, Tel Aviv, ISR · onsite · mid

cpp-langsoftware-engineercustom-asicdistributed-inferencegpu-kernels

posted 5w ago · verified today

Amazon (AWS)hyperscaler

Seattle, Washington, USA · onsite · staff plus

solutions-architectcudagpu-generickubernetes-opscluster-datacenter

posted 10w ago · verified today

Amazon (AWS)hyperscaler

Palo Alto, California, USA · onsite · senior

software-engineergpu-genericinferenceinference-engines

posted 4mo ago · verified today

Amazon (AWS)hyperscaler

Santa Clara, California, USA · onsite · unknown

cudagpu-kernelstriton-langgpu-genericmodel-parallelism

posted 4w ago · verified today

Amazon (AWS)hyperscaler

Houston, Texas, USA · onsite · senior

cudagpu-kernelsmodel-parallelismpytorch-disttraining-frameworks

posted 4w ago · verified today

Amazon (AWS)hyperscaler

Boston, Massachusetts, USA · Seattle, Washington, USA +1 more · onsite · senior

cudacutlass-cutedistributed-inferenceevaluationflash-attention

posted 4w ago · verified today

Perplexityai startup

Belgrade · Berlin +1 more · hybrid · senior

software-engineerinferenceevaluation

posted 18mo ago · verified today

Perplexityai startup

Palo Alto · San Francisco · $220K–$405K/year · unknown

cluster-datacentercpp-langgpu-genericinferencekubernetes-ops

posted 5mo ago · verified today

Perplexityai startup

New York City · Palo Alto +1 more · $220K–$485K/year · mid

cudagpu-kernelsinferenceinference-enginescutlass-cute

posted 5mo ago · verified today

Cerebraschip vendor

Sunnyvale, CA · Toronto, CAN · hybrid · staff plus

amdinference-enginespython-langsoftware-engineervllm-engine

posted 7w ago · verified today

Cerebraschip vendor

Sunnyvale, CA · hybrid · manager

cpp-langeng-managerinferencemlir-llvmpython-lang

posted 7w ago · verified today

Cerebraschip vendor

Sunnyvale, CA · Toronto, CAN · hybrid · unknown

custom-asicperformance-engineercpp-langinferencepython-lang

posted 9w ago · verified today

Cerebraschip vendor

Sunnyvale, CA · Toronto, CAN · hybrid · senior

inferenceinference-engineskubernetes-opsml-platformpython-lang

posted 9w ago · verified today

Cerebraschip vendor

Sunnyvale, CA · hybrid · staff plus

inferencereliability-sresreobservabilitycluster-datacenter

posted 10w ago · verified today

Cerebraschip vendor

Sunnyvale, CA · Toronto, CAN · onsite · mid

cpp-langgo-langinferencekubernetes-opsml-platform

posted 12w ago · verified today