AI infrastructure roles, filterable by the stack you actually work on.

Distributed training · inference serving · GPU fleets · network fabric — aggregated straight from company boards, never a copy of a copy.

1296 open roles · 85 companies · last verified today

100 roles

sglang-engine

Nebiusneocloud

Palo Alto, California, United States · San Francisco Bay Area · $195K–$262K · senior

inferenceinference-enginespython-langquantizationvllm-engine

posted 8w ago · verified today

Nebiusneocloud

Palo Alto, California, United States · San Francisco Bay Area · $195K–$262K · senior

research-engineercudainferencepython-langtriton-lang

posted 8w ago · verified today

Nebiusneocloud

United States · $208K–$261K · staff plus

solutions-architectfine-tuninginferenceinference-engineskv-cache-systems

posted 11w ago · verified today

Nebiusneocloud

Amsterdam, Netherlands · Czech Republic +5 more · remote · unknown

cudagpu-genericgpu-kernelsnccl-libperformance-engineer

posted 4mo ago · verified today

Nebiusneocloud

Remote - United States · remote · $180K–$224K · senior

python-langsglang-enginesolutions-architecttensorrt-stackvllm-engine

posted 4mo ago · verified today

Crusoeneocloud

San Francisco, CA - US · Sunnyvale, CA - US · onsite · staff plus

cpp-langcudadistributed-inferencegpu-genericgpu-kernels

posted 8w ago · verified today

Together AIneocloud

San Francisco · $220K–$280K/year est. · staff plus

inferenceinference-enginespython-langcudaevaluation

posted 4mo ago · verified today

Together AIneocloud

San Francisco · $200K–$290K/year est. · senior

inferenceinference-enginesnvidiasoftware-engineerdistributed-inference

posted 13mo ago · verified today

Together AIneocloud

San Francisco · $200K–$290K/year est. · unknown

research-engineerfine-tuninggo-langinferenceinference-engines

posted 10w ago · verified today