ai-infra-jobs

← all roles

Senior Software Engineer, Datacenter Operations Platform Engineering

Crusoeneocloud

San Francisco, CA - US · $170K–$205K est. · senior · FullTime

Apply at Crusoeoriginal posting on jobs.ashbyhq.com

posted today · first seen today · verified today

Hardwaregpu-generic

Infra layercluster-datacenterreliability-srenetwork-fabricscheduling-orchestration

Role typesoftware-engineer

Skills & stackkubernetes-opsterraform-iacansiblego-langrust-langcpp-langobservabilitypython-lang

Crusoe is on a mission to accelerate the abundance of energy and intelligence. As the only vertically integrated AI infrastructure company built from the ground up, we own and operate each layer of the stack — from electrons to tokens — to power the world's most ambitious AI workloads. When you join Crusoe, you join a team that is building the future, faster.

We're in the midst of the greatest industrial revolution of our time. The demand for AI compute is boundless, and power is a bottleneck. We're solving that — with an energy-first approach that makes AI infrastructure better for the world and faster for the people innovating with AI.

We're looking for problem-solving, opportunity-finding teammates with a sense of urgency, who believe in the scale of our ambition and thrive on a path not fully paved — people who want to grow their careers alongside a team of experts across energy, manufacturing, data center construction, and cloud services.

If you want to do the most meaningful work of your career, help our customers and partners advance their AI strategies, and be part of a high-performing team that believes in each other, come build with us at Crusoe.

About the Role:
We are seeking Senior Software Engineers to design and develop internal datacenter tooling and infrastructure management systems for Crusoe Cloud, a leading cloud provider. You will play a crucial role in building tools for datacenter facilities managers and capacity planners, as well as creating automation software to efficiently bring server hardware, switches, and other infrastructure components online.

You’ll help evaluate and implement tools and frameworks for our internal customers, focusing on reliability, scalability, operational efficiency, and ease of use. This role is central to streamlining infrastructure management processes and enhancing our cloud platform’s overall performance as we dramatically scale our hardware footprint.

Blog Posts About This Team:

What You’ll Be Working On:

  • Designing and developing advanced internal tooling for datacenter facilities managers and capacity planners

  • Creating automation software for rapid deployment and configuration of servers, network switches, power delivery units (PDUs), and coolant delivery units (CDUs)

  • Mentoring junior engineers through design guidance and code reviews to ensure high-quality solutions

  • Innovating and implementing features that streamline infrastructure management and operational capabilities

  • Collaborating with cloud support and operations teams to develop tools that enable growth and empower internal processes

  • Partnering cross-functionally to align goals and optimize resource utilization for improved infrastructure management

  • Leading by example in technical excellence and fostering an environment of innovation in infrastructure and tooling development

What You’ll Bring to the Team:

  • 5+ years of professional software development experience

  • 5+ years of programming experience in at least one modern compiled language (Go, Rust, Java, or C++)

  • 5+ years of experience contributing to architecture and design (patterns, reliability, scaling) of new and existing systems

  • Bachelor’s degree in Computer Science or related field, or 5–8+ years of equivalent experience

  • Strong computer science fundamentals in data structures and algorithms

  • Proven experience building and maintaining scalable, highly available, fault-tolerant distributed systems

  • Solid understanding of infrastructure design and operational trade-offs

  • Familiarity with CI/CD practices and build systems (GitLab CI/CD, CircleCI, GitHub Actions)

  • Familiarity with modern infrastructure tools (Docker, Kubernetes, Ansible, CloudFormation, Terraform)

  • Experience with concurrency, multithreading, and synchronization

  • Experience with Unix/Linux environments

  • Experience with TCP/IP and network programming

  • Excellent communication skills

  • Alignment with company values

Bonus Points:

  • Experience working in large-scale datacenter or cloud environments

  • Hands-on exposure to GPU clusters or high-performance computing environments

  • Background in infrastructure observability or monitoring systems

Benefits:

  • Industry competitive pay

  • Restricted Stock Units in a fast growing, well-funded technology company

  • Health insurance package options that include HDHP and PPO, vision, and dental for you and your dependents

  • Employer contributions to HSA accounts

  • Paid Parental Leave

  • Paid life insurance, short-term and long-term disability

  • Teladoc

  • 401(k) with a 100% match up to 4% of salary

  • Generous paid time off and holiday schedule

  • Cell phone reimbursement

  • Tuition reimbursement

  • Subscription to the Calm app

  • MetLife Legal

  • Company paid commuter benefit; $300 per month

Compensation:

Compensation will be paid in the range of $170,000 - $205,000 + Bonus. Restricted Stock Units are included in all offers. Compensation to be determined by the applicant’s education, experience, knowledge, skills, and abilities, as well as internal equity and alignment with market data.

Crusoe is an Equal Opportunity Employer. Employment decisions are made without regard to race, color, religion, disability, genetic information, pregnancy, citizenship, marital status, sex/gender, sexual preference/ orientation, gender identity, age, veteran status, national origin, or any other status protected by law or regulation.