Senior Software Engineer - Capacity

Snowflake - US-WA-Bellevue - original posting ->
Status
Open
Remote policy
Hybrid
Employment type
Full-time
Salary
200,000-287,500 USD / year
Categories
Engineering
Tech
awsazuregcpsnowflakeprometheusgojavapython
Source
snowflake
First observed
2026-08-21 19:53 UTC
Last seen
2026-08-21 19:53 UTC
Source claims posted
2026-08-21 18:01 UTC
Consecutive misses
0 of 3

What the posting says

At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done.

Senior Software Engineer, Capacity Engineering

At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic, fast-moving environments and approach challenges with an experimental mindset, rapidly testing emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done.

Snowflake’s infrastructure is expanding rapidly across AWS, Azure, and GCP. The Capacity team plays a pivotal role in provisioning the cloud resources essential for Snowflake's operations and ongoing growth. Capacity Engineering accurately models demand, forecasts requirements, and delivers optimal CPU and GPU capacity on schedule. We drive hardware cost-efficiency and price/performance while continually maximizing fleet utilization. To achieve this across all major cloud providers, the team is building a centralized, self-serve internal capacity platform. This software-driven system provides early visibility into supply risks, ensures sufficient lead time for capacity deployment, and maintains high utilization across committed cloud resources.

The technical problem spans the full lifecycle. We model demand and supply as first-class data, reconcile heterogeneous provider telemetry and commitments into a single canonical capacity layer, and make the live state of the fleet legible and actionable in real time. That includes CPU and GPU procurement and reservation lifecycle for AI/ML workloads (training, fine-tuning, and model serving), demand forecasting, cloud resource and cost optimization, hardware evolution analysis as new generations become available (price/performance, cross-family flexibility, migration paths), and the allocation and efficiency systems that close utilization gaps with the teams that own those workloads.

We are actively looking for a senior software engineer. If you love solving problems at scale, prefer to write scalable, reliable, and testable software, are an ace troubleshooter, and are deeply technical, then this is the role for you! Snowflake’s growth and multi-cloud footprint in a constrained capacity environment demand real engineering maturity in the systems that plan and land compute. While the domains below describe the shape of our current goals, the engineer will drive the strategy and deliverables for clear company impact.

AS A SENIOR SOFTWARE ENGINEER AT SNOWFLAKE YOU WILL:

Design and build the capacity platform that unifies CPU and GPU allocation, procurement, reservation lifecycle, and utilization across all three clouds.

Own the canonical capacity data layer: ingest and reconcile demand forecasts, provider supply signals, commitments, and fleet utilization into a single, trustworthy model consumed across the company.

Serve as the liaison to Cloud Service Providers managing and integrating vendor relationships into the capacity planning and procurement workflow.

Build planning and allocation systems that translate demand into hardware requirements (shape, quantity, region, timing) and surface supply risk early, with real-time visibility into fleet and reservation health.

Drive efficiency: instrument utilization across CPU and GPU accelerator workloads, establish price/performance baselines, and build the tooling that recovers stranded capacity and right-sizes commitments.

Integrate hardware evolution into the platform: evaluate new CPU and GPU generations and their price/performance, and build the flexibility (backup and cross-family fallbacks) that keeps plans aligned to the hardware roadmap.

Partner with core services, warehouse, AI/ML, and finance teams to forecast and procure capacity ahead of launches, support AI/ML workloads reliably, and turn insights into procurement and allocation decisions.

Ensure high availability, reliability, and performance of capacity systems by participating in on-call rotations and incident management.

WHAT WE LOOK FOR:

7+ years of industry experience designing, building, and supporting large-scale systems in production.

Hands-on experience working with cloud providers on compute cluster and cloud services provisioning (CPU and/or GPU fleets).

Experience with capacity planning, procurement, resource management, or efficiency work on systems built on large private clouds or public cloud providers.

Deep system and architectural analysis experience to identify actionable performance, availability, and efficiency insights across CPU and GPU accelerator fleets.

Proficiency in programming languages such as Go, Python, or Java.

Excellent problem-solving skills and ability to troubleshoot complex issues in a production environment.

Strong communication skills and the ability to collaborate effectively in a team environment.

BS / MS in Computer Science, Engineering or related fields.

Experience developing or using observability infrastructure such as OpenTelemetry or Prometheus is a plus.

Familiarity with accelerator/GPU fleets, hardware price/performance analysis, or Kubernetes-based compute at scale is a plus.

Prior background working with Modeling, Forecasting, Cloud Spend Optimization, and LLMs is a plus.

Snowflake is growing fast, and we’re scaling our team to help enable and accelerate our growth. We are looking for people who share our values, challenge ordinary thinking, and push the pace of innovation while building a future for themselves and Snowflake.

Snowflake is growing fast, and we’re scaling our team to help enable and accelerate our growth. We are looking for people who share our values, challenge ordinary thinking, and push the pace of innovation while building a future for themselves and Snowflake.

How do you want to make your impact?

For jobs located in the United States, please visit the job posting on the Snowflake Careers Site for salary and benefits information: careers.snowflake.com

Quality

Completeness: 100%

Not enough history yet to judge honesty signals.

Timeline

  1. *
    #255156 2026-08-21 19:53 UTC
    Published