Senior Production Engineer

Status
Open
Remote policy
Not stated
Employment type
Not stated
Salary
Not stated
Categories
Infrastructure
Tech
awskafkapostgresredissnowflakekubernetesterraformgojavapythonsoftwaresenior
Source
arbeitnow
First observed
2026-09-13 13:26 UTC
Last seen
2026-09-13 13:26 UTC
Source claims posted
2026-09-13 12:55 UTC
Consecutive misses
0 of 3

What the posting says

<p><span style="font-size: 12pt;"><strong>About Clear Street:<br></strong></span></p> <div><span style="font-size: 12pt;">Clear Street’s mission is to give every sophisticated investor access to every asset, in every market, through a unified platform built for speed, transparency and scale.</span></div> <div><br><span style="font-size: 12pt;">We give our clients the technology, tools, and service once reserved for the largest institutions, rebuilt with modern infrastructure. Our single, cloud-native, end-to-end capital markets platform powers investor growth today and is transforming how they can interact with markets tomorrow.</span><br><br></div> <div><span style="font-size: 12pt;">For more information, visit&nbsp;<a href="https://clearstreet.io/" target="_blank" data-saferedirecturl="https://www.google.com/url?q=https://clearstreet.io&amp;source=gmail&amp;ust=1766166763873000&amp;usg=AOvVaw3GfALVHjsW8XaQSab-3CbX">https://clearstreet.io</a>.</span></div> <p><strong><span style="font-size: 12pt;">The Role</span></strong></p> <p><br><span style="font-size: 12pt;">As a Production Engineer, you sit at the intersection of software reliability and operational </span><span style="font-size: 12pt;">excellence. You own the health, resilience, and recovery of our production systems—while </span><span style="font-size: 12pt;">spending equal energy innovating solutions that eliminate human toil, reduce incident blast </span><span style="font-size: 12pt;">radius, and raise the reliability bar across the entire platform.&nbsp; </span><span style="font-size: 12pt;">You will partner closely with engineering, operations, and business teams to understand daily </span><span style="font-size: 12pt;">pain points and translate them into lasting automated solutions. Half your time is spent in the </span><span style="font-size: 12pt;">trenches—supporting production, responding to incidents, and deeply understanding how our</span><br><span style="font-size: 12pt;">systems behave under real conditions. The other half is yours to build: automation, tooling, and </span><span style="font-size: 12pt;">observability platforms that make tomorrow&amp;#39;s on-call shift meaningfully easier than today's.&nbsp;&nbsp;</span></p> <p><br><strong><span style="font-size: 12pt;">You will work on challenges like:</span></strong></p> <p><br><span style="font-size: 12pt;">● Design and build comprehensive monitoring and observability platforms that surface the </span><span style="font-size: 12pt;">right signal at the right time—eliminating alert fatigue and accelerating root-cause </span><span style="font-size: 12pt;">analysis.</span><br><span style="font-size: 12pt;">● Develop intelligent automation and self-healing capabilities that diagnose issues, trigger </span><span style="font-size: 12pt;">recovery workflows, and reduce mean time to recovery (MTTR) without manual&nbsp;</span><span style="font-size: 12pt;">intervention.</span><br><span style="font-size: 12pt;">● Analyze incidents, identify systemic trends, and engineer solutions that prevent entire </span><span style="font-size: 12pt;">classes of failures from recurring.</span><br><span style="font-size: 12pt;">● Build reusable runbooks, diagnostic tooling, and recovery playbooks that turn tribal </span><span style="font-size: 12pt;">knowledge into scalable platform capabilities.</span><br><span style="font-size: 12pt;">● Create golden-path operational workflows—making the safest, most reliable path also </span><span style="font-size: 12pt;">the easiest one for engineering teams to follow.</span><br><span style="font-size: 12pt;">● Partner with Platform Engineering to influence CI/CD pipelines, deployment safety, and </span><span style="font-size: 12pt;">infrastructure resilience from a production reliability perspective.</span><br><span style="font-size: 12pt;">● Champion Infrastructure as Code, GitOps, and SRE best practices while helping teams </span><span style="font-size: 12pt;">adopt modern engineering workflows.</span><br><span style="font-size: 12pt;">● Continuously measure production health through SLIs/SLOs/SLAs, and drive </span><span style="font-size: 12pt;">engineering priorities based on reliability data.</span><br><span style="font-size: 12pt;">● Explore emerging technologies—including AI-assisted diagnostics and developer </span><span style="font-size: 12pt;">tooling—that transform how we operate production systems.</span></p> <p><strong><span style="font-size: 12pt;">The Team</span></strong></p> <p><br><span style="font-size: 12pt;">We believe resilient systems are built by engineers who understand them end to end.&nbsp; </span><span style="font-size: 12pt;">Our Production Engineering team is the first and last line of defense for our production platform.&nbsp; </span><span style="font-size: 12pt;">We treat reliability as a product, with uptime and engineer experience as our north stars. We </span><span style="font-size: 12pt;">combine the discipline of SRE with a builder's mindset: when we see a recurring problem, we </span><span style="font-size: 12pt;">build a solution—not a workaround.</span></p> <p><br><span style="font-size: 12pt;">You will work across every engineering and operations team to understand failure modes, quantify </span><span style="font-size: 12pt;">reliability gaps, and build platform capabilities that scale with the organization. Whether it's </span><span style="font-size: 12pt;">reducing MTTR from hours to minutes, building self-service diagnostic tools, or designing </span><span style="font-size: 12pt;">proactive alerting that catches issues before customers notice, your work will have immediate, </span><span style="font-size: 12pt;">measurable impact.</span></p> <p><br><span style="font-size: 12pt;">If you're passionate about making production systems invisible to end users—and you get </span><span style="font-size: 12pt;">energy from both firefighting and building the systems that make fires less likely—you'll thrive </span><span style="font-size: 12pt;">here.</span></p> <p><strong><span style="font-size: 12pt;">What We're Looking For</span></strong></p> <p><br><span style="font-size: 12pt;">We're looking for engineers who combine operational instinct with a builder's discipline.</span></p> <p><br><span style="font-size: 12pt;">You should have:</span><br><span style="font-size: 12pt;">● Strong hands-on Python skills—this is your primary language for automation and tooling.</span><br><span style="font-size: 12pt;">● Experience in SRE, Production Engineering, Platform Engineering, or a related discipline </span><span style="font-size: 12pt;">with direct production ownership.</span><br><span style="font-size: 12pt;">● Proven track record of building automation and diagnostic tooling that improved recovery </span><span style="font-size: 12pt;">times or reduced operational toil.</span><br><span style="font-size: 12pt;">● Deep familiarity with cloud-native technologies—Kubernetes, containers, distributed </span><span style="font-size: 12pt;">systems—and how they fail in production.</span><br><span style="font-size: 12pt;">● Experience with observability platforms such as Datadog, and a strong intuition for what "</span><span style="font-size: 12pt;">good" monitoring looks like.</span><br><span style="font-size: 12pt;">● Exposure to Infrastructure as Code (Terraform) and GitOps-based deployment </span><span style="font-size: 12pt;">workflows (ArgoCD, GitHub Actions, or similar).</span><br><span style="font-size: 12pt;">● Familiarity with the broader technology stack: Java, Go, Kafka, Redis, Snowflake, and </span><span style="font-size: 12pt;">Postgres.</span><br><span style="font-size: 12pt;">● Strong analytical and problem-solving skills—you thrive on ambiguous, high-stakes </span><span style="font-size: 12pt;">production problems.</span><br><span style="font-size: 12pt;">● A product mindset applied to operational tooling: you think about usability, adoption, and </span><span style="font-size: 12pt;">documentation when building internal solutions.</span><br><span style="font-size: 12pt;">● Excellent communication skills and the ability to work fluidly across engineering, </span><span style="font-size: 12pt;">operations, and business stakeholders.</span><br><span style="font-size: 12pt;">● Self-starter mentality—you identify opportunities, take initiative, and deliver with minimal </span><span style="font-size: 12pt;">supervision.</span><br><span style="font-size: 12pt;">● Curiosity and a continuous learning mindset; fintech or financial industry background is a </span><span style="font-size: 12pt;">plus.</span></p> <p><strong><span style="font-size: 12pt;">The Technology You'll Work With</span></strong></p> <p><br><span style="font-size: 12pt;">You'll operate and build on a modern cloud-native platform that includes:</span></p> <p><span style="font-size: 12pt;">● Kubernetes &amp; AWS</span><br><span style="font-size: 12pt;">● Terraform &amp; ArgoCD</span><br><span style="font-size: 12pt;">● GitHub Actions</span><br><span style="font-size: 12pt;">● Kafka, Redis</span><br><span style="font-size: 12pt;">● PostgreSQL &amp; Snowflake</span><br><span style="font-size: 12pt;">● Datadog</span><br><span style="font-size: 12pt;">● Python, Go, Java</span><br><span style="font-size: 12pt;">● gRPC &amp; Protobuf</span><br><span style="font-size: 12pt;">● Internal Platform APIs and Developer Tooling</span></p> <p><strong><span style="font-size: 12pt;">What Success Looks Like</span></strong></p> <p><br><span style="font-size: 12pt;">Within your first year, you'll have made a measurable impact on production reliability. Success l</span><span style="font-size: 12pt;">ooks like:</span></p> <p><br><span style="font-size: 12pt;">● Reducing mean time to detection (MTTD) and mean time to recovery (MTTR) across key </span><span style="font-size: 12pt;">production systems.</span><br><span style="font-size: 12pt;">● Building automation that handles a meaningful percentage of incident scenarios without </span><span style="font-size: 12pt;">human intervention.</span><br><span style="font-size: 12pt;">● Becoming a trusted subject matter expert for core platform components and their failure </span><span style="font-size: 12pt;">modes.</span><br><span style="font-size: 12pt;">● Delivering observability and diagnostic tools that other engineers actually use and </span><span style="font-size: 12pt;">depend on.</span><br><span style="font-size: 12pt;">● Establishing SLO baselines and driving engineering investment based on reliability data.</span><br><span style="font-size: 12pt;">● Spending less of your time—and your teammates time—on repetitive manual toil.</span></p> <p><br><span style="font-size: 12pt;">Your impact won't be measured by the number of incidents you respond to—it will be measured </span><span style="font-size: 12pt;">by how reliably our systems run and how quickly we recover when they don't.</span></p>

Find Jobs in United Kingdom on Arbeitnow

Quality

Completeness: 45%

Not enough history yet to judge honesty signals.

Timeline

  1. *
    #729775 2026-09-13 13:26 UTC
    Published