Senior Production Engineer
- Status
- Open
- Remote policy
- Not stated
- Employment type
- Not stated
- Salary
- Not stated
- Categories
- Infrastructure
- Source
- arbeitnow
- First observed
- 2026-09-13 13:26 UTC
- Last seen
- 2026-09-13 13:26 UTC
- Source claims posted
- 2026-09-13 12:55 UTC
- Consecutive misses
- 0 of 3
What the posting says
<p><span style="font-size: 12pt;"><strong>About Clear Street:<br></strong></span></p> <div><span style="font-size: 12pt;">Clear Street’s mission is to give every sophisticated investor access to every asset, in every market, through a unified platform built for speed, transparency and scale.</span></div> <div><br><span style="font-size: 12pt;">We give our clients the technology, tools, and service once reserved for the largest institutions, rebuilt with modern infrastructure. Our single, cloud-native, end-to-end capital markets platform powers investor growth today and is transforming how they can interact with markets tomorrow.</span><br><br></div> <div><span style="font-size: 12pt;">For more information, visit <a href="https://clearstreet.io/" target="_blank" data-saferedirecturl="https://www.google.com/url?q=https://clearstreet.io&source=gmail&ust=1766166763873000&usg=AOvVaw3GfALVHjsW8XaQSab-3CbX">https://clearstreet.io</a>.</span></div> <p><strong><span style="font-size: 12pt;">The Role</span></strong></p> <p><br><span style="font-size: 12pt;">As a Production Engineer, you sit at the intersection of software reliability and operational </span><span style="font-size: 12pt;">excellence. You own the health, resilience, and recovery of our production systems—while </span><span style="font-size: 12pt;">spending equal energy innovating solutions that eliminate human toil, reduce incident blast </span><span style="font-size: 12pt;">radius, and raise the reliability bar across the entire platform. </span><span style="font-size: 12pt;">You will partner closely with engineering, operations, and business teams to understand daily </span><span style="font-size: 12pt;">pain points and translate them into lasting automated solutions. Half your time is spent in the </span><span style="font-size: 12pt;">trenches—supporting production, responding to incidents, and deeply understanding how our</span><br><span style="font-size: 12pt;">systems behave under real conditions. The other half is yours to build: automation, tooling, and </span><span style="font-size: 12pt;">observability platforms that make tomorrow&#39;s on-call shift meaningfully easier than today's. </span></p> <p><br><strong><span style="font-size: 12pt;">You will work on challenges like:</span></strong></p> <p><br><span style="font-size: 12pt;">● Design and build comprehensive monitoring and observability platforms that surface the </span><span style="font-size: 12pt;">right signal at the right time—eliminating alert fatigue and accelerating root-cause </span><span style="font-size: 12pt;">analysis.</span><br><span style="font-size: 12pt;">● Develop intelligent automation and self-healing capabilities that diagnose issues, trigger </span><span style="font-size: 12pt;">recovery workflows, and reduce mean time to recovery (MTTR) without manual </span><span style="font-size: 12pt;">intervention.</span><br><span style="font-size: 12pt;">● Analyze incidents, identify systemic trends, and engineer solutions that prevent entire </span><span style="font-size: 12pt;">classes of failures from recurring.</span><br><span style="font-size: 12pt;">● Build reusable runbooks, diagnostic tooling, and recovery playbooks that turn tribal </span><span style="font-size: 12pt;">knowledge into scalable platform capabilities.</span><br><span style="font-size: 12pt;">● Create golden-path operational workflows—making the safest, most reliable path also </span><span style="font-size: 12pt;">the easiest one for engineering teams to follow.</span><br><span style="font-size: 12pt;">● Partner with Platform Engineering to influence CI/CD pipelines, deployment safety, and </span><span style="font-size: 12pt;">infrastructure resilience from a production reliability perspective.</span><br><span style="font-size: 12pt;">● Champion Infrastructure as Code, GitOps, and SRE best practices while helping teams </span><span style="font-size: 12pt;">adopt modern engineering workflows.</span><br><span style="font-size: 12pt;">● Continuously measure production health through SLIs/SLOs/SLAs, and drive </span><span style="font-size: 12pt;">engineering priorities based on reliability data.</span><br><span style="font-size: 12pt;">● Explore emerging technologies—including AI-assisted diagnostics and developer </span><span style="font-size: 12pt;">tooling—that transform how we operate production systems.</span></p> <p><strong><span style="font-size: 12pt;">The Team</span></strong></p> <p><br><span style="font-size: 12pt;">We believe resilient systems are built by engineers who understand them end to end. </span><span style="font-size: 12pt;">Our Production Engineering team is the first and last line of defense for our production platform. </span><span style="font-size: 12pt;">We treat reliability as a product, with uptime and engineer experience as our north stars. We </span><span style="font-size: 12pt;">combine the discipline of SRE with a builder's mindset: when we see a recurring problem, we </span><span style="font-size: 12pt;">build a solution—not a workaround.</span></p> <p><br><span style="font-size: 12pt;">You will work across every engineering and operations team to understand failure modes, quantify </span><span style="font-size: 12pt;">reliability gaps, and build platform capabilities that scale with the organization. Whether it's </span><span style="font-size: 12pt;">reducing MTTR from hours to minutes, building self-service diagnostic tools, or designing </span><span style="font-size: 12pt;">proactive alerting that catches issues before customers notice, your work will have immediate, </span><span style="font-size: 12pt;">measurable impact.</span></p> <p><br><span style="font-size: 12pt;">If you're passionate about making production systems invisible to end users—and you get </span><span style="font-size: 12pt;">energy from both firefighting and building the systems that make fires less likely—you'll thrive </span><span style="font-size: 12pt;">here.</span></p> <p><strong><span style="font-size: 12pt;">What We're Looking For</span></strong></p> <p><br><span style="font-size: 12pt;">We're looking for engineers who combine operational instinct with a builder's discipline.</span></p> <p><br><span style="font-size: 12pt;">You should have:</span><br><span style="font-size: 12pt;">● Strong hands-on Python skills—this is your primary language for automation and tooling.</span><br><span style="font-size: 12pt;">● Experience in SRE, Production Engineering, Platform Engineering, or a related discipline </span><span style="font-size: 12pt;">with direct production ownership.</span><br><span style="font-size: 12pt;">● Proven track record of building automation and diagnostic tooling that improved recovery </span><span style="font-size: 12pt;">times or reduced operational toil.</span><br><span style="font-size: 12pt;">● Deep familiarity with cloud-native technologies—Kubernetes, containers, distributed </span><span style="font-size: 12pt;">systems—and how they fail in production.</span><br><span style="font-size: 12pt;">● Experience with observability platforms such as Datadog, and a strong intuition for what "</span><span style="font-size: 12pt;">good" monitoring looks like.</span><br><span style="font-size: 12pt;">● Exposure to Infrastructure as Code (Terraform) and GitOps-based deployment </span><span style="font-size: 12pt;">workflows (ArgoCD, GitHub Actions, or similar).</span><br><span style="font-size: 12pt;">● Familiarity with the broader technology stack: Java, Go, Kafka, Redis, Snowflake, and </span><span style="font-size: 12pt;">Postgres.</span><br><span style="font-size: 12pt;">● Strong analytical and problem-solving skills—you thrive on ambiguous, high-stakes </span><span style="font-size: 12pt;">production problems.</span><br><span style="font-size: 12pt;">● A product mindset applied to operational tooling: you think about usability, adoption, and </span><span style="font-size: 12pt;">documentation when building internal solutions.</span><br><span style="font-size: 12pt;">● Excellent communication skills and the ability to work fluidly across engineering, </span><span style="font-size: 12pt;">operations, and business stakeholders.</span><br><span style="font-size: 12pt;">● Self-starter mentality—you identify opportunities, take initiative, and deliver with minimal </span><span style="font-size: 12pt;">supervision.</span><br><span style="font-size: 12pt;">● Curiosity and a continuous learning mindset; fintech or financial industry background is a </span><span style="font-size: 12pt;">plus.</span></p> <p><strong><span style="font-size: 12pt;">The Technology You'll Work With</span></strong></p> <p><br><span style="font-size: 12pt;">You'll operate and build on a modern cloud-native platform that includes:</span></p> <p><span style="font-size: 12pt;">● Kubernetes & AWS</span><br><span style="font-size: 12pt;">● Terraform & ArgoCD</span><br><span style="font-size: 12pt;">● GitHub Actions</span><br><span style="font-size: 12pt;">● Kafka, Redis</span><br><span style="font-size: 12pt;">● PostgreSQL & Snowflake</span><br><span style="font-size: 12pt;">● Datadog</span><br><span style="font-size: 12pt;">● Python, Go, Java</span><br><span style="font-size: 12pt;">● gRPC & Protobuf</span><br><span style="font-size: 12pt;">● Internal Platform APIs and Developer Tooling</span></p> <p><strong><span style="font-size: 12pt;">What Success Looks Like</span></strong></p> <p><br><span style="font-size: 12pt;">Within your first year, you'll have made a measurable impact on production reliability. Success l</span><span style="font-size: 12pt;">ooks like:</span></p> <p><br><span style="font-size: 12pt;">● Reducing mean time to detection (MTTD) and mean time to recovery (MTTR) across key </span><span style="font-size: 12pt;">production systems.</span><br><span style="font-size: 12pt;">● Building automation that handles a meaningful percentage of incident scenarios without </span><span style="font-size: 12pt;">human intervention.</span><br><span style="font-size: 12pt;">● Becoming a trusted subject matter expert for core platform components and their failure </span><span style="font-size: 12pt;">modes.</span><br><span style="font-size: 12pt;">● Delivering observability and diagnostic tools that other engineers actually use and </span><span style="font-size: 12pt;">depend on.</span><br><span style="font-size: 12pt;">● Establishing SLO baselines and driving engineering investment based on reliability data.</span><br><span style="font-size: 12pt;">● Spending less of your time—and your teammates time—on repetitive manual toil.</span></p> <p><br><span style="font-size: 12pt;">Your impact won't be measured by the number of incidents you respond to—it will be measured </span><span style="font-size: 12pt;">by how reliably our systems run and how quickly we recover when they don't.</span></p>
Find Jobs in United Kingdom on Arbeitnow
Quality
- x Salary range stated weight 35%
- x Remote policy stated weight 20%
- + Location stated weight 15%
- + Organisation stated weight 15%
- + Publication date stated weight 15%
Not enough history yet to judge honesty signals.
Timeline
-
*
#729775 2026-09-13 13:26 UTCPublished