Site Reliability Engineer (AI Start-up)

Status
Open
Remote policy
Not stated
Employment type
Not stated
Salary
100,000-140,000 CHF / year
Tech
awsazuregcppostgresredisdockerkubernetesterraformpythontypescriptdevops
Source
swissdevjobs
First observed
2026-10-03 18:27 UTC
Last seen
2026-10-03 18:27 UTC
Source claims posted
2026-10-01 22:00 UTC
Consecutive misses
0 of 10

What the posting says

Salary: CHF 100'000 - 140'000 per year

Requirements:

Experience designing and operating distributed systems at scale, with a solid understanding of failure modes, capacity planning and the trade-offs between consistency, availability and latency

Strong observability and incident response skills: you build monitoring that catches problems before customers do, lead structured incident responses and drive lasting fixes, not just restarts

Hands-on experience with infrastructure as code, CI/CD pipelines and container orchestration in production

A security mindset: encryption, network isolation and secure handling of enterprise data are a default for you, not an afterthought

Based in, or ready to relocate to, Prague, Berlin, Lisbon, Porto or Valencia, with the right to work in the EU (no visa sponsorship)

Responsibilities:

You'll join a fast-growing, well-funded AI start-up backed by one of Europe's top venture capital firms. They're building an AI operations platform for large enterprises, currently focused on retail and consumer goods. Their AI agents carry out business-critical work wherever it happens, across SAP, spreadsheets, supplier portals, email and APIs, with heavy use of browser and computer automation. Customers rely on these agents for critical workflows, so reliability directly affects their business.

What you'll do

This is not a traditional ops role that follows runbooks. You'll build the reliability practice, not maintain one:

Evolve observability, infrastructure topology and release processes to keep up with high-velocity, AI-assisted development

Make sure misbehaving AI agents are spotted early: understand their failure modes and find the causes of high latency

Build tooling that sounds the alarm before customers notice, so the on-call team can assess impact and mitigate issues in minutes, not hours

Make sure post-mortem actions are followed through across engineering, even if that means adding some friction

Decide what to automate first, balance reliability against shipping speed, and make incident calls with incomplete information

The team

The engineering team has around 20 people and is growing. About half are based at the home base in Prague, and the rest work from small hubs across Europe, all within two hours of Prague time. The culture is built on ownership, direct feedback and customer focus: the team ships weekly, fixes forward, and uses AI tools wherever they help.

Technologies:

AI

AI Agents

AWS

Azure

CI/CD

Cloud

Datadog

GCP

Kubernetes

LLM

Network

REST

SAP

Security

Terraform

Docker

PostgreSQL

Python

Redis

TypeScript

More:

Rockstar Recruiting is hiring on behalf of a fast-growing, well-funded AI start-up, backed by one of Europe's top venture capital firms, that is building an AI operations platform for large enterprises, with customers among Europe's biggest retailers. This is a senior Site Reliability Engineer role for someone who knows what good looks like and wants to build a reliability practice from the ground up: owning observability, incident response, security and the reliability of production AI agents, on a cloud-native stack built on GCP, Kubernetes, Terraform and Datadog.

The team works from small hubs across Europe rather than fully remote, so the role is open to candidates based in, or ready to relocate to, Prague, Berlin, Lisbon, Porto or Valencia, with the right to work in the EU. Please get in touch with Tijana of Rockstar Recruiting to discuss further details: [email protected]

last updated 40 week of 2026

Quality

Completeness: 65%

Not enough history yet to judge honesty signals.

Timeline

  1. *
    #1159131 2026-10-03 18:27 UTC
    Published