Site Reliability Engineer (AI Start-up)
- Status
- Open
- Remote policy
- Not stated
- Employment type
- Not stated
- Salary
- 100,000-140,000 CHF / year
- Source
- swissdevjobs
- First observed
- 2026-10-03 18:27 UTC
- Last seen
- 2026-10-03 18:27 UTC
- Source claims posted
- 2026-10-01 22:00 UTC
- Consecutive misses
- 0 of 10
What the posting says
Salary: CHF 100'000 - 140'000 per year
Requirements:
Experience designing and operating distributed systems at scale, with a solid understanding of failure modes, capacity planning and the trade-offs between consistency, availability and latency
Strong observability and incident response skills: you build monitoring that catches problems before customers do, lead structured incident responses and drive lasting fixes, not just restarts
Hands-on experience with infrastructure as code, CI/CD pipelines and container orchestration in production
A security mindset: encryption, network isolation and secure handling of enterprise data are a default for you, not an afterthought
Based in, or ready to relocate to, Prague, Berlin, Lisbon, Porto or Valencia, with the right to work in the EU (no visa sponsorship)
Responsibilities:
You'll join a fast-growing, well-funded AI start-up backed by one of Europe's top venture capital firms. They're building an AI operations platform for large enterprises, currently focused on retail and consumer goods. Their AI agents carry out business-critical work wherever it happens, across SAP, spreadsheets, supplier portals, email and APIs, with heavy use of browser and computer automation. Customers rely on these agents for critical workflows, so reliability directly affects their business.
What you'll do
This is not a traditional ops role that follows runbooks. You'll build the reliability practice, not maintain one:
Evolve observability, infrastructure topology and release processes to keep up with high-velocity, AI-assisted development
Make sure misbehaving AI agents are spotted early: understand their failure modes and find the causes of high latency
Build tooling that sounds the alarm before customers notice, so the on-call team can assess impact and mitigate issues in minutes, not hours
Make sure post-mortem actions are followed through across engineering, even if that means adding some friction
Decide what to automate first, balance reliability against shipping speed, and make incident calls with incomplete information
The team
The engineering team has around 20 people and is growing. About half are based at the home base in Prague, and the rest work from small hubs across Europe, all within two hours of Prague time. The culture is built on ownership, direct feedback and customer focus: the team ships weekly, fixes forward, and uses AI tools wherever they help.
Technologies:
AI
AI Agents
AWS
Azure
CI/CD
Cloud
Datadog
GCP
Kubernetes
LLM
Network
REST
SAP
Security
Terraform
Docker
PostgreSQL
Python
Redis
TypeScript
More:
Rockstar Recruiting is hiring on behalf of a fast-growing, well-funded AI start-up, backed by one of Europe's top venture capital firms, that is building an AI operations platform for large enterprises, with customers among Europe's biggest retailers. This is a senior Site Reliability Engineer role for someone who knows what good looks like and wants to build a reliability practice from the ground up: owning observability, incident response, security and the reliability of production AI agents, on a cloud-native stack built on GCP, Kubernetes, Terraform and Datadog.
The team works from small hubs across Europe rather than fully remote, so the role is open to candidates based in, or ready to relocate to, Prague, Berlin, Lisbon, Porto or Valencia, with the right to work in the EU. Please get in touch with Tijana of Rockstar Recruiting to discuss further details: [email protected]
last updated 40 week of 2026
Quality
- + Salary range stated weight 35%
- x Remote policy stated weight 20%
- x Location stated weight 15%
- + Organisation stated weight 15%
- + Publication date stated weight 15%
Not enough history yet to judge honesty signals.
Timeline
-
*
#1159131 2026-10-03 18:27 UTCPublished