Infrastructure Operations Engineer
- Status
- Open
- Remote policy
- Not stated
- Employment type
- Not stated
- Salary
- Not stated
- Source
- flyio
- First observed
- 2026-09-04 13:25 UTC
- Last seen
- 2026-09-04 13:25 UTC
- Source claims posted
- 2026-09-04 00:00 UTC
- Consecutive misses
- 0 of 3
What the posting says
Fly.io builds computers for agents. It’s the thing we’re excited about right now and the future we’re building toward: agents that need a real computer, fast to start, isolated, disposable, reachable from anywhere, spun up thousands at a time and gone again. That unlocks a category of work that doesn’t happen on a shared sandbox, and demand for it is a big part of why our fleet is growing the way it is.
The good news, and the reason we can move quickly on it, is that we already built the machine that does this and it has thousands of customers. Our platform transmogrifies Docker containers into Firecracker micro-VMs that run on our own hardware around the world, and connects all of them to a global Anycast network that picks up requests from everywhere and routes them to the nearest VM. That’s what runs apps close to users, and teams run their production workloads on it today. Computers for agents isn’t a separate product bolted on the side, they are born out of the same platform. Our infrastructure needs to support everything, be it an app made by an agent, or a computer that an agent runs on, or the millions of apps that have been running on our platform since before the age of agents.
We run customer apps on our own hardware, in a variety of data centers everywhere on the globe. Keeping that fleet healthy, growing, and ahead of demand is the job of the infrastructure operations team. We’re looking for new members.
About this role
The kernel of this role is incident management. Not in a procedural sense, but in a visceral, hands-on-keyboard, first-and-last responder sense. For this role to be a fit, this stuff needs to be in your blood. Lots of engineers find incident response unpleasant or even paralyzing. You need to find it energizing, like the loose cannon protagonist of a buddy cop movie.
When things go wrong (for you), you’re going to parachute into the middle of it. You’ll inspect physicals and our hosting stack (a mix of standard stuff and a lot of orchestration and systems stuff we built ourselves) and figure out what’s going on. At times, you’ll need to become a world expert in Linux kernel features you only learned about 15 minutes earlier. That’s the gig.
Over the past year, we’ve scaled to thousands of servers with a very small team. While none of us work directly inside data centers, that level of hardware and networking knowledge is genuinely valuable here; we often need to be a step ahead of our upstream partners.
Beyond logging into servers and configuring software, this is a communication role. “Teamwork” is not a platitude here; most incidents require team solutions, and you’ll need to be the force pulling that team together. That includes:
consulting with product teams and figuring out what they actually need
working with our upstream providers to acquire new hardware and chase down issues
getting onto the radar of different platform teams in our organization (Fly Machines, networking, MPG, Sprites) and getting the right people involved in ongoing incidents
much more where that came from.
One big difference between infra at Fly.io and infra at many other places: we have our own Infrastructure as Code stack. You need to be a software developer; you’ll be reading and writing real code in a variety of languages (specifically: we build in Rust, Go, Ruby, and Elixir, on Linux). Willingness to learn any of those languages goes a long way, but you’ll need to have good foundations.
What we’re looking for
Software development experience. Our IaC is our own, written across a variety of languages. You don’t need to already know all of them, but you do need to be willing to learn.
Domain expertise. At minimum, you need to be day-1 comfortable with Linux networking and routing, standard ops telemetry (Prometheus, OpenTelemetry), and Linux performance measurement. You’re a build-your-own-kernel type.
A talented communicator. The role depends on talking to product teams, upstream providers, and whoever’s in the room during an incident, and turning that into the right action.
Hardware and networking experience. We don’t rack servers ourselves, but we work with partners on the physical side, and knowing that world is a real edge.
Operationally calm under fire. When something breaks, you’re the kind of person who jumps in and figures out what’s actually needed.
How We Hire
This is a mid-level, remote, full-time position.
In order to optimize for pay equity, Fly.io doesn’t negotiate salaries. We have standardized salaries for each employee level. The salary for this role is $190k USD, and we offer competitive equity grants with a long exercise window. We provide flexible vacation time (with a minimum), hardware/phone allowances, the standard stuff.
Our hiring process may be a little different from what you’re used to. We respect career experience but we aren’t hypnotized by it, and we’re thrilled at the prospect of discovering new talent. So instead of resumes and interviews, we’re going to show you the kind of work we’re doing and then see if you enjoy actually doing it, with “work-sample challenges”. Unlike a lot of places that assign “take-home problems”, our challenges are the backbone of our whole process; they’re not pre-screeners for an interview gauntlet. (We’re happy to talk, though!)
There’s more about us than you probably want to know at our hiring documentation.
If you’re interested, mail jobs+[email protected]. You can tell us a bit about yourself, if you like. Please include your location, your Github username for work sample access, and your favorite fancy beverage. We probably won’t respond to emails that don’t include all three items.
Quality
- x Salary range stated weight 35%
- x Remote policy stated weight 20%
- x Location stated weight 15%
- + Organisation stated weight 15%
- + Publication date stated weight 15%
Not enough history yet to judge honesty signals.
Timeline
-
*
#570478 2026-09-04 13:25 UTCPublished