Senior Site Reliability / DevOps Engineer– AI Products
- Status
- Open
- Remote policy
- Not stated
- Employment type
- Not stated
- Salary
- Not stated
- Categories
- Engineering
- Source
- tripactions
- First observed
- 2026-08-19 07:55 UTC
- Last seen
- 2026-08-30 12:02 UTC
- Source claims posted
- 2026-07-07 05:33 UTC
- Consecutive misses
- 0 of 3
What the posting says
At Navan, "It's all about the user. All of them." We're passionate about providing a seamless one-stop experience for business travelers, no matter how they travel, where they stay, or where they're going. Navan is building cutting-edge solutions at the intersection of travel, expense, payments, and AI. As a leader in the AI for Travel domain, we are using intelligent, practical AI experiences to make business travel simpler, faster, and more reliable for travelers, travel managers, finance teams, and support teams.
We are seeking a Senior Site Reliability / DevOps Engineer to ensure the scalability, performance, and reliability of our user-facing generative AI features.
In this role, you will bridge the gap between traditional infrastructure and cutting-edge machine learning.
You will build and maintain the high-throughput, low-latency systems required to serve AI models directly to millions of users.
This position is based out of our new Tel Aviv office.
What You'll Do:
Infrastructure Ownership: Design, build, and scale the infrastructure hosting our user-facing AI applications and inference engines.
Performance Optimization: Optimize system latency, specifically targeting Time-to-First-Token (TTFT) and total round-trip time for user requests.
GPU & Resource Orchestration: Manage and scale GPU clusters within Kubernetes to maximize utilization and minimize operational costs.
Resiliency & Fallbacks: Build robust fallback systems, circuit breakers, and rate-limiting infrastructure to handle upstream LLM API failures and traffic spikes.
Monitoring & Observability: Implement deep observability for AI workloads, tracking custom metrics like token usage, model drift, and GPU memory saturation.
What We're Looking For:
SRE Fundamentals: 4+ years of experience in SRE, DevOps, or Production Engineering roles supporting high-traffic, user-facing applications.
LLMOps / AI Infrastructure Expertise: Experience working with AI workloads (such as serving models using vLLM, Server TGI or working with cloud providers like Bedrock, OpenAI, etc.).
Container Orchestration: Strong expertise in Kubernetes (EKS, GKE, or AKS) and infrastructure-as-code (Terraform).
AI/ML Ecosystem: Hands-on experience with inference servers (e.g., vLLM, TGI) and vector databases (e.g., Pinecone, Milvus, Qdrant).
Programming: Proficiency in Python and Go for automation, tooling, and backend optimization.
Cloud Architecture: Deep experience managing cloud compute resources, specifically specialized GPU instances
Models AI and Code:
Ability to build automation processes that not only update code versions, but also support testing and safe deployment of new models (Shadow Deployments, Canary releases for models) Product thinking and user orientation (User-Facing)
Advanced Observability:
Mastery of tools like OpenTelemetry, Prometheus, Datadog or Grafana, with the ability to trace agent-based systems and complex model calls.
Cost & Capacity Optimization:
Ability to manage the high costs of GPU/Inference in a productive architecture without compromising availability or performance.
Empathy for the end-user experience:
Understanding that every millisecond of latency or error in the stream directly impacts customer retention.
Preferred Qualifications
Experience building semantic caching layers to reduce LLM API costs.
Active contributor to open-source LLMOps or MLOps projects. (edited)
Navan uses AI-assisted Automated Employment Decision Tool (Metaview) to assist with evaluating resumes against job qualifications for this role. All final decisions are made by human recruiters and hiring managers.
Human oversight: Metaview does not automatically reject candidates or make final hiring decisions. Our recruiters and hiring managers review all outputs and make the final hiring decision regarding every application.
Your rights: If you prefer to have your application reviewed without AI assistance, you may request a human evaluation by entering your email here. Your decision to do so will not affect how your candidacy is evaluated.
Please refer to our Candidate Privacy Notice for more information about our processing of personal data, and your rights.
Quality
- x Salary range stated weight 35%
- x Remote policy stated weight 20%
- + Location stated weight 15%
- + Organisation stated weight 15%
- + Publication date stated weight 15%
Based on 4 observation(s).
- + Days open - fineOpen for 11 days so far
- + Reopen count - fineNever reopened
- + Salary range removed after publication - fineSalary range has not been removed since publication
- + Salary range narrowed - fineSalary range has not narrowed since publication
- + Missing/reappear cycles - fineNo missing-then-reappeared cycles observed
Timeline
-
*
#177176 2026-08-19 07:55 UTCPublished
-
~
#200370 2026-08-20 01:18 UTCModified
-
source_updated_at
2026-08-17T18:05:15-04:00->2026-08-19T17:27:16-04:00
-
source_updated_at
-
~
#208289 2026-08-20 09:25 UTCModified
-
source_updated_at
2026-08-19T17:27:16-04:00->2026-08-20T04:07:32-04:00
-
source_updated_at
-
~
#463077 2026-08-30 12:02 UTCModified
-
Title
Sr. SRE AI Engineer->Senior Site Reliability / DevOps Engineer– AI Products -
Description
At Navan, "It's all about the user. All of them." We're passionate about providing a seamless one-stop experience for business travelers,...->At Navan, "It's all about the user. All of them." We're passionate about providing a seamless one-stop experience for business travelers,... -
source_updated_at
2026-08-20T04:07:32-04:00->2026-08-30T04:58:51-04:00
-
Title