Senior Site Reliability / DevOps Engineer– AI Products

Navan - Tel-Aviv, Israel - original posting ->
Status
Open
Remote policy
Not stated
Employment type
Not stated
Salary
Not stated
Categories
Engineering
Tech
grafanakubernetesprometheusterraformgopythondevopssenior
Source
tripactions
First observed
2026-08-19 07:55 UTC
Last seen
2026-08-30 12:02 UTC
Source claims posted
2026-07-07 05:33 UTC
Consecutive misses
0 of 3

What the posting says

At Navan, "It's all about the user. All of them." We're passionate about providing a seamless one-stop experience for business travelers, no matter how they travel, where they stay, or where they're going. Navan is building cutting-edge solutions at the intersection of travel, expense, payments, and AI. As a leader in the AI for Travel domain, we are using intelligent, practical AI experiences to make business travel simpler, faster, and more reliable for travelers, travel managers, finance teams, and support teams.

We are seeking a Senior Site Reliability / DevOps Engineer to ensure the scalability, performance, and reliability of our user-facing generative AI features.

In this role, you will bridge the gap between traditional infrastructure and cutting-edge machine learning.

You will build and maintain the high-throughput, low-latency systems required to serve AI models directly to millions of users.

This position is based out of our new Tel Aviv office.

What You'll Do:

Infrastructure Ownership: Design, build, and scale the infrastructure hosting our user-facing AI applications and inference engines.

Performance Optimization: Optimize system latency, specifically targeting Time-to-First-Token (TTFT) and total round-trip time for user requests.

GPU & Resource Orchestration: Manage and scale GPU clusters within Kubernetes to maximize utilization and minimize operational costs.

Resiliency & Fallbacks: Build robust fallback systems, circuit breakers, and rate-limiting infrastructure to handle upstream LLM API failures and traffic spikes.

Monitoring & Observability: Implement deep observability for AI workloads, tracking custom metrics like token usage, model drift, and GPU memory saturation.

What We're Looking For:

SRE Fundamentals: 4+ years of experience in SRE, DevOps, or Production Engineering roles supporting high-traffic, user-facing applications.

LLMOps / AI Infrastructure Expertise: Experience working with AI workloads (such as serving models using vLLM, Server TGI or working with cloud providers like Bedrock, OpenAI, etc.).

Container Orchestration: Strong expertise in Kubernetes (EKS, GKE, or AKS) and infrastructure-as-code (Terraform).

AI/ML Ecosystem: Hands-on experience with inference servers (e.g., vLLM, TGI) and vector databases (e.g., Pinecone, Milvus, Qdrant).

Programming: Proficiency in Python and Go for automation, tooling, and backend optimization.

Cloud Architecture: Deep experience managing cloud compute resources, specifically specialized GPU instances

Models AI and Code:

Ability to build automation processes that not only update code versions, but also support testing and safe deployment of new models (Shadow Deployments, Canary releases for models) Product thinking and user orientation (User-Facing)

Advanced Observability:

Mastery of tools like OpenTelemetry, Prometheus, Datadog or Grafana, with the ability to trace agent-based systems and complex model calls.

Cost & Capacity Optimization:

Ability to manage the high costs of GPU/Inference in a productive architecture without compromising availability or performance.

Empathy for the end-user experience:

Understanding that every millisecond of latency or error in the stream directly impacts customer retention.

Preferred Qualifications

Experience building semantic caching layers to reduce LLM API costs.

Active contributor to open-source LLMOps or MLOps projects. (edited)

Navan uses AI-assisted Automated Employment Decision Tool (Metaview) to assist with evaluating resumes against job qualifications for this role. All final decisions are made by human recruiters and hiring managers.

Human oversight: Metaview does not automatically reject candidates or make final hiring decisions. Our recruiters and hiring managers review all outputs and make the final hiring decision regarding every application.

Your rights: If you prefer to have your application reviewed without AI assistance, you may request a human evaluation by entering your email here. Your decision to do so will not affect how your candidacy is evaluated.

Please refer to our Candidate Privacy Notice for more information about our processing of personal data, and your rights.

Quality

Completeness: 45%
Honesty: 100%

Based on 4 observation(s).

Timeline

  1. *
    #177176 2026-08-19 07:55 UTC
    Published
  2. ~
    #200370 2026-08-20 01:18 UTC
    Modified
    • source_updated_at
      2026-08-17T18:05:15-04:00->2026-08-19T17:27:16-04:00
  3. ~
    #208289 2026-08-20 09:25 UTC
    Modified
    • source_updated_at
      2026-08-19T17:27:16-04:00->2026-08-20T04:07:32-04:00
  4. ~
    #463077 2026-08-30 12:02 UTC
    Modified
    • Title
      Sr. SRE AI Engineer->Senior Site Reliability / DevOps Engineer– AI Products
    • Description
      At Navan, "It's all about the user. All of them." We're passionate about providing a seamless one-stop experience for business travelers,...->At Navan, "It's all about the user. All of them." We're passionate about providing a seamless one-stop experience for business travelers,...
    • source_updated_at
      2026-08-20T04:07:32-04:00->2026-08-30T04:58:51-04:00