Senior AI Infrastructure Engineer, Kubernetes
- Status
- Open
- Remote policy
- Not stated
- Employment type
- Not stated
- Salary
- Not stated
- Categories
- AI Infrastructure Platform
- Source
- arbeitnow
- First observed
- 2026-09-13 13:26 UTC
- Last seen
- 2026-09-13 13:26 UTC
- Source claims posted
- 2026-09-13 12:55 UTC
- Consecutive misses
- 0 of 3
What the posting says
<p><strong>Firmus Technologies</strong></p> <p><span data-contrast="none"><span data-ccp-parastyle="FIR Body" data-ccp-parastyle-defn="{"ObjectId":"baa2b379-a6d3-56f4-af2b-6a08de9434e5|1","ClassId":1073872969,"Properties":[469777841,"Aeonik",469777842,"Aeonik",469777843,"Aeonik",469777844,"Aeonik",469769226,"Aeonik",201342446,"1",201342447,"5",201342448,"1",201342449,"1",201341986,"1",268442635,"20",335551500,"197122",335559740,"264",201341983,"0",335559738,"145",469775450,"FIR Body",201340122,"2",134234082,"true",134233614,"true",469778129,"FIRBody",335572020,"1",469778324,"Body Text"]}">Firmus Technologies is a global </span><span data-ccp-parastyle="FIR Body">leader</span><span data-ccp-parastyle="FIR Body"> pioneering the development </span></span><span data-contrast="none"><span data-ccp-parastyle="FIR Body">and operation of efficient AI infrastructure across Asia Pacific.</span><span data-ccp-parastyle="FIR Body"> </span></span><span data-ccp-props="{"201341983":0,"335559738":145,"335559740":264}"> </span></p> <p><span data-contrast="none"><span data-ccp-parastyle="FIR Body">Founded in Australia in 2019, our mission is to create the most efficient AI infrastructure by </span></span><span data-contrast="none"><span data-ccp-parastyle="FIR Body">combining </span><span data-ccp-parastyle="FIR Body">cutting-edge</span><span data-ccp-parastyle="FIR Body"> technology with a steadfast commitment to sustainability.</span></span><span data-ccp-props="{"201341983":0,"335559738":145,"335559740":264}"> </span></p> <p><span data-contrast="none"><span data-ccp-parastyle="FIR Body">At Firmus, we are unique in our approach. We design, build, and </span><span data-ccp-parastyle="FIR Body">operate</span><span data-ccp-parastyle="FIR Body"> a new class of digital </span></span><span data-contrast="none"><span data-ccp-parastyle="FIR Body">infrastructure – the AI Factory. Through our model-to-grid technology approach, we have pushed </span></span><span data-contrast="none"><span data-ccp-parastyle="FIR Body">the boundaries of multi-generational liquid cooling systems, energy management, AI software </span></span><span data-contrast="none"><span data-ccp-parastyle="FIR Body">orchestration, and construction. For our customers, this approach allows us to make every watt </span></span><span data-contrast="none"><span data-ccp-parastyle="FIR Body">count and deliver low-cost AI tokens globally.</span></span><span data-ccp-props="{"201341983":0,"335559738":145,"335559740":264}"> </span></p> <p> </p> <p><strong>Firmus AI Cloud</strong></p> <p><span data-contrast="none"><span data-ccp-parastyle="FIR Body" data-ccp-parastyle-defn="{"ObjectId":"baa2b379-a6d3-56f4-af2b-6a08de9434e5|1","ClassId":1073872969,"Properties":[469777841,"Aeonik",469777842,"Aeonik",469777843,"Aeonik",469777844,"Aeonik",469769226,"Aeonik",201342446,"1",201342447,"5",201342448,"1",201342449,"1",201341986,"1",268442635,"20",335551500,"197122",335559740,"264",201341983,"0",335559738,"145",469775450,"FIR Body",201340122,"2",134234082,"true",134233614,"true",469778129,"FIRBody",335572020,"1",469778324,"Body Text"]}">Our large-scale GPU cloud platform, Firmus AI Cloud, is purpose-built </span></span><span data-contrast="none"><span data-ccp-parastyle="FIR Body">to deliver energy-efficient AI </span><span data-ccp-parastyle="FIR Body">compute</span><span data-ccp-parastyle="FIR Body"> at scale to customers.</span></span><span data-ccp-props="{"201341983":0,"335559738":145,"335559740":264}"> </span></p> <p><span data-contrast="none"><span data-ccp-parastyle="FIR Body">It empowers developers, enterprises, educational institutions, and government users to train and </span></span><span data-contrast="none"><span data-ccp-parastyle="FIR Body">deploy AI models with unmatched efficiency and cost savings. With an ever-growing suite of services </span></span><span data-contrast="none"><span data-ccp-parastyle="FIR Body">and applications, we are committed to delivering a cloud experience that is market-leading, </span></span><span data-contrast="none"><span data-ccp-parastyle="FIR Body">proprietary, and built to scale.</span></span><span data-ccp-props="{"201341983":0,"335559738":145,"335559740":264}"> </span></p> <p><strong>Role Summary</strong></p> <p>The Senior Kubernetes Engineer, AI Infrastructure owns the technical design and delivery of the backend infrastructure that powers the Firmus Kubernetes platform. This is a hands-on principal-level individual contributor role, responsible for building production-grade cluster lifecycle, control-plane, networking, storage, security, observability, and automation capabilities across GPU-accelerated bare-metal environments.</p> <p>They solve the hardest platform engineering problems, set Kubernetes engineering standards, and provide domain-level technical sign-off for platform designs. They work across AI Platforms, Solutions Architecture & Delivery, networking, security, and operations to create a secure, resilient, multi-tenant platform that can be deployed and operated consistently at AI-factory scale.</p> <p><strong>Key Responsibilities</strong></p> <ul> <li>Define and own the Kubernetes platform reference architecture across management and workload clusters, including control-plane topology, cluster lifecycle, multi-tenancy, workload isolation, and failure-domain design.</li> <li>Build and maintain the backend services, APIs, controllers, operators, and automation required to provision, configure, upgrade, scale, and retire Kubernetes clusters reliably.</li> <li>Engineer repeatable bare-metal Kubernetes deployment and lifecycle workflows using infrastructure-as-code and automated provisioning technologies such as Cluster API, kubeadm, Redfish, PXE, Ironic, or Metal3.</li> <li>Design and operate cluster networking across CNI, ingress, service discovery, DNS, load balancing, network policy, and service mesh; integrate Multus, SR-IOV, BGP, InfiniBand, or RoCE where required for high-performance AI workloads.</li> <li>Define persistent-storage and data-service patterns using CSI, Ceph, local NVMe, object storage, backup and restore, and disaster-recovery mechanisms appropriate for stateful platform and AI workloads.</li> <li>Integrate and productionise NVIDIA GPU and Network Operators, device plugins, drivers, DCGM telemetry, scheduling, quotas, and topology-aware placement for multi-node accelerated workloads.</li> <li>Establish GitOps and CI/CD patterns for platform software, configuration, policy, and release management, with safe testing, progressive rollout, rollback, and upgrade practices.</li> <li>Build platform security into the architecture through identity and access control, RBAC, secrets management, policy-as-code, image and software-supply-chain controls, tenant isolation, and auditable change management.</li> <li>Define service-level objectives and engineer observability for metrics, logs, traces, events, capacity, and performance; lead diagnosis of complex distributed systems failures and eliminate recurring operational toil.</li> <li>Set engineering standards, design patterns, review practices, and operational readiness criteria; mentor senior engineers and resolve cross-team technical decisions while remaining directly involved in implementation.</li> </ul> <p> </p> <p><strong>Skills & Experience</strong></p> <ul> <li>7+ years of progressive infrastructure, systems, or platform engineering experience, including substantial ownership of production Kubernetes platforms and at least 3 years operating at senior staff, principal, or equivalent level.</li> <li>Deep knowledge of Kubernetes internals, including the API server, etcd, scheduler, controller manager, kubelet, admission, CRI, CNI, CSI, reconciliation patterns, cluster performance, upgrades, and control-plane failure modes.</li> <li>Demonstrated experience designing, building, and operating highly available, large scale and multi-cluster Kubernetes platforms on bare metal, private cloud, or hybrid infrastructure.</li> <li>Strong software engineering ability in Go and/or Rust, with practical Python and Bash skills; experience building Kubernetes operators, controllers, admission webhooks, CLIs, or platform services.</li> <li>Expert Linux systems knowledge, including namespaces, cgroups, systemd, kernel, host networking and container runtime behaviour, performance analysis, and low-level troubleshooting.</li> <li>Strong Kubernetes networking expertise across Cilium, Calico, or equivalent CNI implementations, plus load balancing, DNS, ingress, BGP, network policy, and multi-network architectures.</li> <li>Strong infrastructure automation and GitOps experience with tools such as Terraform, Ansible, Argo CD, Flux, GitHub Actions, GitLab CI, or Jenkins.</li> <li>Practical experience with Kubernetes security and governance, including RBAC, OPA Gatekeeper or Kyverno, secrets management, certificate lifecycle, image security, and workload isolation.</li> <li>Experience implementing production observability with Prometheus, Grafana, OpenTelemetry, Loki, Elasticsearch, or equivalent technologies, and using telemetry to manage reliability, capacity, and performance.</li> <li>Experience with GPU-enabled Kubernetes infrastructure, NVIDIA GPU Operator, accelerator scheduling for AI workloads at large scale, RDMA networking, and distributed AI workload requirements.</li> <li>Experience with distributed storage and data services such as Ceph, CSI-backed storage, object storage, backup and restore, and disaster recovery.</li> <li>CKA-level expertise is expected; CKA, CKS, or relevant cloud-native certifications are strongly preferred.</li> <li>Bachelor’s degree in computer science, engineering, or a related discipline, or equivalent depth of practical engineering experience.</li> <li>Clear technical judgement and communication, with a record of influencing architecture across software, networking, security, platform, and operations teams.</li> </ul> <p> </p> <p><strong>Location & Reporting</strong></p> <ul> <li>Australia (Sydney, NSW or Launceston, TAS)</li> <li>Reporting to Head of AI Platform</li> </ul> <p> </p> <p><strong>Employment Basis</strong></p> <p>Full-time</p> <p> </p> <p><strong>Diversity</strong></p> <p>At Firmus, we are committed to building a diverse and inclusive workplace. We encourage applications from candidates of all backgrounds who are passionate about creating a more sustainable future through innovative engineering solutions.</p> <p>Join us in our mission to revolutionize the AI industry through sustainable practices and cutting-edge engineering. Apply now to be part of shaping the future of sustainable AI infrastructure.</p>
Find more English Speaking Jobs in United Kingdom on Arbeitnow
Quality
- x Salary range stated weight 35%
- x Remote policy stated weight 20%
- + Location stated weight 15%
- + Organisation stated weight 15%
- + Publication date stated weight 15%
Not enough history yet to judge honesty signals.
Timeline
-
*
#729770 2026-09-13 13:26 UTCPublished