← MLOps & Model PlatformAll roles
MLOps · Full-timeRemote

Platform SRE (Edge Runtimes & Kubernetes)

StaffApply by 2026-12-31
SREKubernetesCloudflare WorkersObservabilityTerraformIncident Response
Apply now

The work

Halcer is hiring a Platform SRE to keep edge and cluster workloads reliable. You will own SLOs, toil reduction, and safe delivery pipelines for multi-tenant systems — describing incidents by impact class, not by customer name.

You will

  • Define SLIs/SLOs and error budgets for edge and Kubernetes services
  • Reduce operational toil through automation, runbooks, and paved-road CI/CD
  • Lead incident response with blameless reviews and durable fixes
  • Guide application teams on retries, timeouts, and backpressure at the edge

You likely have

  • 7+ years in SRE, platform engineering, or production operations
  • Deep Kubernetes plus familiarity with isolate/edge runtimes such as Cloudflare Workers
  • Infrastructure-as-code fluency (Terraform or equivalent)
  • Clear incident writing and stakeholder updates under pressure

Candidate apply

Apply for Platform SRE (Edge Runtimes & Kubernetes)

Six fields to start. Your profile is saved locally for faster re-apply across roles.

Resume

Upload PDF or Word, or paste your experience. One is enough.

Private POST only — applications are not published on this site. Files go to our hiring pipeline for screening.