← MLOps & Model PlatformAll roles
MLOps · Full-timeRemote
Platform SRE (Edge Runtimes & Kubernetes)
StaffApply by 2026-12-31
SREKubernetesCloudflare WorkersObservabilityTerraformIncident Response
Apply nowThe work
Halcer is hiring a Platform SRE to keep edge and cluster workloads reliable. You will own SLOs, toil reduction, and safe delivery pipelines for multi-tenant systems — describing incidents by impact class, not by customer name.
You will
- Define SLIs/SLOs and error budgets for edge and Kubernetes services
- Reduce operational toil through automation, runbooks, and paved-road CI/CD
- Lead incident response with blameless reviews and durable fixes
- Guide application teams on retries, timeouts, and backpressure at the edge
You likely have
- 7+ years in SRE, platform engineering, or production operations
- Deep Kubernetes plus familiarity with isolate/edge runtimes such as Cloudflare Workers
- Infrastructure-as-code fluency (Terraform or equivalent)
- Clear incident writing and stakeholder updates under pressure
Candidate apply
Apply for Platform SRE (Edge Runtimes & Kubernetes)
Six fields to start. Your profile is saved locally for faster re-apply across roles.
Platform SRE (Edge Runtimes & Kubernetes)
Apply