Staff Platform Engineer
Fliff
Fulltime
Office
With Experience
Europe (European Time Zone Based Candidate)
🥅 sports
Analytics
We’re looking for a Platform Engineer to join our Platform team, which builds shared infrastructure and delivery tooling across Fliff’s products. The team is responsible for Cloud Core, CI/CD, Platform Security, and support triage. In this role, you’ll take ownership of our CI/CD platform and help it scale, making architectural decisions, evaluating trade-offs, and adapting to changing requirements and production challenges. Your work will shape how software delivery and infrastructure operate across the engineering organization. Responsibilities
- Own and evolve our CI/CD platform: GitHub Actions with self-hosted runners (RunsOn), plus an in-progress migration and refactor of our Jenkins setup ("Jenkins Next").
- Build and maintain Infrastructure as Code across AWS using Terraform / Terragrunt / Terramate, following reusable component patterns (CloudPosse-style).
- Design secure-by-default delivery pipelines: branch protection and PR approval rulesets, secrets management (SOPS), and prod-approval flows for critical paths.
- Drive our move from IAM users and long-lived access keys toward IAM roles, and help close foundational security gaps (open ports, unmanaged credentials, missing guardrails).
- Work with Kubernetes (EKS), Helm, and ArgoCD, and support our observability stack (Grafana, Prometheus) and developer platform (Backstage).
- Review and improve existing components — Terraform, pipelines, AWS infra — not only build new ones, and raise the bar on how the team delivers.
- Share the support/triage load and help reduce it through better automation and tooling.
Requirements
- Strong AWS experience and deep hands-on CI/CD expertise (GitHub Actions and/or Jenkins).
- Solid Infrastructure as Code background with Terraform (Terragrunt / Terramate a plus).
- Real Kubernetes / EKS experience, including Helm and GitOps (ArgoCD).
- A security-conscious mindset — comfortable reasoning about IAM, secrets, network exposure, and least-privilege access (a DevSecOps lean is a strong plus).
- The ability to design a solution, then adapt gracefully when constraints change ("now make it global", "now cost is the primary driver") — digging for context rather than getting flustered.
- Sound judgment on edge cases and incidents: traffic spikes 10x, a region goes down — how do you approach the change and handle the outage?
- Effective, deliberate use of AI tooling, with genuine understanding of what it produces and the ability to explain trade-offs it can't.