Site Reliability Engineer
# Site Reliability Engineer **FTMO** · 11000 Praha · `On-site` 🕒 **Статус:** *Опубликовано: сегодня* · *Источник: Indeed* --- ### About the Role Prague, Czech Republic • IT Take advantage of the opportunity and become a part of the fast-growing fintech FTMO.com. We are looking for a new colleague to help to keep the lights on and our customers happy. The ideal candidate is an SRE enthusiast who wants to be a part of a meaningful project with results visible on a global scale. As our first dedicated SRE , you will help us to build the full reliability lifecycle: defining what “reliable enough” means with SLOs and error budgets, making releases safe, adjusting our incident response process so outages are rare/short/learned-from, engineering for resilience and capacity, and driving toil out of operations. This is a hands-on engineering role (IAC in Terraform, CI/CD pipelines) with a strong platform and influence component (setting the reliability standards other teams build on). The first project you would work on is polishing and improving our observability stack and, together with the Cloud Platform team, finishing the migration of our observability to Elastic. Our infrastructure & tech stack Google Cloud Platform Kubernetes (GKE) Elastic Cloud : OTEL + ELK for observability (and Grafana + Prometheus stack as legacy observability solution). KEDA for event-/metric-driven autoscaling. PagerDuty for on-call routing and incident response. GitLab for source control & CI/CD; Helm , Terraform/Terragrunt for IaC Kafka (RedPanda); Cloudflare . What will be your agenda? Setting reliability standards (SLO/SLI & error budgets): Cooperation with Security and Business to establish guidelines on defining SLOs, error-budget policies, and using budgets to inform release and reliability decisions with the developer teams. Incident management & on-call: Revisiting our out-of-hours on-call processes and the incident lifecycle – PagerDuty routing, severity/escalation, comms, and blameless postmortems with tracked follow-ups. Release & deployment safety: Helping to set up progressive delivery (canary/blue-green), automated rollbacks, and health gates tied to SLOs. Secure-first CI/CD templates: Contribute to our secure-by-default pipelines and project templates, together with Cloud Platform, DevEx & Security teams, to ensure the best practices (reliability gates, security gates, standards) are followed. Production-readiness reviews (PRR): Help our developers to set up processes for go-live gates for our new services: SLOs, alerting, observability,.. before production. Resilience & disaster recovery: Ensure our teams have established processes for backup/restore validations, failovers, and DR drills across the critical services/infrastructure. Capacity & performance engineering: Help our developer and infrastructure teams set up autoscaling policies (KEDA), load/performance testing, and reliability-vs-cost trade-offs. Observability & telemetry: Ensuring we collect metrics, logs, and traces from our critical services using the Elastic observability stack, setting instrumentation standards, and dashboard/alert quality. What do we expect from you? Hands-on experience running production services on Kubernetes (ideally GKE), including firefighting real incidents. Hands-on infrastructure as code (IaC) with Terraform (ideally Terragrunt) Knowledge of cloud infrastructure from major cloud providers (ideally GCP). Practical skills with Linux , containers, and Helm charts. Experience defining SLOs/SLIs and error budgets , and using them to inform release decisions. Practical observability skills : metrics, logs, and traces; Prometheus/Grafana or Elastic (ELK), awareness of OpenTelemetry. Building and maintaining CI/CD pipelines and reusable templates. Familiarity with progressive delivery (canary, blue-green) and automated rollbacks . Awareness of event-/metric-driven autoscaling (KEDA) and reliability-vs-cost trade-offs. Why join the FTMO team? We are a Czech fintech that, since 2015, has grown from an idea into a global project . 300+ amazing teammates . We’re a great team who learn from each other every day. How do we work ? We focus on meaningful work and open communication, while only adopting processes that make our lives easier. Prague, Národní třída. Enjoy our modern offices at the Quadrio shopping center, offering beautiful views and excellent accessibility. What if I don’t trade? No worries. We’ll show you what our product is all about and introduce you to the basics of trading. Free fruit, snacks, and coffee are always within reach in the office. How do we promote strong relationships and well-being? Company cottage, team building events, and running club. Flexible hybrid model. We prefer collaborating in person to keep the team spirit high, but we offer the flexibility you need to stay balanced. The benefits mentioned above apply to FTMO on-site employees in our Prague office. Company cottage in Krkonose mountains Hardware suitable for your position Motivational bonuses for outstanding results VIP discounts in Quadrio shopping centre Teambuildings for developing relationships Welcome Pack for a pleasant start Sick days when you’re not feeling well Multisport card for unlimited activities Snacks and beverages from healthy treats to sweets Relax days 5 days off for your well-being Courses and education suitable for your position Modern office building in the very centre of Prague
Наблюдалась 2026-09-21, впервые 2026-09-21, источник — Indeed.