openqareer

Senior DevOps Engineer

Shopse Digital Finance · Mumbai, Maharashtra

# Senior DevOps Engineer **Shopse Digital Finance** · Mumbai, Maharashtra · `On-site` 🕒 **Статус:** *Опубликовано: сегодня* · *Источник: Indeed* --- ### About the Role Key Responsibilities 1. Infrastructure Management (AWS) - Own and manage the entire AWS infrastructure — EC2, RDS, S3, ECS/EKS, VPC, IAM, CloudFront, Route 53, and related services. - Design for high availability and fault tolerance; ensure infra can handle payment-grade uptime requirements. - Right-size and optimise infrastructure for cost without compromising reliability. - Maintain environment parity across dev, staging, and production. 2. CI/CD Pipelines - Build, maintain, and continuously improve CI/CD pipelines across all services supporting a daily release cadence. - Manage deployment triggering across multiple AWS availability zones; handle traffic routing between zones including manual intervention when required. - Ensure fast, reliable, and safe deployments with rollback capabilities and pre/post-release health checks. - Work closely with the engineering team to reduce deployment friction and maintain release velocity. - Manage branching strategies, environment promotion, and deployment gates. 3. Monitoring, Alerting & Incident Response - Own site monitoring end-to-end — set up and maintain CloudWatch and Site24x7 (or equivalent) for uptime checks, dashboards, log aggregation, and alerting. SENIOR DEVOPS ENGINEER Shopse | Full-Time | Own the Infra. Keep Us Running. - Configure intelligent alerting rules that surface real issues without creating noise; maintain runbooks for common alert scenarios. - Be the first responder for production alerts — acknowledge, triage, and resolve or escalate within defined SLAs. Proactive action on alerts is a core expectation of this role. - Conduct root cause analysis (RCA) for all production incidents and drive fixes to prevent recurrence. - Maintain an on-call schedule; this role requires availability outside business hours for critical alerts. 4. System Automation - Automate repetitive operational tasks — provisioning, scaling, patching, backups, and configuration management. - Manage infrastructure-as-code using Terraform or equivalent; no manual console changes in production. - Automate monitoring setup, alerting thresholds, and runbook execution wherever possible. 5. Security & Compliance - Implement and maintain security rules, firewall policies, network ACLs, and IAM roles with least-privilege principles. - Manage VPN setup and access controls for internal teams and vendors. - Manage access provisioning and deprovisioning for all team members across AWS and related services. - Ensure infrastructure compliance with PCI DSS and ISO 27001 requirements — work closely with the InfoSec Lead on audit evidence and remediation. - Manage SSL/TLS certificates, key rotation, secrets management (AWS Secrets Manager / Vault), and encryption at rest and in transit. 6. Releases & Server Operations - Coordinate and execute daily production releases — validate pre-release checklists, trigger deployments across zones, manage traffic routing, and monitor post-release health. - Handle server maintenance, OS upgrades, dependency patching, and scheduled downtime windows. - Manage database backups, restoration drills, and disaster recovery procedures. - Maintain up-to-date infrastructure documentation and runbooks. 7. AI-Augmented Development Infrastructure - Design and manage infrastructure for AI-assisted development workflows — including on-demand provisioning and teardown of ephemeral EC2 or container instances used by AI dev tools such as Claude Code. - Integrate AI dev tooling into CI/CD pipelines — enabling automated code generation, review, and testing stages that spin up isolated compute, execute tasks, and clean up on completion. - Implement IAM policies, network boundaries, and cost guardrails for ephemeral AI development instances. - Build monitoring and observability into AI-augmented pipelines — tracking instance lifecycle, run times, failure rates, and compute costs. Requirements Essential Experience - 5–8 years of hands-on DevOps or infrastructure engineering experience. - Deep, practical AWS expertise — EC2, RDS, ECS/EKS, VPC, IAM, ALB, Route 53, CloudWatch. Ability to architect, troubleshoot, and optimise independently. - Proven experience building and maintaining CI/CD pipelines (GitHub Actions, GitLab CI, Jenkins, or equivalent) including multi-zone deployments. What we really mean by 'reliable and dependable' In a payments business, infrastructure failures have real, immediate consequences for merchants and customers. We need someone who treats a production alert at 11 PM the same way they would at 11 AM — someone who does not need to be chased, who over-communicates during incidents, and who takes pride in a green dashboard. If that describes you, we want to talk. - Hands-on experience with monitoring stacks — CloudWatch, Site24x7, Grafana, Prometheus, Datadog, or equivalent. - Infrastructure-as-code using Terraform, CloudFormation, or Pulumi. - Strong scripting skills — Python, Bash, or equivalent. - Container experience — Docker and Kubernetes (EKS or self-managed). - Experience managing VPN solutions (OpenVPN, AWS Client VPN, WireGuard, or equivalent). - Access management experience — provisioning, deprovisioning, and periodic reviews across cloud and tooling. Nice to Have - Experience integrating AI dev tools (Claude Code, GitHub Copilot, Cursor, or similar) into CI/CD pipelines. - Familiarity with ephemeral environment patterns for automated testing or AI-assisted workflows. - AWS certifications — Solutions Architect, DevOps Engineer Professional, or Security Specialty. - Familiarity with PCI DSS infrastructure requirements and ISO 27001 technical controls. - Experience in a fintech, payments, or BFSI environment.

Наблюдалась 2026-10-06, впервые 2026-10-06, источник — Indeed.

Открыть у работодателя