Sr Availability TPM, DCC Availability
# Sr Availability TPM, DCC Availability **Amazon Data Services, Inc.** · Herndon, Virginia, USA · `On-site` · `full-time` 🕒 **Статус:** *Опубликовано: сегодня* · *Источник: Amazon* --- ### Top Skills & Match 🎯 **Ключевой стек роли:** `[Project/Program/Product Management--Technical]` `[Technical Program Management]` --- ### About the Role AWS Infrastructure Services owns the design, planning, delivery, and operation of all AWS global infrastructure. In other words, we’re the people who keep the cloud running. We support all AWS data centers and all of the servers, storage, networking, power, and cooling equipment that ensure our customers have continual access to the innovation they rely on. We work on the most challenging problems, with thousands of variables impacting the supply chain — and we’re looking for talented people who want to help. You’ll join a diverse team of software, hardware, and network engineers, supply chain specialists, security experts, operations managers, and other vital roles. You’ll collaborate with people across AWS to help us deliver the highest standards for safety and security while providing seemingly infinite capacity at the lowest possible cost for our customers. And you’ll experience an inclusive culture that welcomes bold ideas and empowers you to own them to completion. The Availability team is looking for a Sr Availability TPM to serve as a technical resource and leader. The builder will be responsible for driving large-scale global projects and initiatives, that directly impact capacity delivery and availability for our customers. If you like to interact with customers and a diverse array of stakeholders, track and analyze performance data, and develop new processes and methods to drive improvements, we’d like to meet you. Your work will help ensure delivering high availability for AWS customers. The Alarm Reduction program objective is to sustain a fleet-wide reduction in high-severity infrastructure alarms across data center sites globally, coordinating remediation across engineering and operations teams and closing the feedback loop upstream so alarm quality is addressed in design, new-product readiness, and commissioning rather than remediated after go-live. The program comprises three main areas: analysis and prioritization, cross-team remediation, and upstream process integration. Analysis and prioritization involves maintaining a fleet-wide view of top alarm drivers, identifying common trends and systemic failure modes, and maintaining the alarm reduction metrics and dashboards used in recurring business reviews. Cross-team remediation involves coordinating Field Engineering, controls/automation, data center operations, and the operations center against prioritized alarm sources, and driving remediation plans through to measurable reduction across both new and existing sites. Upstream process integration focuses on feeding systemic alarm findings back into controls design standards, integrating alarm review into the new-product readiness process, and improving the pre-turnover commissioning hand-off so new sites arrive without alarm issues already present. Key job responsibilities Our Availability TPMs are individuals who demonstrate initiative and proactively seek solutions to problems. • Own and deliver large-scale and complex global engineering and operational programs and initiatives that directly impact capacity delivery and availability for our customers. • Partnering with and influence the direction of multiple engineering and operations teams within and outside of AWS to deliver complex/cross-functional projects • Mentor, train, and develop career progression for members of the organization . • Obsess over team learning and development, both from a technical/functional and soft skills (critical thinking, emotional intelligence, and adaptability) development perspective. • Develop, improve, and share operational best practices across the region and with peers globally. - 8+ years of technical product or program management experience - 7+ years of working directly with engineering teams experience - Experience managing programs across cross functional teams, building processes and coordinating release schedules - Experience with data center critical infrastructure, controls/SCADA systems, or alarm/event management
- Project/Program/Product Management--Technical
- Technical Program Management
Наблюдалась 2026-09-29, впервые 2026-09-29, источник — Amazon.