Platform and DevOps · Level 4 of 5
Staff Platform Engineer job description
This is what Platform and DevOps teams expect from a Staff Platform Engineer. 36 skills, each with the mastery level set for this rung, and 4 certifications required from here on. It is the same framework Competrace ships to new customers, so you can read it here and import it as-is.
Sets platform direction across teams and leads incidents.
Staff Platform Engineer only.
Platform and DevOps — Staff Platform Engineer Sets platform direction across teams and leads incidents. REQUIRED SKILLS Cloud Infrastructure - Cloud Landing Zone Design — You can redesign part of the landing zone to fix a structural problem, such as a blast-radius or billing boundary, and roll the change out without breaking existing workloads. - Infrastructure Provisioning Automation — You can plan a multi-module change that touches several environments, sequence it to avoid an outage, and set the conventions your team writes infrastructure code to. - Network and Cloud Security Configuration — You can lead a network or access redesign spanning several services, and you are the one called in when a security review finds a serious gap. - Secrets and Credential Rotation — You can eliminate a category of long-lived secrets by moving it to short-lived or dynamically issued credentials, and get other teams to adopt the pattern. - Disaster Recovery and Backup Planning — You can coordinate a multi-service failover drill, find the dependency nobody had written down, and get it fixed before it causes a real outage. - Infrastructure Vulnerability and Patch Management — You can drive remediation of a vulnerability that spans many teams' systems, and change the pipeline so the same class of gap does not recur. CI/CD and Delivery - CI Pipeline Engineering — You can rebuild delivery for a repository where builds are slow, flaky or tangled across many services, and you set the pipeline conventions your team works to. You are who people ask when a green build ships a broken artifact. - Release Automation and Deployment Strategy — You can design rollouts under awkward constraints such as stateful data, schema migrations or a change that spans several services, so a bad version reaches few users. You are called in when a deployment strands a service half-upgraded. - Artifact and Package Registry Management — You can design artifact management across many teams, including signing and provenance, so any deployed build traces back to the commit it came from. You are who people ask when an image cannot be matched to any source. - Environment and Configuration Management — You can untangle environments that have drifted apart over years and bring them back under source control without downtime. You are who people ask when a change works everywhere except production. Container Platform - Container Orchestration — You can design multi-cluster or multi-region topology, and lead the response when the orchestration layer itself is the thing failing. - Container Image Build and Hardening — You can harden images for services with awkward runtime needs and shrink the attack surface across a whole fleet without breaking builds. You are who people ask when an image passes the scanner but still fails a security review. - Service Mesh Operations — You can plan a mesh rollout or upgrade across many services without dropping traffic, and settle what the mesh should own versus the application. You are called in when latency or a failure sits in the sidecar rather than the code. SRE and Reliability - On-Call Incident Command — You can lead the response to a severe, multi-team incident under real pressure, and keep the communication honest when the news is bad. - Service Level Objective Management — You can set SLOs across a portfolio of services, resolve a dispute between teams about what the objective should be, and retire an SLO that no longer reflects user experience. - Alert Design and Noise Reduction — You can audit an alerting setup across several services, cut its noise significantly, and get the team to trust paging again. - Postmortem and Incident Learning — You can spot the same contributing cause recurring across several postmortems, and get the underlying fix prioritised over the next incident. - Chaos Engineering and Failure Testing — You can build the case for a failure-testing programme across several services, and get sceptical teams to opt in. Cost and DevEx - Cloud Cost Management — You can take on the spend nobody owns, such as data transfer and shared clusters, where architecture, commitments and usage all interact. You are who people ask when the bill jumps and nobody can say why. - Capacity Forecasting and Right-Sizing — You can plan capacity across a fleet whose workloads compete for one pool, and size for events such as a launch or a migration with no usable history. You are who people ask when a service falls over at a load the forecast said it would survive. - Platform Self-Service Tooling — You can make self-service work for tasks that are genuinely hard to make safe, such as production access or schema changes, keeping the guardrails without making the path unusable. You are who people ask when teams route around the platform. - Developer Platform Documentation and Enablement — You can plan enablement for a large platform change so teams migrate without queueing at your desk, and restructure documentation that has grown into a maze. You are who people ask when a rollout stalls because nobody understands the new path. Delivery - Project Management — You can run work spanning several teams, negotiate scope and sequencing with their owners, and maintain one shared plan that all of them actually use. - Planning & Estimation — You can estimate work spanning several teams, name the assumptions each figure rests on, and re-cut the plan as those assumptions break. - Ownership & Accountability — You can hold accountability for outcomes delivered mostly by other people, absorbing the blame when it fails and passing on the credit when it works. - Quality Focus — You can design the quality practice for complex work owned by several teams, and you anticipate the failure modes that only appear once systems interact. Craft - Problem Solving — You can solve problems in domains where you are not the expert, and your solutions hold up on cost, performance, and maintainability at once. - Domain Expertise — You can bring in practice from outside the organisation and make it work here, and your judgement demonstrably improves the projects you touch. - Continuous Learning — You can judge which new ideas are worth the team's time and which are not, and you make room for the people around you to learn too. Communication - Communication — You can bring disagreeing groups to a shared understanding, and colleagues come to you for help framing a difficult or sensitive message. - Collaboration — You can align teams with competing priorities on a common goal, surfacing the conflict early instead of letting it harden into resentment. - Technical Writing — You can own the documentation of a large project, coordinating contributions so the work can be maintained by people who never built it. - Stakeholder Management — You can hold senior and external relationships, negotiate between competing demands, and deliver unwelcome news without losing trust. Leadership - Leadership — You can lead across team boundaries, build the credibility that makes people follow you by choice, and create room for others to lead. - Mentoring — You can develop other mentors, coach people through career decisions rather than tasks, and lift the capability of a whole team. - Strategic Thinking — You can set direction for an area, choose deliberately what not to do, and defend that choice when it is challenged. REQUIRED CERTIFICATIONS - HashiCorp Certified: Terraform Associate (required from Platform Engineer) - Certified Kubernetes Administrator (CKA) (required from Senior Platform Engineer) - Google Professional Cloud Architect (required from Senior Platform Engineer) - AWS Certified DevOps Engineer - Professional (required from Staff Platform Engineer)
Import this exact framework into your own org
Create a free account and Platform and DevOps lands in your org as a department: all 36 skills, with the mastery expected at each of your 5 career levels — already filled in. Rename or delete anything you don't want.