Platform and DevOps · Level 3 of 5
Senior Platform Engineer job description
This is what Platform and DevOps teams expect from a Senior Platform Engineer. 36 skills, each with the mastery level set for this rung, and 3 certifications required from here on. It is the same framework Competrace ships to new customers, so you can read it here and import it as-is.
Owns reliability and delivery for a platform area end to end.
Senior Platform Engineer only.
Platform and DevOps — Senior Platform Engineer Owns reliability and delivery for a platform area end to end. REQUIRED SKILLS Cloud Infrastructure - Cloud Landing Zone Design — You can extend the landing zone with a new network segment or account type, and you judge a requested exception to the guardrails on its risk rather than on precedent. - Infrastructure Provisioning Automation — You can plan a multi-module change that touches several environments, sequence it to avoid an outage, and set the conventions your team writes infrastructure code to. - Network and Cloud Security Configuration — You can design the network segmentation and access boundaries for a new service, and tighten an existing configuration you find is over-permissive. - Secrets and Credential Rotation — You can design the rotation schedule and blast-radius for a class of secrets, and lead the response when a credential is suspected to have leaked. - Disaster Recovery and Backup Planning — You can write and test a disaster recovery plan for a service, including its recovery time and recovery point objectives, and run the drill yourself. - Infrastructure Vulnerability and Patch Management — You can run a patch cycle across a fleet with minimal disruption, and negotiate an exception for a system that genuinely cannot be patched on schedule. CI/CD and Delivery - CI Pipeline Engineering — You can rebuild delivery for a repository where builds are slow, flaky or tangled across many services, and you set the pipeline conventions your team works to. You are who people ask when a green build ships a broken artifact. - Release Automation and Deployment Strategy — You can design rollouts under awkward constraints such as stateful data, schema migrations or a change that spans several services, so a bad version reaches few users. You are called in when a deployment strands a service half-upgraded. - Artifact and Package Registry Management — You can run a registry for a set of services unaided: version policy, promotion between repositories, scanning gates and retention rules that stop storage growing without limit. - Environment and Configuration Management — You can keep a set of environments reproducible and consistent on your own, including secret handling and promoting a configuration change safely up to production. You rebuild an environment from its definitions, not from memory. Container Platform - Container Orchestration — You can design multi-cluster or multi-region topology, and lead the response when the orchestration layer itself is the thing failing. - Container Image Build and Hardening — You can build minimal, reproducible images for a new service unaided, choosing the base, trimming layers and keeping the build cache useful. You decide what belongs in the image and what belongs at runtime. - Service Mesh Operations — You can write traffic policy for a service on your own, including routing rules, retries, circuit breaking and mutual TLS across namespaces. You debug a request the mesh dropped from the proxy logs and metrics rather than guessing. SRE and Reliability - On-Call Incident Command — You can lead the response to a severe, multi-team incident under real pressure, and keep the communication honest when the news is bad. - Service Level Objective Management — You can set SLOs for a service with its stakeholders, and use a burning error budget to justify slowing down feature work. - Alert Design and Noise Reduction — You can audit an alerting setup across several services, cut its noise significantly, and get the team to trust paging again. - Postmortem and Incident Learning — You can facilitate a postmortem for a complex incident, keep it blameless when tempers are short, and push back on an action item that will not actually prevent recurrence. - Chaos Engineering and Failure Testing — You can run failure experiments against production with real safeguards, and turn a surprising result into a concrete reliability fix. Cost and DevEx - Cloud Cost Management — You can attribute a platform's spend to teams and features, forecast the month and cut cost without weakening reliability. You judge for yourself when a saving is not worth the risk it introduces. - Capacity Forecasting and Right-Sizing — You can forecast demand from usage trends and seasonality and provision to match, including autoscaling rules that hold at peak. You set the headroom yourself and can defend the figure. - Platform Self-Service Tooling — You can design and ship a golden path for a common task end to end, from provisioning to first deploy, so a product team runs it without you. You know when a request should become a tool and when it should stay a one-off. - Developer Platform Documentation and Enablement — You can own the documentation for a part of the platform, from onboarding through reference to troubleshooting, and keep it current as the platform ships. You answer the repeat questions in the docs instead of in the support channel. Delivery - Project Management — You can run a multi-person project end to end: you set the scope, track the dependencies between the people involved, and re-plan when reality moves. - Planning & Estimation — You can estimate a whole project including its risks and unknowns, split it into milestones, and hold the estimate up under challenge. - Ownership & Accountability — You can hold accountability for outcomes delivered mostly by other people, absorbing the blame when it fails and passing on the credit when it works. - Quality Focus — You can design the quality practice for complex work owned by several teams, and you anticipate the failure modes that only appear once systems interact. Craft - Problem Solving — You can solve problems in domains where you are not the expert, and your solutions hold up on cost, performance, and maintainability at once. - Domain Expertise — You can be the person the team consults on your area, and you follow where the field is moving and apply it where it pays off. - Continuous Learning — You can judge which new ideas are worth the team's time and which are not, and you make room for the people around you to learn too. Communication - Communication — You can bring disagreeing groups to a shared understanding, and colleagues come to you for help framing a difficult or sensitive message. - Collaboration — You can align teams with competing priorities on a common goal, surfacing the conflict early instead of letting it harden into resentment. - Technical Writing — You can own the documentation of a large project, coordinating contributions so the work can be maintained by people who never built it. - Stakeholder Management — You can win support for a proposal, reset expectations when the plan changes, and refuse a request with a reason the other side accepts. Leadership - Leadership — You can take charge of work with no clear owner, motivate people who do not report to you, and make calls others are willing to follow. - Mentoring — You can mentor someone over months, tell them the uncomfortable thing they need to hear, and adapt how you teach to how they learn. - Strategic Thinking — You can look a year ahead in your area, name what will matter by then, and turn that into work people can pick up now. REQUIRED CERTIFICATIONS - HashiCorp Certified: Terraform Associate (required from Platform Engineer) - Certified Kubernetes Administrator (CKA) (required from Senior Platform Engineer) - Google Professional Cloud Architect (required from Senior Platform Engineer)
Import this exact framework into your own org
Create a free account and Platform and DevOps lands in your org as a department: all 36 skills, with the mastery expected at each of your 5 career levels — already filled in. Rename or delete anything you don't want.