Platform and DevOps · Level 2 of 5
Platform Engineer job description
This is what Platform and DevOps teams expect from a Platform Engineer. 36 skills, each with the mastery level set for this rung, and 1 certification required from here on. It is the same framework Competrace ships to new customers, so you can read it here and import it as-is.
Operates infrastructure and CI/CD independently day to day.
Platform Engineer only.
Platform and DevOps — Platform Engineer Operates infrastructure and CI/CD independently day to day. REQUIRED SKILLS Cloud Infrastructure - Cloud Landing Zone Design — You can provision a new project or account from the standard landing zone template, applying the required guardrails and tagging without help. - Infrastructure Provisioning Automation — You can design a reusable module that other teams adopt, structure state so changes stay isolated, and safely refactor a module already running in production. - Network and Cloud Security Configuration — You can configure a security group or firewall rule for a routine request, following the least-privilege pattern the team already uses. - Secrets and Credential Rotation — You can design the rotation schedule and blast-radius for a class of secrets, and lead the response when a credential is suspected to have leaked. - Disaster Recovery and Backup Planning — You can verify that backups for a service are actually completing and restorable, not just scheduled, and escalate when they are not. - Infrastructure Vulnerability and Patch Management — You can run a patch cycle across a fleet with minimal disruption, and negotiate an exception for a system that genuinely cannot be patched on schedule. CI/CD and Delivery - CI Pipeline Engineering — You can design the build, test and packaging stages for a new service on your own, including caching and parallelism, and keep run times honest. You fix the pipeline when it breaks rather than routing around it. - Release Automation and Deployment Strategy — You can automate a rolling deployment to the pattern your team already uses, including health checks and an automatic rollback trigger. You test the rollback rather than assuming it works. - Artifact and Package Registry Management — You can set up publishing and versioning for a new component to the scheme your team follows, and act on the vulnerability findings a scanner raises against images you own. - Environment and Configuration Management — You can keep a set of environments reproducible and consistent on your own, including secret handling and promoting a configuration change safely up to production. You rebuild an environment from its definitions, not from memory. Container Platform - Container Orchestration — You can write a manifest for a new workload, set sensible resource requests and limits, and debug a crash loop from the events and logs. - Container Image Build and Hardening — You can build minimal, reproducible images for a new service unaided, choosing the base, trimming layers and keeping the build cache useful. You decide what belongs in the image and what belongs at runtime. - Service Mesh Operations — You can read the mesh configuration for a service and explain what traffic it allows. You apply a policy someone else has written and confirm the sidecar picked it up, with a reviewer checking your change. SRE and Reliability - On-Call Incident Command — You can triage an incident, communicate status at a steady cadence, and mitigate a known failure mode without waiting for direction. - Service Level Objective Management — You can propose an SLO for a simple service from its actual user impact, and track the error budget it consumes week to week. - Alert Design and Noise Reduction — You can design a new alert around user-facing symptoms rather than internal signals, and remove an alert that nobody has acted on in months. - Postmortem and Incident Learning — You can facilitate a postmortem for a complex incident, keep it blameless when tempers are short, and push back on an action item that will not actually prevent recurrence. - Chaos Engineering and Failure Testing — You can take part in a planned failure test on a system, following the test plan someone else wrote. Cost and DevEx - Cloud Cost Management — You can tag resources so spend lands against the right team, and clear the obvious waste: idle instances, over-provisioned storage, test environments nobody switched off. You raise a rising line on a chart before month end. - Capacity Forecasting and Right-Sizing — You can right-size instances and container requests for services you own from their observed usage, and flag a limit that is throttling a workload. You base the change on measurement, not on the number that felt safe. - Platform Self-Service Tooling — You can build a service template or a small command to the patterns your team already uses, and make its failure messages say what to do next. You take the feedback from the engineers using it and act on it. - Developer Platform Documentation and Enablement — You can own the documentation for a part of the platform, from onboarding through reference to troubleshooting, and keep it current as the platform ships. You answer the repeat questions in the docs instead of in the support channel. Delivery - Project Management — You can break a small piece of work into tasks, sequence them, and run it to a date you agreed, escalating risks before they become slips. - Planning & Estimation — You can estimate your own tasks with reasonable accuracy and deliver at a steady enough pace that other people can plan around you. - Ownership & Accountability — You can pick up a problem that has no obvious owner and make yourself the accountable party for it, including the parts nobody enjoys. - Quality Focus — You can define what good enough means for a project, put the checks in place to prove it, and hold back a release that misses the bar. Craft - Problem Solving — You can break a large, ambiguous problem into tractable pieces, weigh the options against evidence, and explain the tradeoff you chose. - Domain Expertise — You can work confidently across the systems and tools your team uses daily, and you can explain how your work serves the team's goals. - Continuous Learning — You can pick up an unfamiliar area fast enough to be useful in it, and you turn what you learned into something others can reuse. Communication - Communication — You can explain a complex topic to people with very different backgrounds, adjusting the detail to the audience without talking down to them. - Collaboration — You can build working relationships beyond your own team and get things done through people who do not report to you. - Technical Writing — You can write for a defined audience, whether a proposal, a runbook, or a decision record, and make the reasoning as clear as the conclusion. - Stakeholder Management — You can identify who is affected by your work and keep them updated at a level of detail and a cadence that suits them. Leadership - Leadership — You can identify a problem and propose a way forward unprompted, and you take a small leading role such as running a working group. - Mentoring — You can onboard someone onto your team's work, answer their questions patiently, and give feedback specific enough to act on. - Strategic Thinking — You can connect your team's work to its goals and question a task that does not appear to serve them. REQUIRED CERTIFICATIONS - HashiCorp Certified: Terraform Associate (required from Platform Engineer)
Import this exact framework into your own org
Create a free account and Platform and DevOps lands in your org as a department: all 36 skills, with the mastery expected at each of your 5 career levels — already filled in. Rename or delete anything you don't want.