Platform Engineering
Best practices for building Internal Developer Platforms (IDPs) that reduce cognitive load, accelerate delivery, and create golden paths for development teams.
IDP Architecture Layers
A well-designed IDP separates concerns into distinct layers. Each layer abstracts complexity from the one above it.
Developer Interface (Portal / CLI / API)
|
Orchestration Layer (Workflows, Templates, Scaffolding)
|
Integration Layer (APIs, Plugins, Connectors)
|
Resource Layer (Infrastructure, Services, Tools)
| Layer |
Purpose |
Components |
Owned By |
| Developer Interface |
Self-service entry point |
Backstage portal, CLI tools, API gateway |
Platform team |
| Orchestration |
Workflow automation, templating |
Scaffolder, Terraform modules, Crossplane |
Platform team |
| Integration |
Connect tools and services |
Backstage plugins, API adapters, webhooks |
Platform + tool owners |
| Resource |
Actual infrastructure and services |
Kubernetes, databases, CI/CD, monitoring |
Infrastructure team |
| Governance |
Policy enforcement and compliance |
OPA, Kyverno, cost policies, security scans |
Security + platform team |
Platform Team Topology and Responsibilities
Team Structure
| Role |
Responsibility |
Focus Area |
| Platform Product Manager |
Roadmap, prioritization, user research |
Developer needs, adoption metrics |
| Platform Engineer |
IDP core, golden paths, automation |
Infrastructure abstraction, tooling |
| Developer Advocate |
Documentation, onboarding, feedback loops |
DevEx, training, communication |
| SRE/Reliability Lead |
Platform reliability, SLOs, incident response |
Uptime, performance, observability |
| Security Engineer |
Policy-as-code, compliance automation |
Guardrails, scanning, access control |
Interaction Model
Stream-Aligned Teams (consumers)
|
| self-service requests
v
Platform Team (enablers)
|
| golden paths, templates, APIs
v
Infrastructure / Cloud (resources)
Platform teams operate as enabling teams (Team Topologies model). They reduce cognitive load on stream-aligned teams by providing curated, opinionated abstractions.
Backstage: Service Catalog and Developer Portal
catalog-info.yaml -- Service Registration
Every service registers itself in the catalog via a catalog-info.yaml at the repo root.
apiVersion: backstage.io/v1alpha1
kind: Component
metadata:
name: payment-service
description: Handles payment processing and refunds
annotations:
github.com/project-slug: acme-corp/payment-service
backstage.io/techdocs-ref: dir:.
pagerduty.com/service-id: P1234ABC
grafana/dashboard-selector: "payment-service"
tags:
- java
- spring-boot
- payments
links:
- url: https://grafana.internal/d/payments
title: Dashboard
icon: dashboard
- url: https://runbooks.internal/payments
title: Runbook
icon: docs
spec:
type: service
lifecycle: production
owner: team-payments
system: checkout-system
providesApis:
- payment-api
consumesApis:
- inventory-api
- notification-api
dependsOn:
- resource:payments-db
- component:auth-service
---
apiVersion: backstage.io/v1alpha1
kind: API
metadata:
name: payment-api
description: Payment processing REST API
spec:
type: openapi
lifecycle: production
owner: team-payments
system: checkout-system
definition:
$text: ./openapi.yaml
Service Catalog API -- Querying Components
# List all components owned by a team
curl -s "https://backstage.internal/api/catalog/entities?filter=kind=component,spec.owner=team-payments" \
-H "Authorization: Bearer $BACKSTAGE_TOKEN" | jq '.[] | {name: .metadata.name, lifecycle: .spec.lifecycle}'
# Find all services consuming a specific API
curl -s "https://backstage.internal/api/catalog/entities?filter=kind=component,spec.consumesApis=payment-api" \
-H "Authorization: Bearer $BACKSTAGE_TOKEN" | jq '.[] | .metadata.name'
# Get component details with relations
curl -s "https://backstage.internal/api/catalog/entities/by-name/component/default/payment-service" \
-H "Authorization: Bearer $BACKSTAGE_TOKEN" | jq '{
name: .metadata.name,
owner: .spec.owner,
apis: .spec.providesApis,
dependencies: .spec.dependsOn
}'
Golden Path Templates
Golden paths are opinionated, pre-configured templates that encode best practices. They give teams a paved road to production.
Backstage Software Template
apiVersion: scaffolder.backstage.io/v1beta3
kind: Template
metadata:
name: spring-boot-service
title: Spring Boot Microservice
description: Creates a production-ready Spring Boot service with CI/CD, monitoring, and database
tags:
- java
- spring-boot
- recommended
spec:
owner: platform-team
type: service
parameters:
- title: Service Details
required:
- name
- owner
- description
properties:
name:
title: Service Name
type: string
pattern: '^[a-z][a-z0-9-]*$'
ui:autofocus: true
owner:
title: Owner Team
type: string
ui:field: OwnerPicker
ui:options:
catalogFilter:
kind: Group
description:
title: Description
type: string
javaVersion:
title: Java Version
type: string
default: '21'
enum: ['17', '21']
- title: Infrastructure
properties:
database:
title: Database
type: string
default: postgresql
enum: [postgresql, mysql, none]
cacheLayer:
title: Cache Layer
type: string
default: none
enum: [redis, none]
messageBroker:
title: Message Broker
type: string
default: none
enum: [kafka, rabbitmq, none]
steps:
- id: fetch-template
name: Fetch Skeleton
action: fetch:template
input:
url: ./skeleton
values:
name: ${{ parameters.name }}
owner: ${{ parameters.owner }}
description: ${{ parameters.description }}
javaVersion: ${{ parameters.javaVersion }}
database: ${{ parameters.database }}
- id: create-repo
name: Create Repository
action: publish:github
input:
repoUrl: github.com?owner=acme-corp&repo=${{ parameters.name }}
description: ${{ parameters.description }}
defaultBranch: main
protectDefaultBranch: true
requireCodeOwnerReviews: true
- id: register-catalog
name: Register in Catalog
action: catalog:register
input:
repoContentsUrl: ${{ steps['create-repo'].output.repoContentsUrl }}
catalogInfoPath: /catalog-info.yaml
- id: create-argocd-app
name: Create ArgoCD Application
action: argocd:create-resources
input:
appName: ${{ parameters.name }}
repoUrl: ${{ steps['create-repo'].output.remoteUrl }}
output:
links:
- title: Repository
url: ${{ steps['create-repo'].output.remoteUrl }}
- title: Service in Catalog
url: ${{ steps['register-catalog'].output.entityRef }}
- title: CI/CD Pipeline
url: ${{ steps['create-repo'].output.remoteUrl }}/actions
Golden Path Coverage Matrix
| Category |
What the Golden Path Provides |
Without Golden Path |
| Repository |
Pre-configured with CI/CD, linting, CODEOWNERS |
Manual setup, inconsistent configs |
| CI/CD |
Working pipeline from day one |
Copy-paste from other repos, broken configs |
| Observability |
Dashboards, alerts, SLOs pre-configured |
No monitoring until first incident |
| Security |
Dependency scanning, SAST, secrets detection |
Added retroactively (if at all) |
| Documentation |
ADR template, README structure, API docs |
Empty README, no docs |
| Infrastructure |
Terraform modules, Kubernetes manifests |
Hand-crafted YAML, drift between envs |
| Testing |
Test framework, coverage gates, fixtures |
Ad-hoc test setup, no coverage requirements |
Developer Experience Metrics (SPACE Framework)
Measure platform effectiveness using the SPACE framework. Never rely on a single dimension.
| Dimension |
What It Measures |
Example Metrics |
Collection Method |
| Satisfaction |
How developers feel about the platform |
NPS score, satisfaction survey (1-5) |
Quarterly survey |
| Performance |
Outcome of developer work |
Deployment frequency, change failure rate |
DORA metrics pipeline |
| Activity |
Volume of actions |
Scaffolding requests, API calls, portal visits |
Platform telemetry |
| Communication |
Quality of collaboration |
Time to first response on platform support |
Ticketing system |
| Efficiency |
Flow and minimal friction |
Time from commit to deploy, onboarding time |
Pipeline metrics |
Key DevEx Metrics Dashboard
# Platform DevEx Metrics -- collected via platform telemetry
metrics:
onboarding:
time_to_first_deploy:
target: "< 2 hours"
description: "Time from new hire to first successful deployment"
source: "scaffolder + pipeline timestamps"
time_to_first_commit:
target: "< 4 hours"
description: "Time from repo creation to first merged commit"
source: "github events"
self_service:
template_adoption_rate:
target: "> 80%"
description: "Percentage of new services using golden path templates"
source: "backstage scaffolder logs"
self_service_resolution_rate:
target: "> 70%"
description: "Percentage of requests resolved without platform team intervention"
source: "support tickets vs portal actions"
reliability:
platform_availability:
target: "99.9%"
description: "Uptime of developer portal, CI/CD, and artifact registry"
source: "synthetic monitoring"
mean_time_to_recovery:
target: "< 30 minutes"
description: "Time to restore platform services after incident"
source: "incident management system"
delivery:
deployment_frequency:
target: "multiple per day per team"
description: "How often teams deploy to production"
source: "deployment pipeline events"
lead_time_for_changes:
target: "< 1 day"
description: "Time from commit to production"
source: "git + pipeline timestamps"
Self-Service Portal Workflow
Request Flow Architecture
Developer submits request via Portal UI
|
v
Request Validation (schema check, policy check)
|
v
Approval Gate (if required by policy)
| |
| auto | manual
v v
Orchestration Engine (executes workflow)
|
+---> Provision Infrastructure (Terraform/Crossplane)
+---> Configure CI/CD (GitHub Actions / ArgoCD)
+---> Register in Catalog (Backstage)
+---> Set Up Monitoring (Grafana / PagerDuty)
+---> Notify Team (Slack / Email)
|
v
Verification (health checks, smoke tests)
|
v
Developer notified -- ready to use
Self-Service Capability Matrix
| Capability |
Automation Level |
Approval Required |
Typical Time |
| Create new service |
Fully automated |
No |
5 minutes |
| Provision database |
Fully automated |
No (dev/staging), Yes (prod) |
10 minutes |
| Add CI/CD pipeline |
Fully automated |
No |
2 minutes |
| Request cloud credentials |
Semi-automated |
Yes (security review) |
1 hour |
| Create new environment |
Fully automated |
No (non-prod), Yes (prod) |
15 minutes |
| Add monitoring/alerts |
Fully automated |
No |
5 minutes |
| Resize infrastructure |
Semi-automated |
Yes (cost review > threshold) |
30 minutes |
| Decommission service |
Automated with safeguards |
Yes (owner confirmation) |
10 minutes |
Platform Engineering Maturity Model
| Level |
Name |
Characteristics |
Capabilities |
| 0 |
Ad Hoc |
No platform, tribal knowledge |
Teams manage their own infra |
| 1 |
Reactive |
Shared scripts, wiki docs |
Basic CI/CD, manual provisioning |
| 2 |
Standardized |
Golden paths, basic portal |
Service templates, catalog, basic self-service |
| 3 |
Optimized |
Full IDP, metrics-driven |
Self-service everything, DevEx metrics, policy-as-code |
| 4 |
Strategic |
Platform as product, innovation |
API-first platform, marketplace, continuous feedback |
Maturity Assessment Checklist
Level 1 --> Level 2:
[x] Service catalog exists and is maintained
[x] At least 3 golden path templates available
[x] Basic developer portal deployed
[x] CI/CD standardized across teams
Level 2 --> Level 3:
[x] Self-service for >80% of common requests
[x] SPACE metrics collected and reviewed monthly
[x] Policy-as-code enforced (not advisory)
[x] Platform team has dedicated product manager
[x] Internal SLOs defined for platform services
Level 3 --> Level 4:
[x] API-first platform (all capabilities programmable)
[x] Internal developer marketplace for plugins/extensions
[x] Continuous developer experience research program
[x] Platform economics model (cost attribution per team)
[x] Platform contributes to organizational strategy
Anti-Patterns
| Anti-Pattern |
Problem |
Fix |
| Build it and they will come |
No adoption without developer input |
Treat platform as product; user research before building |
| Ticket-ops disguised as platform |
Self-service portal that just creates tickets |
Automate end-to-end; tickets are a smell, not a solution |
| Mandating platform use |
Forced adoption breeds resentment and workarounds |
Make the golden path the easiest path, not the only path |
| One-size-fits-all templates |
Overly rigid templates that don't fit team needs |
Composable templates with sensible defaults and escape hatches |
| No feedback loops |
Platform team builds in isolation |
Regular surveys, office hours, embedded rotations with teams |
| Ignoring developer experience |
Technically correct but painful to use |
Measure DevEx metrics, optimize for developer happiness |
| Platform team as bottleneck |
All changes go through platform team |
Self-service with guardrails; teams should not wait on platform |
| Over-abstracting too early |
Complex abstraction layers before understanding needs |
Start with concrete solutions, abstract when patterns emerge |
| Neglecting documentation |
Powerful platform nobody knows how to use |
Docs-as-code, TechDocs in Backstage, examples for everything |
| No platform SLOs |
Platform reliability treated as best-effort |
Define and publish SLOs; platform is a product with SLAs |
| Shadow platforms |
Teams build their own tooling around the platform |
Understand why and address gaps; shadow platforms reveal unmet needs |
| Gold plating the portal |
Spending months on portal UI before delivering value |
Ship incrementally; a working CLI beats a beautiful but empty portal |
Platform Engineering Checklist
1---2name: platform-engineering-33description: Provides platform engineering best practices for Internal Developer Platforms (IDPs), golden paths, service catalogs, and developer experience. Use when building developer platforms, configuring Backstage, designing self-service workflows, or when user mentions 'platform engineering', 'backstage', 'golden path', 'IDP', 'developer portal', 'service catalog', 'DevEx', 'platform team', 'self-service'.4---56# Platform Engineering78Best practices for building Internal Developer Platforms (IDPs) that reduce cognitive load, accelerate delivery, and create golden paths for development teams.910## IDP Architecture Layers1112A well-designed IDP separates concerns into distinct layers. Each layer abstracts complexity from the one above it.1314```15Developer Interface (Portal / CLI / API)16 |17 Orchestration Layer (Workflows, Templates, Scaffolding)18 |19 Integration Layer (APIs, Plugins, Connectors)20 |21 Resource Layer (Infrastructure, Services, Tools)22```2324| Layer | Purpose | Components | Owned By |25|-------|---------|------------|----------|26| Developer Interface | Self-service entry point | Backstage portal, CLI tools, API gateway | Platform team |27| Orchestration | Workflow automation, templating | Scaffolder, Terraform modules, Crossplane | Platform team |28| Integration | Connect tools and services | Backstage plugins, API adapters, webhooks | Platform + tool owners |29| Resource | Actual infrastructure and services | Kubernetes, databases, CI/CD, monitoring | Infrastructure team |30| Governance | Policy enforcement and compliance | OPA, Kyverno, cost policies, security scans | Security + platform team |3132## Platform Team Topology and Responsibilities3334### Team Structure3536| Role | Responsibility | Focus Area |37|------|---------------|------------|38| Platform Product Manager | Roadmap, prioritization, user research | Developer needs, adoption metrics |39| Platform Engineer | IDP core, golden paths, automation | Infrastructure abstraction, tooling |40| Developer Advocate | Documentation, onboarding, feedback loops | DevEx, training, communication |41| SRE/Reliability Lead | Platform reliability, SLOs, incident response | Uptime, performance, observability |42| Security Engineer | Policy-as-code, compliance automation | Guardrails, scanning, access control |4344### Interaction Model4546```47Stream-Aligned Teams (consumers)48 |49 | self-service requests50 v51Platform Team (enablers)52 |53 | golden paths, templates, APIs54 v55Infrastructure / Cloud (resources)56```5758Platform teams operate as **enabling teams** (Team Topologies model). They reduce cognitive load on stream-aligned teams by providing curated, opinionated abstractions.5960## Backstage: Service Catalog and Developer Portal6162### catalog-info.yaml -- Service Registration6364Every service registers itself in the catalog via a `catalog-info.yaml` at the repo root.6566```yaml67apiVersion: backstage.io/v1alpha168kind: Component69metadata:70 name: payment-service71 description: Handles payment processing and refunds72 annotations:73 github.com/project-slug: acme-corp/payment-service74 backstage.io/techdocs-ref: dir:.75 pagerduty.com/service-id: P1234ABC76 grafana/dashboard-selector: "payment-service"77 tags:78 - java79 - spring-boot80 - payments81 links:82 - url: https://grafana.internal/d/payments83 title: Dashboard84 icon: dashboard85 - url: https://runbooks.internal/payments86 title: Runbook87 icon: docs88spec:89 type: service90 lifecycle: production91 owner: team-payments92 system: checkout-system93 providesApis:94 - payment-api95 consumesApis:96 - inventory-api97 - notification-api98 dependsOn:99 - resource:payments-db100 - component:auth-service101102---103apiVersion: backstage.io/v1alpha1104kind: API105metadata:106 name: payment-api107 description: Payment processing REST API108spec:109 type: openapi110 lifecycle: production111 owner: team-payments112 system: checkout-system113 definition:114 $text: ./openapi.yaml115```116117### Service Catalog API -- Querying Components118119```bash120# List all components owned by a team121curl -s "https://backstage.internal/api/catalog/entities?filter=kind=component,spec.owner=team-payments" \122 -H "Authorization: Bearer $BACKSTAGE_TOKEN" | jq '.[] | {name: .metadata.name, lifecycle: .spec.lifecycle}'123124# Find all services consuming a specific API125curl -s "https://backstage.internal/api/catalog/entities?filter=kind=component,spec.consumesApis=payment-api" \126 -H "Authorization: Bearer $BACKSTAGE_TOKEN" | jq '.[] | .metadata.name'127128# Get component details with relations129curl -s "https://backstage.internal/api/catalog/entities/by-name/component/default/payment-service" \130 -H "Authorization: Bearer $BACKSTAGE_TOKEN" | jq '{131 name: .metadata.name,132 owner: .spec.owner,133 apis: .spec.providesApis,134 dependencies: .spec.dependsOn135 }'136```137138## Golden Path Templates139140Golden paths are opinionated, pre-configured templates that encode best practices. They give teams a paved road to production.141142### Backstage Software Template143144```yaml145apiVersion: scaffolder.backstage.io/v1beta3146kind: Template147metadata:148 name: spring-boot-service149 title: Spring Boot Microservice150 description: Creates a production-ready Spring Boot service with CI/CD, monitoring, and database151 tags:152 - java153 - spring-boot154 - recommended155spec:156 owner: platform-team157 type: service158159 parameters:160 - title: Service Details161 required:162 - name163 - owner164 - description165 properties:166 name:167 title: Service Name168 type: string169 pattern: '^[a-z][a-z0-9-]*$'170 ui:autofocus: true171 owner:172 title: Owner Team173 type: string174 ui:field: OwnerPicker175 ui:options:176 catalogFilter:177 kind: Group178 description:179 title: Description180 type: string181 javaVersion:182 title: Java Version183 type: string184 default: '21'185 enum: ['17', '21']186187 - title: Infrastructure188 properties:189 database:190 title: Database191 type: string192 default: postgresql193 enum: [postgresql, mysql, none]194 cacheLayer:195 title: Cache Layer196 type: string197 default: none198 enum: [redis, none]199 messageBroker:200 title: Message Broker201 type: string202 default: none203 enum: [kafka, rabbitmq, none]204205 steps:206 - id: fetch-template207 name: Fetch Skeleton208 action: fetch:template209 input:210 url: ./skeleton211 values:212 name: ${{ parameters.name }}213 owner: ${{ parameters.owner }}214 description: ${{ parameters.description }}215 javaVersion: ${{ parameters.javaVersion }}216 database: ${{ parameters.database }}217218 - id: create-repo219 name: Create Repository220 action: publish:github221 input:222 repoUrl: github.com?owner=acme-corp&repo=${{ parameters.name }}223 description: ${{ parameters.description }}224 defaultBranch: main225 protectDefaultBranch: true226 requireCodeOwnerReviews: true227228 - id: register-catalog229 name: Register in Catalog230 action: catalog:register231 input:232 repoContentsUrl: ${{ steps['create-repo'].output.repoContentsUrl }}233 catalogInfoPath: /catalog-info.yaml234235 - id: create-argocd-app236 name: Create ArgoCD Application237 action: argocd:create-resources238 input:239 appName: ${{ parameters.name }}240 repoUrl: ${{ steps['create-repo'].output.remoteUrl }}241242 output:243 links:244 - title: Repository245 url: ${{ steps['create-repo'].output.remoteUrl }}246 - title: Service in Catalog247 url: ${{ steps['register-catalog'].output.entityRef }}248 - title: CI/CD Pipeline249 url: ${{ steps['create-repo'].output.remoteUrl }}/actions250```251252### Golden Path Coverage Matrix253254| Category | What the Golden Path Provides | Without Golden Path |255|----------|------------------------------|---------------------|256| Repository | Pre-configured with CI/CD, linting, CODEOWNERS | Manual setup, inconsistent configs |257| CI/CD | Working pipeline from day one | Copy-paste from other repos, broken configs |258| Observability | Dashboards, alerts, SLOs pre-configured | No monitoring until first incident |259| Security | Dependency scanning, SAST, secrets detection | Added retroactively (if at all) |260| Documentation | ADR template, README structure, API docs | Empty README, no docs |261| Infrastructure | Terraform modules, Kubernetes manifests | Hand-crafted YAML, drift between envs |262| Testing | Test framework, coverage gates, fixtures | Ad-hoc test setup, no coverage requirements |263264## Developer Experience Metrics (SPACE Framework)265266Measure platform effectiveness using the SPACE framework. Never rely on a single dimension.267268| Dimension | What It Measures | Example Metrics | Collection Method |269|-----------|-----------------|-----------------|-------------------|270| **S**atisfaction | How developers feel about the platform | NPS score, satisfaction survey (1-5) | Quarterly survey |271| **P**erformance | Outcome of developer work | Deployment frequency, change failure rate | DORA metrics pipeline |272| **A**ctivity | Volume of actions | Scaffolding requests, API calls, portal visits | Platform telemetry |273| **C**ommunication | Quality of collaboration | Time to first response on platform support | Ticketing system |274| **E**fficiency | Flow and minimal friction | Time from commit to deploy, onboarding time | Pipeline metrics |275276### Key DevEx Metrics Dashboard277278```yaml279# Platform DevEx Metrics -- collected via platform telemetry280metrics:281 onboarding:282 time_to_first_deploy:283 target: "< 2 hours"284 description: "Time from new hire to first successful deployment"285 source: "scaffolder + pipeline timestamps"286287 time_to_first_commit:288 target: "< 4 hours"289 description: "Time from repo creation to first merged commit"290 source: "github events"291292 self_service:293 template_adoption_rate:294 target: "> 80%"295 description: "Percentage of new services using golden path templates"296 source: "backstage scaffolder logs"297298 self_service_resolution_rate:299 target: "> 70%"300 description: "Percentage of requests resolved without platform team intervention"301 source: "support tickets vs portal actions"302303 reliability:304 platform_availability:305 target: "99.9%"306 description: "Uptime of developer portal, CI/CD, and artifact registry"307 source: "synthetic monitoring"308309 mean_time_to_recovery:310 target: "< 30 minutes"311 description: "Time to restore platform services after incident"312 source: "incident management system"313314 delivery:315 deployment_frequency:316 target: "multiple per day per team"317 description: "How often teams deploy to production"318 source: "deployment pipeline events"319320 lead_time_for_changes:321 target: "< 1 day"322 description: "Time from commit to production"323 source: "git + pipeline timestamps"324```325326## Self-Service Portal Workflow327328### Request Flow Architecture329330```331Developer submits request via Portal UI332 |333 v334Request Validation (schema check, policy check)335 |336 v337Approval Gate (if required by policy)338 | |339 | auto | manual340 v v341Orchestration Engine (executes workflow)342 |343 +---> Provision Infrastructure (Terraform/Crossplane)344 +---> Configure CI/CD (GitHub Actions / ArgoCD)345 +---> Register in Catalog (Backstage)346 +---> Set Up Monitoring (Grafana / PagerDuty)347 +---> Notify Team (Slack / Email)348 |349 v350Verification (health checks, smoke tests)351 |352 v353Developer notified -- ready to use354```355356### Self-Service Capability Matrix357358| Capability | Automation Level | Approval Required | Typical Time |359|------------|-----------------|-------------------|--------------|360| Create new service | Fully automated | No | 5 minutes |361| Provision database | Fully automated | No (dev/staging), Yes (prod) | 10 minutes |362| Add CI/CD pipeline | Fully automated | No | 2 minutes |363| Request cloud credentials | Semi-automated | Yes (security review) | 1 hour |364| Create new environment | Fully automated | No (non-prod), Yes (prod) | 15 minutes |365| Add monitoring/alerts | Fully automated | No | 5 minutes |366| Resize infrastructure | Semi-automated | Yes (cost review > threshold) | 30 minutes |367| Decommission service | Automated with safeguards | Yes (owner confirmation) | 10 minutes |368369## Platform Engineering Maturity Model370371| Level | Name | Characteristics | Capabilities |372|-------|------|----------------|--------------|373| 0 | Ad Hoc | No platform, tribal knowledge | Teams manage their own infra |374| 1 | Reactive | Shared scripts, wiki docs | Basic CI/CD, manual provisioning |375| 2 | Standardized | Golden paths, basic portal | Service templates, catalog, basic self-service |376| 3 | Optimized | Full IDP, metrics-driven | Self-service everything, DevEx metrics, policy-as-code |377| 4 | Strategic | Platform as product, innovation | API-first platform, marketplace, continuous feedback |378379### Maturity Assessment Checklist380381```382Level 1 --> Level 2:383 [x] Service catalog exists and is maintained384 [x] At least 3 golden path templates available385 [x] Basic developer portal deployed386 [x] CI/CD standardized across teams387388Level 2 --> Level 3:389 [x] Self-service for >80% of common requests390 [x] SPACE metrics collected and reviewed monthly391 [x] Policy-as-code enforced (not advisory)392 [x] Platform team has dedicated product manager393 [x] Internal SLOs defined for platform services394395Level 3 --> Level 4:396 [x] API-first platform (all capabilities programmable)397 [x] Internal developer marketplace for plugins/extensions398 [x] Continuous developer experience research program399 [x] Platform economics model (cost attribution per team)400 [x] Platform contributes to organizational strategy401```402403## Anti-Patterns404405| Anti-Pattern | Problem | Fix |406|--------------|---------|-----|407| Build it and they will come | No adoption without developer input | Treat platform as product; user research before building |408| Ticket-ops disguised as platform | Self-service portal that just creates tickets | Automate end-to-end; tickets are a smell, not a solution |409| Mandating platform use | Forced adoption breeds resentment and workarounds | Make the golden path the easiest path, not the only path |410| One-size-fits-all templates | Overly rigid templates that don't fit team needs | Composable templates with sensible defaults and escape hatches |411| No feedback loops | Platform team builds in isolation | Regular surveys, office hours, embedded rotations with teams |412| Ignoring developer experience | Technically correct but painful to use | Measure DevEx metrics, optimize for developer happiness |413| Platform team as bottleneck | All changes go through platform team | Self-service with guardrails; teams should not wait on platform |414| Over-abstracting too early | Complex abstraction layers before understanding needs | Start with concrete solutions, abstract when patterns emerge |415| Neglecting documentation | Powerful platform nobody knows how to use | Docs-as-code, TechDocs in Backstage, examples for everything |416| No platform SLOs | Platform reliability treated as best-effort | Define and publish SLOs; platform is a product with SLAs |417| Shadow platforms | Teams build their own tooling around the platform | Understand why and address gaps; shadow platforms reveal unmet needs |418| Gold plating the portal | Spending months on portal UI before delivering value | Ship incrementally; a working CLI beats a beautiful but empty portal |419420## Platform Engineering Checklist421422- [ ] Platform team established with clear product ownership423- [ ] Developer portal deployed (Backstage or equivalent)424- [ ] Service catalog populated with all production services425- [ ] At least 3 golden path templates available and documented426- [ ] Self-service provisioning for common infrastructure (databases, queues, caches)427- [ ] CI/CD pipelines standardized and available via templates428- [ ] Observability stack integrated (dashboards auto-created with new services)429- [ ] Security scanning built into golden paths (not bolted on after)430- [ ] DevEx metrics defined and collected (SPACE framework dimensions)431- [ ] Feedback mechanism active (surveys, office hours, Slack channel)432- [ ] Platform SLOs defined and monitored433- [ ] Documentation maintained in developer portal (TechDocs)434- [ ] Onboarding time measured and optimized (target: first deploy < 2 hours)435- [ ] Cost visibility per team/service available through platform436- [ ] Platform roadmap published and informed by developer feedback437- [ ] Escape hatches documented for when golden paths don't fit438- [ ] Platform reliability meets or exceeds published SLOs