Common Pitfalls in Platform Engineering
- Platform engineering success is not achieved through code, but rather through factors beyond technical implementation 0s.
- Many organizations begin their platform engineering journey by deploying a developer portal, such as Backstage, because it is perceived as a "sexy" UI popularized by Spotify 1m15s.
- A common challenge with tools like Backstage is that they are empty upon installation, requiring significant time, financial investment, and multiple engineers to configure before they provide any perceived value 2m15s.
- Internal development platforms often fail because they are treated as projects with deadlines rather than as ongoing products 2m35s.
- Primary reasons for platform failure include poor adoption rates, unmet management expectations, a lack of alignment with the needs of development teams, and a failure to deliver promised value 2m35s.
- Platforms frequently lack a clearly defined purpose or a high-level vision, such as specific goals to increase development speed, reduce costs, or address technical debt 3m5s.
- Over-engineering is a common issue where tools are integrated into a platform simply because they are popular or were seen at conferences, rather than because they address a verified need 3m25s.
- A practice described as "random engineering" occurs when engineers select tools based on personal preference or limited exposure rather than a systematic evaluation of available alternatives 3m50s.
The Dangers of Infrastructure-First Thinking
- Infrastructure-first thinking is identified as a primary cause of failure in platform engineering, largely because many practitioners transition into the field from cloud or infrastructure backgrounds. 0s
- Infrastructure-first approaches are characterized by being technology-centered, tool-oriented, and focused on architectural decisions or designing from the infrastructure layer upward. 15s
- A significant drawback of this mindset is that it prioritizes the technology stack and the creation of "awesome" infrastructure over the needs and perspectives of the end users. 35s
- Many engineers lack the natural inclination or professional experience to engage in demand management, which involves interviewing users to identify and challenge their actual requirements. 55s
Organizational Challenges and Industry Hype
- The "shift left" movement is criticized for effectively pushing complex responsibilities—such as CI/CD, observability, tracing, security, and compliance—onto developers rather than solving the underlying problems. 1m20s
- Organizations often struggle with scaling these processes because multiple teams, such as those handling sales applications, web portals, or backend logistics, frequently operate in silos. 1m45s
- Large enterprises experience cyclical waves of harmonization where teams attempt to standardize tools and approaches, only to have these efforts disrupted by reorganizations. 2m0s
- These organizational shifts often lead to high staff turnover, resulting in the loss of institutional knowledge and the introduction of new personnel who may not be aligned with previous standards. 2m15s
- Industry hypes, such as the focus on IoT, cloud-native containerization, and current trends in AI and machine learning, often lead to wasted resources when technologies are adopted without a clear business case. 2m25s
- Many initiatives driven by industry trends fail to deliver success or return on investment because teams often attempt to implement these technologies in their own fragmented ways. 2m45s
Sustainability and Product Mindset
- Organizational environments frequently undergo cycles of disruption and harmonization due to shifting trends and hypes, leading to recurring periods of duplicated effort 0s.
- Operational tasks, release management, user management, and internal communication regarding new features often result in the creation of temporary, "patch" solutions 15s.
- While terminology has evolved from sysadmins to DevOps teams, the underlying structure often remains problematic, with a heavy emphasis on operations and high-pressure, on-call requirements for staff 42s.
- Patch solutions are inherently unsustainable and represent a cycle of breaking and rebuilding that organizations should avoid 1m15s.
- Platform engineering is distinct from DevOps and aims to provide a product that developers choose to use voluntarily, rather than a system that must be enforced by management 1m30s.
- Enforcing the use of a platform through mandates or threats is ineffective and often triggers resistance from internal users 1m50s.
- A successful platform requires a product mindset that identifies and understands the needs of its internal customers, which include developers, security personnel, compliance officers, product owners, and business stakeholders 2m15s.
- Platforms can provide value to non-technical stakeholders, such as VPs, by delivering automated reports on usage and deployment metrics directly to their inboxes 2m35s.
- Infrastructure that lacks user adoption and a product-focused approach is often perceived as unappealing and fails to become an indispensable tool for the organization 2m55s.
Balancing Platform Dimensions
- Achieving high adoption rates for a platform is a complex challenge that requires a strategic perspective, such as redefining the iron triangle 3m15s.
- Platform engineering success relies on balancing three core dimensions: feasibility (building the thing right), desirability (building what the customer wants), and viability (building the right thing) 0s.
- These three dimensions function like an iron triangle where pulling in one direction often necessitates sacrificing another, as the balance is constrained by factors such as budget, time, and available personnel 15s.
- While large platforms may eventually reach a threshold where they can fulfill all three dimensions simultaneously, it is generally impossible to build a perfect platform before reaching that scale 42s.
Establishing Guiding Principles
- A lack of guiding frameworks often leads to decisions being made based on the subjective preferences of individuals rather than strategic alignment 1m25s.
- Establishing clear principles provides a necessary framework for decision-making, helping to avoid issues like the arbitrary addition of tools that do not solve actual problems 1m45s.
- Eliminating waste is a critical principle; failing to focus on this from the beginning often leads to overspending, which consumes budgets that could otherwise be used for further innovation 2m6s.
- Principles can be framed as "we will" statements to foster team unity and positive framing, such as focusing on driving transparency for resource consumption rather than simply "eliminating waste" 2m45s.
- Alternatively, principles can be framed as "you must" guidelines that provide direction through best practices and transparency, such as offering fine-grained access to operational data without being overly prescriptive 3m15s.
- Principles should be distinguished from directives, as directives are mandatory instructions that do not allow for the personal alignment required by a principle 0s.
- A balanced approach to platform engineering involves mixing "we will" principles, which align with team interests and unity, with "we must" principles, which are used to guide people in a specific direction 15s.
Drivers of Platform Development
- Strategic drivers often originate from management or external partnerships, sometimes forcing teams to adopt tools that are not ideal for the platform but are required for business or financial reasons 42s.
- Personal drivers, including the moral or ethical concerns of team members, can significantly impact platform development 1m15s.
- When team members refuse to work with specific tools due to personal or political objections, it can harm the development and motivation of the team 1m35s.
- Market drivers include external factors such as changes in open-source licensing, which can force shifts in platform direction and cause operational difficulties 2m15s.
- Regional sovereignty requirements, such as those currently seen in Europe and Germany, can influence technology choices, sometimes leading to debates about the geographic location of open-source contributors 2m35s.
- Data, insights, and analyst reports often serve as drivers for management decisions, though these reports are frequently influenced by the financial relationships between vendors and reporting organizations 3m5s.
Defining Platform Purpose
- Many platforms fail to define a clear purpose or statement explaining why the platform is necessary, which is a critical missing component in many observed projects 3m35s.
- Platform engineering initiatives often face skepticism regarding their business value, cost, and technical feasibility, leading to concerns that they may be perceived as unrealistic "dreamland" projects 0s.
- Defining a clear purpose for the platform is essential, as it provides the necessary justification for why the platform exists and who it serves 35s.
- The "platform engineering purpose canvas" is a tool designed to help organizations identify their purpose by analyzing roles, actions, targets, and organizational values 55s.
- Many employees and companies lack a mutual understanding of their respective values, which can hinder the alignment necessary for successful platform development 1m15s.
- A product mindset, which includes identifying stakeholders and assessing strengths and development areas, is a critical component of the purpose-defining process 1m30s.
- Investing time—such as one or two days—with architects and product owners to discuss the platform's purpose can prevent the development of tools that ultimately go unused 1m45s.
- Failing to define a clear purpose and strategy can lead to a "lose-lose-lose" scenario where the platform is abandoned, the customer loses money, and the responsible parties face negative professional consequences 2m0s.
- Establishing a clear purpose enables better communication with internal teams and external stakeholders, fostering a "win-win-win" outcome 2m20s.
Measuring Platform Success
- A common failure in platform engineering is the lack of measurement, specifically the failure to establish key performance indicators (KPIs) before the platform is implemented 2m35s.
- Measuring platform performance only after implementation makes it impossible to objectively compare success against a baseline, leaving organizations to rely on subjective feelings 2m50s.
- The recommended approach is to start with a "simplest viable platform" and introduce new features incrementally, ensuring continuous improvement 3m5s.
- Measuring metrics like monthly active users provides concrete data on platform success 3m25s.
- Centralizing resources, such as documentation, can significantly improve user adoption by replacing fragmented, difficult-to-find information with a single, accessible location 3m40s.
- Achieving transparency regarding platform usage, such as tracking engagement with a service catalog, is vital for understanding how the platform is being utilized 4m5s.
Metrics and Performance Evaluation
- Platform engineering requires a shift in mindset from a purely engineering perspective to a product-oriented approach that prioritizes understanding user needs rather than simply releasing features 0s.
- Time to market remains a relevant performance metric for internal platforms, defined as the duration between identifying a problem and deploying an internal solution 15s.
- Shadow IT, such as unauthorized servers or the use of private credit cards for cloud services, serves as a key indicator of whether a platform is desired by users 35s.
- A successful platform can lead to a significant reduction in shadow IT expenses, as demonstrated by a large telecommunications customer that decreased external IT spending by providing internal solutions to identified problems 55s.
- DORA metrics are often misused by management as a tool to pressure teams to increase deployment speed or feature output, which can negatively impact code quality 1m25s.
- While DORA metrics can help raise awareness regarding the need for improvements, they are frequently counterproductive when used to evaluate team performance, as they often ignore the complexities of development work 2m5s.
- SPACE metrics are presented as a more comprehensive alternative to DORA, as they evaluate satisfaction, performance, activity, communication, and efficiency across individuals, teams, and systems 2m25s.
- Metrics like code review velocity can be misleading or ineffective if applied without context, as they fail to account for the complexity or volume of the work being reviewed 2m55s.
- Metrics that measure code volume and review time are inherently correlated, as shipping more code naturally requires more time for review 0s.
- The SPACE framework is a useful alternative for measurement, though it requires significant time to set up properly and must be adjusted periodically 15s.
- Effective measurement requires dedicated team members who continuously evaluate if the current KPIs remain relevant or if they should be dropped or postponed 25s.
- DevEx metrics are considered highly helpful because they focus on the flow state and the overall experience of platform end users 45s.
- Quantitative metrics for flow state include tracking reopened commits, change requests, and the number of commits per pull request 58s.
- Quantitative data must be supplemented by personal feedback gathered through direct interaction with users to determine if they feel comfortable and if the platform is truly helpful 1m5s.
- Relying solely on numbers can be misleading, as two metrics that appear correlated may only be linked by coincidence 1m30s.
Cognitive Load and Feedback Loops
- Cognitive load is defined as any additional task or requirement that complicates a person's work and distracts them from solving their primary problem 1m45s.
- Frequent context switching, such as attending multiple short meetings, significantly increases cognitive load because the brain continues processing information long after a disruption occurs 2m5s.
- It can take up to one hour for the brain to process information following a five-minute disruption 2m20s.
- Transitions between meetings are often difficult because the brain struggles to immediately shift focus to a completely different topic, often carrying the emotional state or stress from a previous interaction into the next 2m40s.
- High cognitive load prevents individuals from moving forward effectively because the brain remains stuck on previous tasks or problems 3m15s.
- Feedback loops regarding platform performance can be established through mechanical methods, such as using thumbs-up features after deployments, or through subjective interviews 3m25s.
- Written feedback is considered a valuable source of information for improving platform development 3m45s.
User Research and Process Improvement
- Technical debt is a primary factor contributing to the failure of platform adoptions 0s.
- Effective platform development requires user research, which involves engaging directly with the intended users to understand their needs before building 15s.
- When conducting user research, it is recommended to ask five specific questions: what tools are currently in use, what the current workflow looks like, what the perceived migration effort is, what existing gaps are present, and what the ideal workflow would be 35s.
- The focus of user research should be on the functionality required by users rather than the specific tools themselves, as the goal is to gain a comprehensive understanding of their day-to-day work 55s.
- Many organizational problems stem from people and processes rather than technology, meaning a platform cannot solve every operational issue 1m35s.
- Implementing effective feedback loops is difficult, and many companies fail to conduct them properly by relying on infrequent, annual surveys 1m55s.
- Feedback should be collected by the team responsible for the product, rather than by a distant organizational department, to ensure the data is relevant and actionable 2m25s.
- Relying on one-time feedback snapshots is unreliable because individual responses are often heavily influenced by a person's immediate, daily experience rather than a long-term assessment of the tools 2m45s.
- Continuous feedback collection is necessary to obtain a complete and accurate picture of user experience, as human perception of tool quality can fluctuate daily 3m5s.
- Effective feedback collection requires asking open-ended questions rather than simple yes or no inquiries to gain actionable insights into user requirements and future project needs 0s.
Managing Technical Debt and Evolution
- Managing technical debt is complicated by the sunk cost fallacy, where organizations feel compelled to maintain useless products due to accounting and tax-related pressures 35s.
- A recommended approach to minimizing technical debt is to start with a smallest viable platform and incrementally add features, though the nature of Kubernetes often leads to the accumulation of debt over time 1m15s.
- Platform engineering should be viewed as a living, malleable entity rather than a static structure, requiring constant maintenance, updates, and the occasional removal of tools 1m35s.
- Communicating the need to discard significant past investments to management is difficult, necessitating clear justifications based on whether a tool causes more harm than it provides value 2m0s.
- Establishing formal deprecation criteria from the beginning is essential for determining when a tool or platform component has become obsolete or detrimental 2m15s.
- Subjective user feedback, such as complaints regarding workflows, user interfaces, or command-line tools, should be used alongside metrics to identify when a component must be replaced 2m25s.
Culture and Long-Term Strategy
- Platform engineering is not defined by specific technologies like Kubernetes or Backstage, nor is it merely a rebranding of DevOps or infrastructure departments 2m55s.
- Platform engineering is fundamentally defined by the people involved, the company culture, and the purpose of the platform rather than its technical layers. 0s
- Successful platform engineering involves understanding and shaping company culture, often fostering inner source or open source practices to encourage external contributions. 0s
- A clear purpose and the execution of effective user research are essential components of a successful platform. 23s
- Technical debt management and the acceptance of constant change are necessary, as no platform component remains static indefinitely. 33s
- Building a community around a platform is a significant factor in increasing adoption rates and user satisfaction. 50s
- Inviting others to contribute to the platform and acknowledging their feedback creates value and fosters a collaborative environment. 1m5s
- Challenging Conway’s Law—the tendency for system design to mirror organizational communication structures—is difficult but effective. 1m23s
- Rather than allowing the platform to follow existing company processes, it is more effective to build an ideal workflow and adjust the company to follow that platform workflow. 1m30s
- Implementing a new platform workflow is a long-term journey that requires change management and can take several years. 1m35s
- Platform success is not determined by the specific technology or tools used, provided the tools function correctly and meet user needs. 1m45s
- Focusing exclusively on technical coolness while ignoring people, culture, and processes will lead to the failure of platform engineering activities. 2m5s








