All issues

Field Notes from Delivery · Issue 02

Nobody Owns the Pipeline. That's a DevOps Problem

ISSUE 02: This Week: The DEPENDENCY that is NOT Visible on the Planning Board. Your backlog can be immaculate, and your delivery can still stall on a shared pipeline nobody owns.

By Vikas Agarwal ·

Nobody Owns the Pipeline. That's a DevOps Problem

ISSUE 02: This Week: The DEPENDENCY that is NOT Visible on the Planning Board.

Your backlog can be immaculate, and your delivery can still stall on a shared pipeline nobody owns.

Four days into a two-week delivery cycle, one of my client's teams got completely stuck, blocked by a staging environment another team was using. It showed up on neither team's board.

That's the pattern I keep running into when clients tell me their dependency management is broken. It usually isn't at the level they're looking. Planning is clean. The dependency board is full of arrows between teams. What's missing could be the following dependencies

  1. a shared pipeline,
  2. a shared test environment,
  3. a shared release train,
  4. a feature flag another team owns and forgot to mention.

Did you forget to track a queue?

Why does this always keep happening

There's a fifty-year-old explanation. In 1968, Melvin Conway published a paper arguing that any organization that designs a system will produce a design whose structure copies the organization's own communication structure. If your teams are organized around features and your infrastructure is organized around shared services, the mismatch between those two structures is exactly where dependencies go to hide. Nobody put the staging environment on a board in my client's case. It simply wasn't in anyone's job description to.

Team Topologies, Matthew Skelton and Manuel Pais's 2019 book on organizing teams for fast flow, gives this problem an actual name. They split teams into four types, and the one that matters here is the platform team, whose job isn't to build features but to give other teams self-service access to what Skelton and Pais, borrowing a phrase from technologist Evan Bottcher, call "a compelling internal product." In their model, the healthiest relationship between a platform team and everyone consuming it is what they call X-as-a-Service: a well-defined interface with minimal back-and-forth. A dependency board manages something closer to their "Collaboration" mode. That works fine for new, ambiguous work between two teams. It stops working once the relationship is supposed to be routine.

The data backs this up too. Accelerate, the 2018 book built on years of DORA's State of DevOps research, found that loosely coupled teams and architecture, measured by whether a team can test and deploy without coordinating with another team, was one of the strongest predictors of software delivery performance the researchers identified, and the finding held even for teams running on decades-old mainframes, not only teams running microservices.

1CASE: SEGMENT, 2018
2Segment ran into a version of this at scale. Engineers had built 140-plus microservices around shared libraries, and under deadline pressure, individual teams would update the shared code for their own service without touching anyone else's copy. Within two years, engineer Alexandra Noonan wrote in the company's public engineering postmortem, the libraries had drifted so far apart that nobody could safely change shared code without risking someone else's service. Segment's actual fix was structural: consolidate the shared libraries into one monorepo with a single owned version of each dependency, so the quiet drift couldn't happen again.
3CASE: SPOTIFY
4Spotify eventually ran the platform-as-product idea for real. The internal developer portal its platform team built to manage exactly this class of problem, later open-sourced as Backstage, is now a Cloud Native Computing Foundation project that other companies run their own platforms on.

The pattern underneath all three comes down to the same missing thing: someone structurally responsible for the shared resource everyone depended on, regardless of which framework happened to be running the planning on top of it.

What could be done to fix it?

Platform-owning teams build a specific set of things to make self-service real.

Infrastructure and pipeline configuration should be version-controlled code, reviewed the same way as application code. The Cloud Native Computing Foundation's GitOps Working Group formalizes this as four principles:

  1. the system's state is declarative,
  2. versioned and immutable,
  3. pulled automatically rather than pushed,
  4. and continuously reconciled against what's actually running.

The shared pipeline stops being something anyone changes from their own laptop. Every change goes through a pull request with a named reviewer, same as any other code.

Ephemeral, on-demand environments remove the scarcity: instead of a single shared environment every team schedules around, each pull request or branch gets its own disposable copy.

For dependencies on one team's API and another team's use of it, a consuming team writes a contract describing exactly what it expects from an API, essentially recorded request and response examples, and that contract gets verified against the real provider inside the provider's own pipeline. Neither side needs the other's environment running to know whether they're still compatible. DORA's own guidance on loosely coupled teams calls out this practice by name, along with backward-compatible API versioning and the strangler-fig pattern for slowly migrating off systems that were never built to be decoupled.

Feature flags decouple deploying code from releasing it, so a team can merge to trunk continuously and turn a feature on whenever it's ready, independent of what anyone else is doing. A "golden path", Spotify's term for a pre-built, opinionated way to do something common like standing up a new service, replaces what one of their engineers called "rumour-driven development" (where the only way to learn how to do something was to go find whoever did it last time) with a route that doesn't require asking anyone.

Internal platform adoption is correlated with real gains. A 10% increase in team performance and an 8% boost in individual productivity, but it also found throughput and change stability both dipped when platform use was made mandatory across the board.

What change it can bring

No matter which framework sits on top of it, technical dependencies get planned at the environment and pipeline layer during whatever planning ritual the team already runs. The shared platform gets a backlog and a named owner. It is no longer nobody's job to say no.

I built ProvCraft's "Delivery Without Silos" program around exactly this gap, because it's the one that keeps showing up after a team is already certified and the org is still asking why the training didn't fix the dependency chaos. It doesn't, on its own. Training teaches the planning practices. Somebody still has to own the seams between teams, and no certificate hands that job to anyone.

If your dependency board looks complete and your delivery still gets blocked by something that isn't on it, that's usually where the real problem is.

Further reading

#ProvCraft #DevOps #PlatformEngineering #AgileDelivery #TeamTopologies #SoftwareDelivery

This issue was also published on LinkedIn. Join the conversation there or subscribe to Field Notes from Delivery.

More from Field Notes from Delivery