Skip to main content
DevOps10 min read

Keeping the Mainline Green Is Now a Platform Engineering Problem

As organizations consolidate services, libraries, and infrastructure into polyglot monorepos, CI reliability becomes a shared platform concern rather than a team-by-team workflow issue. Lessons from Uber’s approach to merge queues and mainline health show how engineering leaders can reduce integration debt with better queueing, ownership, and CI partitioning.

A red mainline used to be a local inconvenience. In a large polyglot monorepo, it can become an organization-wide delivery outage.

When hundreds or thousands of engineers share the same repository, every flaky test, slow build, missing owner, or incompatible dependency upgrade compounds. Keeping the main branch green stops being just a CI configuration problem and becomes a platform engineering responsibility.

Context: monorepos make integration visible

Keeping the Mainline Green Is Now a Platform Engineering Problem
Keeping the Mainline Green Is Now a Platform Engineering Problem

The InfoQ presentation, “Keeping the Mainline Green Across Diverse Language Monorepos,” discusses how Uber maintains green mainlines across massive monorepos that include many languages, services, and teams. The focus is on merge queues, CI reliability, and mainline health at a scale where traditional pull request workflows start to break down.

That context matters because many organizations are moving in a similar direction, even if they are not operating at Uber’s scale. Services, shared libraries, platform code, infrastructure-as-code, mobile clients, data pipelines, and internal tooling are increasingly being consolidated into larger repositories. The motivations are practical: shared standards, easier code discovery, coordinated refactoring, centralized dependency management, and better visibility into cross-service changes.

But consolidation has a cost. The larger the repository, the more fragile the integration point becomes. A change that passes in isolation may fail after another change lands. A test suite that is reliable for one team may become flaky under parallel load. A dependency upgrade may be safe for one service but incompatible with another language runtime or build path. The mainline becomes the place where every hidden coupling is eventually exposed.

In smaller repositories, teams often absorb that pain manually. Someone reruns CI. Someone reverts a bad commit. Someone posts in Slack asking who owns a failing test. In a large monorepo, that model does not scale. The mainline needs engineered guardrails.

Why “green mainline” is a platform concern

A green mainline means developers can branch, build, test, and release from a trustworthy source of truth. When the mainline is unhealthy, the entire engineering system slows down.

Developers lose confidence in CI results. Reviewers become more cautious. Release trains wait for unrelated failures. Platform teams spend time diagnosing whether a failure belongs to infrastructure, test ownership, dependency drift, or an application team. Eventually, teams normalize red builds, which is one of the most expensive cultural failures in software delivery.

This is why mainline health belongs in the platform engineering conversation. Platform teams are already responsible for developer experience, build systems, internal tooling, and paved roads. In a polyglot monorepo, the merge path is one of the most critical paved roads in the company.

The goal is not only to block bad code. It is to create a reliable integration system that gives engineers fast, trustworthy feedback while preserving delivery velocity.

Merge queues: serializing risk without stopping delivery

At the center of the InfoQ presentation is the role of merge queues. A merge queue controls how approved changes enter the mainline. Instead of allowing every approved pull request to merge immediately, the queue validates changes in an ordered, controlled way.

That distinction is important. In a busy repository, two changes can each pass CI independently but fail when combined. This is especially common when changes touch shared libraries, build configuration, generated code, protocol definitions, dependency manifests, or runtime assumptions. A merge queue reduces this integration risk by testing changes in a context closer to the actual post-merge state.

For engineering leaders, the key insight is that merge queues are not just a tooling feature. They are a policy layer for integration. They encode how much confidence the organization requires before code enters the mainline.

A mature merge queue strategy usually answers questions such as:

  • Which checks are required before a change can enter the queue?
  • Which checks run while the change is in the queue?
  • Can independent changes be batched safely?
  • What happens when a batch fails?
  • How are flaky tests distinguished from legitimate failures?
  • Who owns queue health when throughput drops?

Without those answers, a merge queue can become just another bottleneck. With them, it becomes a reliability mechanism for continuous integration at scale.

Polyglot monorepos need CI partitioning

Polyglot repositories introduce another layer of complexity. A single repository may contain Go services, Java libraries, Python data jobs, TypeScript frontends, Kotlin mobile code, Terraform modules, Helm charts, and generated API clients. Running every test for every change is usually too slow. Running too few tests creates risk.

CI partitioning is the practice of dividing validation into meaningful slices: by language, ownership, dependency graph, service boundary, risk level, or change type. The objective is to run the right tests at the right time, not necessarily all tests all the time.

For example, a documentation-only change should not wait behind a full distributed integration suite. A change to a shared authentication library probably should trigger broader downstream validation. A change to a Terraform module may need policy checks and plan validation rather than application unit tests. A protocol buffer change may need compatibility checks across multiple generated clients.

This is where platform engineering and software maintenance intersect. Good CI partitioning depends on accurate metadata: dependency graphs, ownership files, service catalogs, build manifests, and language-specific package relationships. If that metadata is stale, the CI system either over-tests, which slows delivery, or under-tests, which lets breakages through.

Modernization programs often focus on upgrading frameworks, replacing legacy runtimes, or consolidating services. Those efforts should also modernize the metadata and CI topology around the code. Otherwise, the repository may be newer, but the integration process remains fragile.

Ownership rules turn failures into accountable work

A flaky mainline often exposes an ownership problem. The failing test may be obvious, but the responsible team may not be. In a large monorepo, unclear ownership turns CI failures into archaeology.

Ownership rules help route failures to the right people quickly. They can be based on directory structure, service metadata, CODEOWNERS files, build targets, or internal catalogs. The implementation matters less than the principle: every meaningful part of the codebase should have an accountable owner, and CI should use that information automatically.

This is especially important for shared components. Libraries, build plugins, base containers, schemas, and infrastructure modules can affect many teams. When those assets are treated as “everyone’s code,” they often become no one’s responsibility. Assigning explicit ownership improves both maintenance and upgrade readiness.

Ownership also supports better escalation. If a queue is blocked by a test owned by the payments platform, the system should make that visible. If a shared test has been flaky for two weeks, it should become tracked reliability work, not background noise.

The broader cloud native community has been emphasizing similar ownership themes. Recent CNCF content around contributor, maintainer, and infrastructure engineering journeys highlights that sustainable platforms depend on clear stewardship, not just tooling. Internal platforms need the same mindset: someone must maintain the paths everyone else depends on.

Flaky tests are integration debt

Flaky tests are often tolerated because each individual failure seems small. But in a merge queue, flakiness directly reduces throughput. A single unreliable test can invalidate a batch, delay unrelated changes, and train developers to rerun instead of investigate.

That makes flakiness a form of integration debt. Like technical debt, it accumulates interest. The more engineers depend on the same queue, the more expensive each unreliable signal becomes.

Teams should treat flaky tests with the same seriousness as production defects in critical paths. That does not mean every flaky test requires an emergency response, but it does mean flakiness needs visibility, ownership, and prioritization.

Useful practices include:

  • Tracking flaky test frequency over time
  • Quarantining known flaky tests without silently ignoring them
  • Assigning owners and due dates for remediation
  • Separating infrastructure failures from product failures
  • Measuring queue time lost to nondeterministic tests
  • Reviewing the most expensive flaky tests in engineering health meetings

The key is to avoid allowing “rerun CI” to become the default operating model. Reruns may be necessary, but they should produce data that helps the platform improve.

Practical implications for engineering teams

Organizations do not need Uber-scale infrastructure to apply these lessons. The same patterns are useful for any team whose repository is large enough that mainline failures create cross-team drag.

1. Define mainline health as an explicit platform metric

Track more than pass/fail status. Useful metrics include merge queue wait time, CI duration by partition, failure rate by test suite, flaky test rate, revert frequency, and time to restore a broken mainline. These metrics show whether the integration platform is helping or slowing delivery.

2. Start with a merge queue for high-risk areas

You do not need to put the entire repository behind a sophisticated queue on day one. Start with shared libraries, infrastructure modules, release branches, or directories with high change volume. Expand as you learn where queueing improves reliability without creating unnecessary latency.

3. Invest in dependency-aware CI

Path-based CI is a good start, but it is often too crude for large codebases. Use build graph data, package dependencies, service catalogs, and ownership metadata to select more accurate checks. The better the dependency model, the more confidently you can avoid wasteful full-repo validation.

4. Make ownership machine-readable

If ownership only exists in people’s heads, CI cannot use it. Adopt CODEOWNERS, service metadata, catalog annotations, or similar mechanisms. Then connect ownership to review routing, failure notifications, merge queue diagnostics, and modernization planning.

5. Treat CI modernization as part of code modernization

When upgrading languages, frameworks, or build tools, update the CI strategy too. A Java version upgrade, for example, may require new cache behavior, new test partitions, or updated compatibility checks. A move from several repositories into a monorepo may require new queueing and ownership rules before the migration is complete.

6. Create a policy for flaky tests

Decide when a flaky test is quarantined, who owns the fix, how long quarantine is allowed, and how teams see the cost of flakiness. Without policy, unreliable tests become permanent fixtures.

The modernization angle: integration systems age too

Software maintenance is not only about source code. Build pipelines, test infrastructure, dependency metadata, and merge policies also age. Many organizations modernize application architecture while leaving the integration path unchanged. That creates a mismatch: modern code flowing through legacy delivery mechanics.

Vibgrate’s perspective is that modernization should include the systems that keep change safe. If a repository is becoming more centralized, more polyglot, or more business-critical, the CI and merge strategy should evolve with it. Otherwise, the mainline becomes the bottleneck that prevents teams from realizing the benefits of consolidation.

A green mainline is not a vanity metric. It is a prerequisite for confident upgrades, large-scale refactoring, dependency cleanup, and platform migration. Teams can move faster when they trust that the repository tells the truth.

Conclusion: the mainline is shared infrastructure

The lesson from Uber’s work, as discussed in the InfoQ presentation, is not simply that large companies need advanced CI. It is that mainline health becomes a platform capability when enough teams depend on the same codebase.

As monorepos grow more polyglot and more central to delivery, keeping the mainline green requires intentional queueing, clear ownership, reliable test signals, and dependency-aware CI. The future of software maintenance will not be defined only by cleaner code. It will also be defined by healthier integration systems that let teams change that code safely, continuously, and at scale.

Vibgrate CLI

See a real scan run

A replay of the actual CLI running against our test repositories — live progress, real findings, a genuine DriftScore. Nothing executes in your browser.

Replay
demo@vibgrate — bash
❯ npx @vibgrate/cli scan
 
╭──────────────────────────────────────────╮
│ Vibgrate Drift Report │
╰──────────────────────────────────────────╯
 
── node-turborepo (node) .
Runtime: >=18.0.0 (6 majors behind)
Frameworks:
Turbo: 1.13.4 → 2.11.5 (1 behind)
TypeScript: 5.9.3 → 7.0.2 (2 behind)
Dependencies:
1 current 1 1-behind 3 2+ behind 1 unknown
 
── @repo/admin (node) apps/admin
Frameworks:
TanStack Query: 5.104.0 → 5.104.0 (current)
React: 18.3.1 → 19.3.0 (1 behind)
React DOM: 18.3.1 → 19.3.0 (1 behind)
TypeScript: 5.9.3 → 7.0.2 (2 behind)
Vite: 5.4.21 → 8.3.1 (3 behind)
Dependencies:
3 current 9 1-behind 3 2+ behind 4 unknown
 
── @repo/api (node) apps/api
Frameworks:
Express: 4.22.3 → 5.2.1 (1 behind)
TypeScript: 5.9.3 → 7.0.2 (2 behind)
Vitest: 1.6.1 → 5.0.2 (4 behind)
Dependencies:
7 current 4 1-behind 4 2+ behind 4 unknown
 
── @repo/web (node) apps/web
Frameworks:
Next.js: 14.2.35 → 16.3.7 (2 behind)
React: 18.3.1 → 19.3.0 (1 behind)
React DOM: 18.3.1 → 19.3.0 (1 behind)
TypeScript: 5.9.3 → 7.0.2 (2 behind)
Dependencies:
2 current 6 1-behind 3 2+ behind 5 unknown
 
── @repo/config (node) packages/config
Frameworks:
TypeScript: 5.9.3 → 7.0.2 (2 behind)
Dependencies:
2 current 2 1-behind 5 2+ behind 0 unknown
 
── @repo/database (node) packages/database
Frameworks:
Prisma: 5.22.0 → 7.10.0 (2 behind)
TypeScript: 5.9.3 → 7.0.2 (2 behind)
Dependencies:
1 current 0 1-behind 3 2+ behind 1 unknown
 
── @repo/types (node) packages/types
Frameworks:
TypeScript: 5.9.3 → 7.0.2 (2 behind)
Dependencies:
0 current 0 1-behind 1 2+ behind 1 unknown
 
── @repo/ui (node) packages/ui
Frameworks:
React: 18.3.1 → 19.3.0 (1 behind)
TypeScript: 5.9.3 → 7.0.2 (2 behind)
React: 18.3.1 → 19.3.0 (1 behind)
Dependencies:
1 current 4 1-behind 1 2+ behind 1 unknown
 
── @repo/utils (node) packages/utils
Frameworks:
TypeScript: 5.9.3 → 7.0.2 (2 behind)
Vitest: 1.6.1 → 5.0.2 (4 behind)
Dependencies:
0 current 1 1-behind 2 2+ behind 1 unknown
 
Tech Stack
Frontend: React, React DOM
Meta-frameworks: Next.js
Bundlers: tsx, Turbo, Vite
CSS / UI: Autoprefixer, PostCSS, Tailwind CSS
Backend: Express
ORM / Database: Prisma, Prisma Client
Testing: Vitest
Lint & Format: ESLint, ESLint Prettier, ESLint React, Prettier, typescript-eslint
 
Services & Integrations
Auth: JWT 9.0.3
Databases: Prisma 5.22.0
 
TypeScript
v5.3.3 · strict ✔ · MIXED · target: ES2022
 
Build & Deploy
Package Managers: pnpm
Monorepo: npm-workspaces, pnpm-workspaces, turbo
 
Product Purpose Signals
Frameworks: react, nextjs
Evidence: 177
Top Signals:
- [heading] Dashboard (apps/admin/src/pages/Dashboard.tsx)
- [title] Revenue Overview (apps/admin/src/pages/Dashboard.tsx)
- [copy] workspace:* (packages/ui/package.json)
- [copy] ./dist (packages/ui/tsconfig.json)
- [copy] ./src/index.ts (packages/ui/package.json)
- [copy] @repo/config/tsconfig-base.json (packages/ui/tsconfig.json)
- [copy] @repo/ui (packages/ui/package.json)
- [copy] #3b82f6 (apps/admin/src/pages/Dashboard.tsx)
Unknowns:
- No pricing or billing evidence found.
- No integrations/connectors evidence found.
- No route structure evidence found.
 
Security Posture
Lockfile ✖ · .env ✔ · node_modules ✔
 
Platform
Native modules: turbo
 
Code Quality
Files: 36 · Functions: 183 · Avg complexity: 2.62 · Avg length: 21.13 lines
Max nesting: 2 · Circular deps: 0 · Dead code: 0%
God files: apps/admin/src/pages/Products (448 lines)
 
Database Schema
postgresql · 8 models · 1 enum
Models: Address, CartItem, Category, Order, OrderItem (+3 more)
 
Findings (16 errors, 11 warnings)
✖ Node.js runtime ">=18.0.0" reached end-of-life on 2025-04-30 (latest: 24.0.0).
vibgrate/runtime-eol in .
⚠ TypeScript is 2 major versions behind (current: 5.9.3, latest: 7.0.2).
vibgrate/framework-major-lag in .
✖ 60% of dependencies are 2+ major versions behind in node-turborepo.
vibgrate/dependency-rot in .
✖ @types/node is 6 major versions behind (spec: ^20.11.0, latest: 26.6.3).
vibgrate/dependency-major-lag in .
⚠ TypeScript is 2 major versions behind (current: 5.9.3, latest: 7.0.2).
vibgrate/framework-major-lag in apps/admin
✖ Vite is 3 major versions behind (current: 5.4.21, latest: 8.3.1).
vibgrate/framework-major-lag in apps/admin
✖ vite is 3 major versions behind (spec: ^5.0.12, latest: 8.3.1).
vibgrate/dependency-major-lag in apps/admin
⚠ TypeScript is 2 major versions behind (current: 5.9.3, latest: 7.0.2).
vibgrate/framework-major-lag in apps/api
✖ Vitest is 4 major versions behind (current: 1.6.1, latest: 5.0.2).
vibgrate/framework-major-lag in apps/api
✖ @types/node is 6 major versions behind (spec: ^20.11.0, latest: 26.6.3).
vibgrate/dependency-major-lag in apps/api
✖ vitest is 4 major versions behind (spec: ^1.2.1, latest: 5.0.2).
vibgrate/dependency-major-lag in apps/api
⚠ Next.js is 2 major versions behind (current: 14.2.35, latest: 16.3.7).
vibgrate/framework-major-lag in apps/web
⚠ TypeScript is 2 major versions behind (current: 5.9.3, latest: 7.0.2).
vibgrate/framework-major-lag in apps/web
✖ @types/node is 6 major versions behind (spec: ^20.11.0, latest: 26.6.3).
vibgrate/dependency-major-lag in apps/web
⚠ TypeScript is 2 major versions behind (current: 5.9.3, latest: 7.0.2).
vibgrate/framework-major-lag in packages/config
✖ 56% of dependencies are 2+ major versions behind in @repo/config.
vibgrate/dependency-rot in packages/config
✖ eslint-plugin-react-hooks is 3 major versions behind (spec: ^4.6.0, latest: 7.1.1).
vibgrate/dependency-major-lag in packages/config
⚠ Prisma is 2 major versions behind (current: 5.22.0, latest: 7.10.0).
vibgrate/framework-major-lag in packages/database
⚠ TypeScript is 2 major versions behind (current: 5.9.3, latest: 7.0.2).
vibgrate/framework-major-lag in packages/database
✖ 75% of dependencies are 2+ major versions behind in @repo/database.
vibgrate/dependency-rot in packages/database
⚠ TypeScript is 2 major versions behind (current: 5.9.3, latest: 7.0.2).
vibgrate/framework-major-lag in packages/types
✖ 100% of dependencies are 2+ major versions behind in @repo/types.
vibgrate/dependency-rot in packages/types
⚠ TypeScript is 2 major versions behind (current: 5.9.3, latest: 7.0.2).
vibgrate/framework-major-lag in packages/ui
⚠ TypeScript is 2 major versions behind (current: 5.9.3, latest: 7.0.2).
vibgrate/framework-major-lag in packages/utils
✖ Vitest is 4 major versions behind (current: 1.6.1, latest: 5.0.2).
vibgrate/framework-major-lag in packages/utils
✖ 67% of dependencies are 2+ major versions behind in @repo/utils.
vibgrate/dependency-rot in packages/utils
✖ vitest is 4 major versions behind (spec: ^1.2.1, latest: 5.0.2).
vibgrate/dependency-major-lag in packages/utils
 
╭──────────────────────────────────────────╮
│ Top Priority Actions │
╰──────────────────────────────────────────╯
 
1. Upgrade EOL runtime in node-turborepo
End-of-life runtimes no longer receive security patches and block ecosystem upgrades.
./.
>=18.0.0 → 24.0.0 (6 majors behind)
Impact: −10 drift points (runtime & EOL)
 
2. Fix security posture: no lockfile found
Without a lockfile, installs are non-deterministic. Run the install command to generate one and commit it.
./
Missing: package-lock.json, pnpm-lock.yaml, or yarn.lock
 
3. Upgrade Vitest 1.6.1 → 5.0.2 in @repo/api (+2 more)
4 major versions behind. Major framework drift increases breaking change risk and blocks access to security fixes and performance improvements.
./apps/api
Vitest: 1.6.1 → 5.0.2 (4 majors behind)
./packages/utils
Vitest: 1.6.1 → 5.0.2 (4 majors behind)
./apps/admin
Vite: 5.4.21 → 8.3.1 (3 majors behind)
Impact: −5–15 drift points
 
4. Reduce dependency rot in @repo/types (100% severely outdated)
1 of 1 dependencies are 2+ majors behind. Run `npm outdated` and prioritise packages with known CVEs or breaking API changes.
./packages/types
typescript: 5.9.3 → 7.0.2 (2 majors behind)
Impact: −5–10 drift points
 
5. Reduce dependency rot in @repo/database (75% severely outdated)
3 of 4 dependencies are 2+ majors behind. Run `npm outdated` and prioritise packages with known CVEs or breaking API changes.
./packages/database
@prisma/client: 5.22.0 → 7.10.0 (2 majors behind)
prisma: 5.22.0 → 7.10.0 (2 majors behind)
typescript: 5.9.3 → 7.0.2 (2 majors behind)
Impact: −5–10 drift points
 
╭──────────────────────────────────────────╮
│ Architecture Layers │
╰──────────────────────────────────────────╯
 
Archetype: nextjs (80% confidence)
Files classified: 24 (11 unclassified)
Folders classified: 8
apps/admin/src presentation 100% 4 files
apps/admin/src/pages presentation 100% 2 files
apps/api/src/middleware middleware 100% 2 files
apps/api/src/routes routing 100% 2 files
apps/web/src/app presentation 100% 4 files
apps/web/src/app/products presentation 100% 2 files
apps/web/src/app/products/[id] presentation 100% 1 file
packages/ui/src presentation 100% 6 files
Unclassified source (sample): 11
 
presentation 15 files drift ████████████████████ 100 risk high
routing 4 files drift ████████████████████ 100 risk high
middleware 2 files drift ███████▍░░░░░░░░░░░░ 37 risk moderate
config 2 files drift ░░░░░░░░░░░░░░░░░░░░ 0 risk none
shared 1 file drift ████████████████████ 100 risk high
 
╭──────────────────────────────────────────╮
│ DriftScore Summary │
╰──────────────────────────────────────────╯
 
DriftScore: 70/100
Risk Level: HIGH
Projects: 9
Classified: 8 nano · 1 micro · 0 small · 0 standard
Billable: 0.42 · 9 detected → 0.42 billable projects (micro-project pricing)
0.1 micro · 0.32 nano
These fractions add up across repositories, then round down to whole billable projects.
 
Score Breakdown
Runtime: ████████████████████ 100
Frameworks: ███████████▊░░░░░░░░ 59
Dependencies: ██████▌░░░░░░░░░░░░░ 33
EOL Risk: ████████████████████ 100
 
Scanned at 2026-09-30T06:02:13.873Z · 7.2s · 286 files scanned · 56 workspace files · 27 dirs
❯
Press Run to start.