PLATFORM ENGINEERING WORKSHOP · AZURE / AKS / ARGO CD

AI increases the cost of poor judgment

How to make
good judgments

A practical workshop for SREs responsible for Azure infrastructure, a fleet of AKS clusters, and GitOps delivery paths.

THE SHIFT

AI removed the implementation speed limit.

Before

Throughput limited damage.

Discussion, typing, and PR review created friction before a risky idea reached production.

1 meaningful change at a time

Now

Output can outrun understanding.

An agent can create policies, Helm values, roles, and pipelines faster than a team can verify their interactions.

13 unfamiliar changes waiting for review

The new bottleneck is not “can we make it work?” It is “can we explain, operate, and reverse it?”

THE JOB

Functional is a floor.
Judgment is the differentiator.

01

Frame the problem

Know the constraint, the owner, and the consequence before choosing a tool.

02

Demand evidence

Prefer observability, constraints, and experiments over a confident recommendation.

03

Protect reversibility

Make the next change small enough to review, operate, and undo.

“The agent produced it” is not a design decision.

WHY THIS MATTERS TO SRE

In a platform, one shortcut can become a fleet-wide contract.

GITOne PRvalues, manifests, Terraform
ARGO CDReconcilesync, drift, promotion
AKSExecuteidentity, policy, networking
AZUREPersistcost, access, reliability
Changes are reconciled repeatedlyFailures cross team boundariesUndo is often harder than apply

A REPEATABLE LOOP

Good judgment turns speed into a controlled experiment.

  1. 1FrameWhat outcome, constraint, and blast radius?
  2. 2CompareWhat is the simplest viable alternative?
  3. 3ProveWhat evidence would change our mind?
  4. 4ConstrainWhat is the smallest reversible move?
  5. 5OwnWho operates the result when it fails at 02:00?

Use this loop before prompting, while reviewing, and during incident follow-up.

CLEAN ARCHITECTURE

Clean means decisions flow toward stable boundaries.

APPLICATION TEAMSIntent

Image, configuration, service contract

declares
ARGO CDDelivery

Desired state, promotion, reconciliation

targets
PLATFORM SREGuardrails

Cluster capabilities, policy, observability

runs on
AZURE FOUNDATIONConstraints

Identity, network, cost, regional resilience

High-level policy does not depend on deployment detailBoundaries have named ownersInterfaces are explicit and small

A HEALTHY PLATFORM CONTRACT

Application teams declare intent.
The platform owns capabilities and constraints.

Engineering partner

Owns

  • Workload manifests and app configuration
  • Resource needs and SLO intent
  • Application rollback decisions

Explicit interface

Provides

  • Approved deployment contract
  • Documented promotion path
  • Observable service boundaries

Platform SRE

Owns

  • Cluster capability and tenancy model
  • Policy, identity, network, telemetry
  • Platform rollback and fleet safety

ANTI-PATTERN 01

The platform repo becomes a junk drawer.

platform/clusters/apps/experiments/legacy/fix-final-v4/values-prod-really/One merge boundary. Many unrelated owners.

Why it feels fast

Everything is close. Anyone can patch it. The agent can find and edit files.

Why it fails

Promotion rules blur, review loses context, and a local application need becomes shared platform behavior.

Repository proximity is not a design boundary.

SMELLS TO CATCH EARLY

Configuration smells are code smells with production credentials.

01

Copy-paste environments

Drift is the only promotion mechanism.

02

Magic sync order

Correctness depends on undocumented Argo timing.

03

Cluster-admin “fix”

Broad access compensates for an unclear contract.

04

Values-file archaeology

No one can explain which override wins.

05

Shared secret ownership

Rotation has no single accountable operator.

06

Permanent feature flag

A temporary bypass quietly becomes architecture.

TOOL SELECTION

Every new tool should retire a constraint—not create a mystery.

Before addingKafka · serverless · a controller · an AI agent loop

Need: What measured constraint cannot the current system meet?

Cost: Who owns on-call, upgrades, permissions, and failure modes?

Exit: How does this get removed if the assumption is wrong?

“The model recommended it” is a hypothesis. It is not evidence.

AI-ASSISTED DELIVERY

Keep the decision record separate from the agent transcript.

Fragile

Prompt → code → merge

Context, alternatives, and caveats disappear inside a long chat link.

Reviewable

Decision → bounded task → evidence → merge

PR explains intent, affected boundaries, proof, alerts, owner, and rollback.

Decision record, in the PR: why this design · alternatives rejected · risk and blast radius · verification · rollback owner

REVIEW FOR REDUCTION

Refuse PRs that ask reviewers to reverse-engineer the design.

01

Scope

One intent; no surprise refactor, migration, and capability change in one PR.

02

Trace

Show the path from Git to Argo to the cluster to the Azure resource.

03

Prove

State pre-production evidence, production signals, and the exact rollback.

If a reviewer cannot locate the control path, the author has not finished the change.

WORKSHOP · 12 MINUTES

Review this change as if it lands in your platform repo tomorrow.

1

Read the scenario.

2

Find the decisions hidden as implementation.

3

Choose the smallest safer next move.

Work in pairs or trios. You are not trying to perfect the design—only make the next decision trustworthy.

SCENARIO

“Make payments deploy faster across all production clusters.”

AI-generated PR summary

Creates a shared Argo CD ApplicationSet, enables automatic sync and prune, adds a cluster-wide role for deployment jobs, and turns off a network policy that blocks a new dependency.

Claimed benefit: fewer manual promotion steps and faster recovery.

Known context

  • Three production clusters have different tenant mixes.
  • Payments has one SLO and its own on-call rotation.
  • The new dependency is only needed by payments.
  • There is no tested rollback for an automated prune.

LOOK FOR THE SMELLS

What is being smuggled in as “faster deployment”?

A

Boundary

Which capability is shared across tenants, and which is application-specific?

B

Evidence

What proves that manual promotion—not a release contract—is the actual constraint?

C

Reversibility

How does a bad sync or prune get contained to one app and one cluster?

D

Ownership

Who responds when the shared controller changes a payment workload at night?

Choose one question your current review template would not force an author to answer.

DEBRIEF

Make the next move smaller, explicit, and observable.

Do not merge

Fleet-wide automation plus elevated access plus policy bypass.

It crosses three boundaries with no proof or containment.

Approve a bounded experiment

One app, one non-production cluster, a least-privilege role, and a measured promotion path.

Predefine health signals, a stop condition, and the owner who reverts the Argo change.

Separate: payments’ deployment contract from platform-wide Argo capability.

Restore: network policy with an explicit, narrow egress rule—not a global bypass.

Prove: reduced recovery time and correct rollback before promotion to a second cluster.

USE THIS IN YOUR NEXT REVIEW

The judgment checklist

01

Problem: What outcome and constraint are we solving?

02

Boundary: Who owns the interface and the 02:00 consequence?

03

Evidence: Why this over the simpler alternative?

04

Smallest move: What is safely reviewable now?

05

Signals: What proves health or harm after reconciliation?

06

Recovery: How do we stop, roll back, or repair?

Put these six questions in every AI-assisted Azure, AKS, and GitOps change.

THE COMMITMENT

Move fast.
Leave a system people can still understand.

AI can accelerate implementation.

Good judgment chooses the direction, limits the blast radius, and keeps the platform operable.