PLATFORM ENGINEERING WORKSHOP · AZURE / AKS / ARGO CD
AI increases the cost of poor judgment
How to make good judgments
A practical workshop for SREs responsible for Azure infrastructure, a fleet of AKS clusters, and GitOps delivery paths.
?
45-minute facilitated session · N speaker notes · ? help
THE SHIFT
AI removed the implementation speed limit.
Before
Throughput limited damage.
Discussion, typing, and PR review created friction before a risky idea reached production.
1 meaningful change at a time
Now
Output can outrun understanding.
An agent can create policies, Helm values, roles, and pipelines faster than a team can verify their interactions.
13 unfamiliar changes waiting for review
The new bottleneck is not “can we make it work?” It is “can we explain, operate, and reverse it?”
THE JOB
Functional is a floor. Judgment is the differentiator.
01
Frame the problem
Know the constraint, the owner, and the consequence before choosing a tool.
02
Demand evidence
Prefer observability, constraints, and experiments over a confident recommendation.
03
Protect reversibility
Make the next change small enough to review, operate, and undo.
“The agent produced it” is not a design decision.
WHY THIS MATTERS TO SRE
In a platform, one shortcut can become a fleet-wide contract.
GITOne PRvalues, manifests, Terraform
ARGO CDReconcilesync, drift, promotion
AKSExecuteidentity, policy, networking
AZUREPersistcost, access, reliability
Changes are reconciled repeatedlyFailures cross team boundariesUndo is often harder than apply
A REPEATABLE LOOP
Good judgment turns speed into a controlled experiment.
1FrameWhat outcome, constraint, and blast radius?
2CompareWhat is the simplest viable alternative?
3ProveWhat evidence would change our mind?
4ConstrainWhat is the smallest reversible move?
5OwnWho operates the result when it fails at 02:00?
Use this loop before prompting, while reviewing, and during incident follow-up.
CLEAN ARCHITECTURE
Clean means decisions flow toward stable boundaries.
APPLICATION TEAMSIntent
Image, configuration, service contract
declares
ARGO CDDelivery
Desired state, promotion, reconciliation
targets
PLATFORM SREGuardrails
Cluster capabilities, policy, observability
runs on
AZURE FOUNDATIONConstraints
Identity, network, cost, regional resilience
High-level policy does not depend on deployment detailBoundaries have named ownersInterfaces are explicit and small
A HEALTHY PLATFORM CONTRACT
Application teams declare intent. The platform owns capabilities and constraints.
Engineering partner
Owns
Workload manifests and app configuration
Resource needs and SLO intent
Application rollback decisions
Explicit interface
Provides
Approved deployment contract
Documented promotion path
Observable service boundaries
Platform SRE
Owns
Cluster capability and tenancy model
Policy, identity, network, telemetry
Platform rollback and fleet safety
ANTI-PATTERN 01
The platform repo becomes a junk drawer.
platform/clusters/apps/experiments/legacy/fix-final-v4/values-prod-really/One merge boundary. Many unrelated owners.
Why it feels fast
Everything is close. Anyone can patch it. The agent can find and edit files.
Why it fails
Promotion rules blur, review loses context, and a local application need becomes shared platform behavior.
Repository proximity is not a design boundary.
SMELLS TO CATCH EARLY
Configuration smells are code smells with production credentials.
01
Copy-paste environments
Drift is the only promotion mechanism.
02
Magic sync order
Correctness depends on undocumented Argo timing.
03
Cluster-admin “fix”
Broad access compensates for an unclear contract.
04
Values-file archaeology
No one can explain which override wins.
05
Shared secret ownership
Rotation has no single accountable operator.
06
Permanent feature flag
A temporary bypass quietly becomes architecture.
TOOL SELECTION
Every new tool should retire a constraint—not create a mystery.
Before addingKafka · serverless · a controller · an AI agent loop
Need: What measured constraint cannot the current system meet?
Cost: Who owns on-call, upgrades, permissions, and failure modes?
Exit: How does this get removed if the assumption is wrong?
“The model recommended it” is a hypothesis. It is not evidence.
AI-ASSISTED DELIVERY
Keep the decision record separate from the agent transcript.
Fragile
Prompt → code → merge
Context, alternatives, and caveats disappear inside a long chat link.
Reviewable
Decision → bounded task → evidence → merge
PR explains intent, affected boundaries, proof, alerts, owner, and rollback.
Decision record, in the PR: why this design · alternatives rejected · risk and blast radius · verification · rollback owner
REVIEW FOR REDUCTION
Refuse PRs that ask reviewers to reverse-engineer the design.
01
Scope
One intent; no surprise refactor, migration, and capability change in one PR.
02
Trace
Show the path from Git to Argo to the cluster to the Azure resource.
03
Prove
State pre-production evidence, production signals, and the exact rollback.
If a reviewer cannot locate the control path, the author has not finished the change.
WORKSHOP · 12 MINUTES
Review this change as if it lands in your platform repo tomorrow.
1
Read the scenario.
2
Find the decisions hidden as implementation.
3
Choose the smallest safer next move.
Work in pairs or trios. You are not trying to perfect the design—only make the next decision trustworthy.
SCENARIO
“Make payments deploy faster across all production clusters.”
AI-generated PR summary
Creates a shared Argo CD ApplicationSet, enables automatic sync and prune, adds a cluster-wide role for deployment jobs, and turns off a network policy that blocks a new dependency.
Claimed benefit: fewer manual promotion steps and faster recovery.
Known context
Three production clusters have different tenant mixes.
Payments has one SLO and its own on-call rotation.
The new dependency is only needed by payments.
There is no tested rollback for an automated prune.
LOOK FOR THE SMELLS
What is being smuggled in as “faster deployment”?
A
Boundary
Which capability is shared across tenants, and which is application-specific?
B
Evidence
What proves that manual promotion—not a release contract—is the actual constraint?
C
Reversibility
How does a bad sync or prune get contained to one app and one cluster?
D
Ownership
Who responds when the shared controller changes a payment workload at night?
Choose one question your current review template would not force an author to answer.
DEBRIEF
Make the next move smaller, explicit, and observable.
Do not merge
Fleet-wide automation plus elevated access plus policy bypass.
It crosses three boundaries with no proof or containment.
Approve a bounded experiment
One app, one non-production cluster, a least-privilege role, and a measured promotion path.
Predefine health signals, a stop condition, and the owner who reverts the Argo change.
Separate: payments’ deployment contract from platform-wide Argo capability.
Restore: network policy with an explicit, narrow egress rule—not a global bypass.
Prove: reduced recovery time and correct rollback before promotion to a second cluster.
USE THIS IN YOUR NEXT REVIEW
The judgment checklist
01
Problem: What outcome and constraint are we solving?
02
Boundary: Who owns the interface and the 02:00 consequence?
03
Evidence: Why this over the simpler alternative?
04
Smallest move: What is safely reviewable now?
05
Signals: What proves health or harm after reconciliation?
06
Recovery: How do we stop, roll back, or repair?
Put these six questions in every AI-assisted Azure, AKS, and GitOps change.
THE COMMITMENT
Move fast. Leave a system people can still understand.
AI can accelerate implementation.
Good judgment chooses the direction, limits the blast radius, and keeps the platform operable.
Next step: adopt the checklist in your next platform PR review.