
Session 4 — AI-Engineering Governance
Session 4 — AI-Engineering Governance
A team ships fast for three months. In month four the requirement changes — loan applications now need video verification of the borrower. Two weeks of work. Not because the change is hard, but because the main workload is understanding the code the AI wrote.
That is the failure Session 4 is about. The team moved quickly and left no trail to come back along.
About this write-up
These are my notes from Day 4 of the AI-Driven Software Development Life-Cycle course, taught by Dr. Bùi Thị Mai Anh, School of Information and Communication Technology (SOICT), Hanoi University of Science and Technology, July 2026. The framing and the workshop design are hers; the summary and wording here are mine. The prompts reproduced below are taken from the course workbook — they are the instructor's, not mine, and I have translated them into English.
It follows Session 1, Session 2 and Session 3. A Vietnamese version of this article is also available.
Context debt
The course gives the condition a name:
Context debt is the state where technical and business knowledge is no longer held in engineering artifacts, and survives only in people's memory or in chat history with an AI.
Technical debt you can see in the code. Context debt is invisible until someone asks a question and nobody can answer it.
Three questions that expose it:
"Why is this premium?" — "The PO said so back in March." — "Is it documented?" — "No." When business logic lives only in the heads of people who were there, or inside the code, the team is vibe coding whether or not it thinks of itself that way.
"What does verified mean?" The rule is Verified = OCR AND valid ID AND face match AND blacklist check. The unit tests check the implementation, not that rule. Behaviour tests protect the business; implementation tests only protect the current shape of the code.
"The PO changed the rule two months in." Originally: disburse only after the customer is verified. Now: VIP customers get disbursed first and complete verification within 24 hours. With good traceability you know exactly which use cases, criteria, tests and modules that touches. Without it, you go looking.
The failure taxonomy
Failures get sorted into groups, because each group needs a different fix.
Context loss
| Failure | What it looks like |
|---|---|
| Requirement drift | The code changed; the requirement did not |
| Incomplete spec | AI cannot produce behaviour the spec never described — every spec should carry a ##History |
| Domain language | Naming so thin that the model has nothing to reason from |
The domain-language example is the sharpest thing on the slides. Compare two prompts:
A: Create a service to manage
tbl_dh_v2with columnsid,ma_kh,tong_tien,tt,ngay_tao,ngay_cap_nhat.
B: Create a service to manage
Order. AnOrderhas manyOrderItem, belongs to aCustomer, has aPaymentonce paid, and has statesdraft,awaiting_payment,paid,cancelled.
Same table. The second one gives the model semantics to reason with; the first gives it six opaque strings. And a related trap: the same term used across different specs gets understood differently each time, which is why a short clarifying note next to a term earns its space.
Engineering failures
| Failure | Signal | Consequence |
|---|---|---|
| Wrong AI decision | A rule exists in the code but not in the spec | Wrong business logic — with tests passing |
| Architecture erosion | The same business rule duplicated across features | Single source of truth quietly gone |
| False test confidence | Tests exercise implementation, not the business rule | Green suite, unprotected behaviour |
On erosion, the point is specific: AI accelerates architecture erosion when each feature is generated independently without being bound by the same architecture and engineering context.
The case that actually happened
Here is what makes this session land. While preparing the sample solution for this very course, the scenario Day 1 warned about happened for real.
The day03 branch was implemented against an outdated single-file spec — a draft that called itself "Frozen Business Rules" but had been superseded — instead of the official nine-file Specification Package v1.0 FROZEN.
What came out of it:
- Wrong data model — no
medical_priority, nooffer_eventsaudit table - Wrong permissions — patients self-registering, when the spec says staff only
- Plain FIFO candidate selection instead of medical priority
- A Notification Service that the real spec explicitly forbids, built anyway because the stale draft asked for it
Nobody caught it. And the reason nobody caught it is the part worth internalising:
The old draft's test suite passed 100% — because it validated the system against exactly the same misunderstanding the system was built from. A self-consistent suite is not a correct one. Only an independent audit that compared the code directly against the frozen spec found it.
Then a second, deeper audit — after the rewrite, with the suite at 37/37 green — found five more real reliability holes (a missing transaction, an untested safety switch, dead code) and three acceptance criteria that no test touched at all.
Spec-centric, and the waterfall question
The response to all of this is to put the spec at the centre: the spec is where the team agrees what it is building, AI writes code and tests and docs around it, and when something is wrong you fix the spec, then fix the code.
Which invites the obvious objection: is this just waterfall again?
The answer the course gives is a distinction worth keeping. Spec-centric work takes what waterfall was right about — requirements that are clear, traceable and testable — and drops what made it painful: rigid sign-off gates, designing every detail before starting, and changes that cost too much to make.
The six governance principles
1 · Requirement-centric. The alternative is prompt-centric, and the slide's line for it is memorable: the conversation drifts away, and the business rule stays behind in a prompt. Prompts are temporary and unmanaged; project knowledge should not be scattered across chat logs. The rule: spec drives code, not code drives spec — if a new concept shows up in the code, it should have shown up in the spec first.
2 · AI-assisted. AI fits work with a clear shape: CRUD, mapping, generating tests from existing acceptance criteria, generating docs from code and spec. AI does not get to decide things with large business consequences — cancelling a loan, refund policy. The human decides what is business-correct; AI executes it consistently.
3 · Iterative improvement. The spec does not have to be perfect on day one. It is a living document. Update the spec first, then have AI generate the code.
4 · Test-protected. Tests are the fence that lets AI rewrite code without drifting behaviour. Concretely: every acceptance criterion should have at least one test, and the test name or annotation should carry the use-case ID and the criterion ID.
5 · Stakeholder-centric. A spec is only worth something once someone who understands the business context and business impact has confirmed it. Per sprint: one approval from business, one from the tech lead. A feature needs both views — is the spec implemented correctly and consistently with the architecture? and is the spec true to reality, and is it missing a rule so obvious nobody wrote it down?
6 · Traceable. From code you can find the spec, from the spec the test, from the test the requirement. Traceability has to run three ways — business requirement ⇔ use case ⇔ code and test — so that when the PO changes a requirement, the team knows exactly which use cases, criteria, test cases and modules are affected.
The workshop: what the exercises actually ask you to do
Sixty minutes, and the object under audit is the team's own Day 3 output. Day 1 asked people to predict where context would break. Day 4 checks whether it actually did.
AI plays AI-Assisted Auditor — the same role used to produce the audit described above. It may cross-check artifacts between phases, compare code against frozen business rules, find dead code via grep, and draft root-cause labels and governance rules.
It is explicitly not accountable for: deciding whether a finding is severe enough to block release, deciding which governance rule is realistic for the team, or confirming that an artifact "matches" when that needs knowledge of the original business intent — only the humans who were in the Day 2 requirements meeting know that.
The stated principle is Human-on-the-loop this time, with one hard condition: never let AI grade something "safe" without evidence pointing at a specific file, line or artifact.
Step 0 — Readiness check. Confirm the five artifacts exist: the Day 1 SDLC artifact map, the Day 2 FROZEN specification package, the Day 3 development package, the Day 3 QA package, and the running source and test suite. The discussion question is good: if one of these is missing, which audit step below weakens first?
Step 1 — SDLC context-drift audit. An audit of information, not code. The central question: did the person or AI in the later phase read the latest version of the earlier phase's artifact? — the exact question the real case answered "no". Every break needs the two phases it sits between, the concrete consequence, and evidence quoted from a file and line.
Prompt — Context-Drift Audit
You are an AI-Assisted Auditor reviewing the SDLC process of a team that has just
finished Days 1-3 of the MedBook case study.
Based on:
- The team's SDLC Artifact Map (Day 1)
- The Specification Package used as input for Day 3 (Day 2)
- The Development Package and the actual source code produced (Day 3)
Compare them and answer: does the Day 3 source code match the CORRECT latest
version of the Day 2 spec? Give concrete evidence (table names, endpoint names,
business rules) - do not conclude vaguely that it "seems to match". Was any Day 2
artifact replaced or updated afterwards while Day 3 kept using the old version?
Was any Day 1 decision (Human-AI Responsibility Matrix, checkpoints) skipped in
the actual Day 3 implementation?
For each break you find, record: which two phases it sits between, the concrete
consequence, and the evidence (quote the file and line, or the artifact excerpt).
Do not infer without evidence - mark it "Needs confirmation with the team".Step 2 — Code and test-coverage failure analysis. Read the actual source and test suite — not old reports, not inferences from file names — and map every business rule to the file implementing it and the test verifying it. If part of a rule has no test touching it, that is a gap, not "covered". Unreferenced functions get confirmed as dead code with grep, not by guessing. And code with a comment explaining what it does but not why this approach rather than the repo's usual one gets flagged as missing context.
Prompt — Code & Coverage Audit
You are an AI-Assisted Auditor. Read the team's actual source code and test suite
directly (do not use old reports, do not infer from file names) and compare them
against the correct FROZEN Business Rules.
For each Business Rule:
- Point to the exact file/function implementing it.
- Point to the exact test (if any) verifying it, naming the test.
- If part of the rule has no test touching it, record that as a "gap", not
"covered".
- If the code has a function that is never called anywhere (confirm with grep,
do not guess), list it as dead code.
- If a piece of code has a comment explaining its MEANING but not the REASON for
choosing this approach over the convention the rest of the repo uses, mark it
as "missing context".
Present the result as a table: Business Rule -> implementing file -> verifying
test -> status (Has test / Indirect / No test).Step 3 — Root cause classification. Every finding gets exactly one of five labels.
Prompt — Root Cause Classification
For the list of findings from Step 1 and Step 2, assign exactly one of the five
root-cause labels (Spec wrong or missing · AI invented it when context was
missing · Human review missed it · No process checkpoint · Technical/concurrency)
and a severity (Critical / Major / Minor).
Briefly explain why you chose that label rather than another - in particular, be
clear about the difference between "AI invented it" and "Spec wrong or missing"
(AI invented it is when the spec said NOTHING and AI decided anyway; Spec wrong
or missing is when the spec DID say something but it was wrong or already
superseded).
Do not file a finding under a lighter category just to tidy up the list.The workbook is specific about the distinction people get wrong: AI invented it is when the spec said nothing and AI decided anyway; spec wrong or missing is when the spec did say something but it was wrong or stale. It also warns against quietly filing a finding under a lighter label to tidy up the list.
Step 4 — Governance charter. From whichever root cause dominates, write three to five rules. Each rule must state the mandatory action as a condition that can be true or false, name which root cause it blocks, say how compliance is verified — automated check or a human-signed checklist — and name an owner by role.
The bar is set by a counter-example: "review more carefully" is not a governance rule. It cannot be verified and nobody is accountable when it is skipped. The workbook's sample rule shows the shape:
Before starting a new coding task, AI must quote the exact filename and FROZEN date of the spec it is using as input, and a human must confirm it is the latest version before code generation is allowed.
And the closing question of the whole day: if you had applied your own governance charter from the start of Day 3, would the real failure have happened?
Prompt — Governance Charter
You are a Technical Lead designing governance for your team's AI-Assisted
Development process, based on the root-cause classification from Step 3.
For each root cause that accounts for a significant share, propose one concrete
governance rule containing:
- Rule - the mandatory action, phrased as a condition that can be true or false
(not a piece of advice).
- Which root cause it blocks - point back to the exact label from Step 3.
- How it is verified - by an automated tool (test, CI, lint) or by a checklist a
human signs off?
- Owner - which role (BA / Dev / QA / Tech Lead) is accountable for this rule not
being skipped?
Reference example (do not copy verbatim, it must match your team's real
findings): "Before starting a new coding task, AI must quote the exact filename
and FROZEN date of the spec it is using as input, and a human must confirm it is
the latest version before code generation is allowed."
Do not propose vague, unmeasurable rules (e.g. "review more carefully", "be
careful when prompting").What I took away
Context debt is invisible by construction. Technical debt shows up in the code; context debt only shows up when someone asks why, and the answer is a person's memory or a chat log that has scrolled away.
A green test suite is evidence about consistency, not correctness. If the tests came from the same reading of the spec as the code, they will always agree with it. The only check that catches that is comparing code to the spec.
Name the root cause or you cannot fix it. "It was a bug" and "the process never required anyone to confirm which spec version was used" lead to completely different next actions.
A governance rule that cannot be checked is a slogan. Verifiable condition, named owner, defined check. Anything short of that is a note to be more careful, which is what everyone was already trying to do when the failure happened.