
Session 5 — AI-Driven Workflow
Session 5 — AI-Driven Workflow
The last slide of the course puts it in one line:
AI helps us execute faster. Humans help us decide better.
And the first content slide sets up why the previous four days were necessary:
The final challenge is not using AI better. It is designing an engineering system that AI can take part in safely.
About this write-up
These are my notes from Day 5 — the final session — of the AI-Driven Software Development Life-Cycle course, taught by Dr. Bùi Thị Mai Anh, School of Information and Communication Technology (SOICT), Hanoi University of Science and Technology, July 2026. The framing is hers; the summary and wording here are mine.
Unlike Days 2–4 this session ships no workbook — the practical half is a capstone project and presentation, so there are no prompts to reproduce here.
It follows Session 1, Session 2, Session 3 and Session 4. A Vietnamese version of this article is also available.
AI-assisted is a loop; AI-driven is a system
The distinction the session opens with is sharper than the usual buzzword pair.
AI-assisted activities are a loop one person runs: prompt → generate → review → test → fix, and back around. It works. It just does not accumulate anything.
An AI-driven engineering system is defined by five properties instead of five verbs:
| Property | What it means |
|---|---|
| Shared context | One context, consistent across the whole system |
| Traceability | End-to-end links: requirement → design → code → test |
| Quality gates | Quality criteria at each stage |
| Human decisions | People make the decisions that matter |
| Continuous update | Updated continuously from feedback and real data |
Read the two lists side by side and the shift is obvious: the first is about how you work with AI, the second is about what the team's system guarantees regardless of who is prompting.
A change request, end to end
The workflow is given as ten steps: understand → build context → update spec → impact analysis → plan → implement → verify → human review → release → update shared knowledge.
What matters is where the controls sit. Not on all ten:
- Human confirmation before the spec is updated, and again before the plan is executed
- Artifacts → evidence through plan, implement and verify: code, tests, migration and docs on one side; test results, static analysis, security checks and traceability on the other
- Quality gate before release
And step ten is easy to skip and shouldn't be: update shared knowledge. The loop closes back into the context that the next change request will read.
Quality gate ≠ tests passed
This is the sentence from Session 5 I would put on a wall:
Quality gate ≠ tests passed. Quality gate = the evidence is sufficient and trustworthy.
The gate has three parts. First, evidence collected — and each item must have a concrete artifact and a concrete result, not a claim: code (PR #123), tests (TC-01, TC-02…), migration (Migration #45), docs (ADR-015), test results (all passed), static analysis (no critical issues), security checks (no vulnerabilities), traceability (UC-042 / AC-01).
Second, the gate checks — completeness (are all the artifacts mandatory for this kind of change present?), consistency (do spec ↔ code ↔ test ↔ docs agree?), quality (do the standards hold?), traceability (can every change be traced back to a spec, AC or UC?). That yields PASS or FAIL.
Third, the human decision — approve (merge, deploy, update status) or request changes (missing evidence, inconsistency, standards not met).
The rule joining them: a change proceeds only when it passes the gate and a human approves. Two conditions, not one.
The autonomy boundary
Inside the boundary, AI is genuinely autonomous — it plans, implements, tests, analyses, fixes and retests, repeating on its own until the gate passes. Nobody is reviewing each iteration.
What is fenced is the two ends. Before: a human defines intent, sets constraints and policies, and approves the plan. After: a human reads the gate result and decides what happens next.
Four principles hold the fence up: clear boundary between AI execution and human decision, full observability (every AI action logged, measurable, traceable), traceable accountability (every important decision has a human owner), and constraints first (AI acts only within already-approved intent, policy and constraints).
Which is the session's other memorable line: automate execution, not accountability.
The related point on checkpoints is a practical relief: you do not human-review every AI output. You review at decision boundaries — business approval after requirements, technical decision after architecture, code review after implementation, quality gate before release. Four checkpoints across five phases.
How the roles change
| Role | Becomes | Owns |
|---|---|---|
| PO / BA | Define & validate | Business intent, business rules, acceptance criteria, domain exceptions |
| Tech Lead | Guard architecture | Technical review, boundaries, ADRs, the AI decision boundary |
| Developer | Orchestrate implementation | Context, AI guidance, code, keeping test ↔ AC aligned |
| QA | Validate behaviour | Testability, edge cases, behaviour validation |
The summary line: human responsibility shifts upward toward intent, decision and validation, while AI takes on more of the execution. Note what happened to "developer" — the job is described as orchestrating, and the deliverable includes keeping tests tied to acceptance criteria.
The repo is the team's operating system
One use case, three views, three directories:
/specs/UC-042 → INTENT why are we doing this? what must be true?
/src/place-order-qr → IMPLEMENTATION how is it done?
/tests/place-order-qr → EVIDENCE how do we know it is right?Intent defines what "correct" means. Implementation realises the intent. Evidence proves it. The formulation on the slide: spec – code – test mirror each other.
If you want one structural change to take away from five days, it is this one. It is also the thing that makes the Session 4 failure impossible to hide: a use case whose evidence directory has nothing matching an acceptance criterion is visibly incomplete.
Five review rules for the AI era
1 · Spec PR ≠ Code PR. Two different kinds of change, two different reviews, two different goals. A spec PR changes business rules, acceptance criteria, exceptions and scope; it is reviewed by PO/BA and the tech lead, and the goal is is the intent right and complete? A code PR changes source, tests, refactors and docs; it is reviewed by developer peers, the tech lead and QA, and the goal is is the code correct, clean, safe and efficient? Business intent changing is not the same event as code changing.
2 · Trace every change: PR ↔ UC ID. Every implementation change must answer "which use case am I changing the system for?" The chain runs UC-042 → spec → BR/AC → PR #247 → code → tests — and crucially it runs backwards too. When a production incident hits, you walk code → PR → UC → business intent to find the cause, understand the real intent, and assess impact.
3 · Spec is the review baseline. Every behaviour in the code must be justified by the spec or by a human decision. The worked example is excellent:
The spec says VIP customers get deferred verification and non-VIP get normal verification. It says nothing about how long the deferral lasts. The generated code says:
if (customer.isVIP()) {
verificationMode = DEFERRED;
verificationDeadline = now.plusHours(24); // where did 24 come from?
}Behaviour in code: VIP → 24h. Found in the spec? No. Decision source? AI inferred it. Verdict: unjustified business decision. The code compiles, reads naturally, and would pass any review that only asked "does this look right".
4 · Verify business behaviour: AC ↔ Test. Do not ask "are there tests?" Ask "does every important acceptance criterion have a test proving it?" In the example, AC-01 and AC-02 map to passing tests and AC-03 — past 24 hours → suspend — maps to nothing. A gap, found by mapping rather than by counting.
The naming convention that makes the mapping mechanical:
UC-042: Deferred Verification
AC-01: VIP → Deferred Verification
└── TC-01 [UC-042][AC-01]: VIP gets deferred verification
AC-02: Non-VIP → Reject
└── TC-02 [UC-042][AC-02]: Non-VIP is rejected
AC-03: Past 24h → Suspend
├── TC-03 [UC-042][AC-03]: past 24h suspends
└── TC-04 [UC-042][AC-03]: within 24h does not suspendPut the IDs in the test name and the coverage question becomes greppable instead of a judgement call.
5 · Know who decided. Review the origin of the decision, not just the artifact.
The reviewer's checklist is four questions: is this logic in the spec? If not, is there an ADR recording it? If neither — did AI infer it? And if AI inferred it, it must be clarified, confirmed, or given an ADR.
The mental model, and where the series ends
The five rules collapse into one chain:
Intent → Implementation → Evidence → Gate → Decision.
Identify the right problem and goal. AI executes to produce code, tests and docs. Collect the evidence that proves the behaviour. Check quality, traceability and evidence at the gate. A human decides what happens next.
Which closes the arc the five days actually traced. Session 1 diagnosed context loss between phases. Session 2 turned context into a spec. Session 3 spent that spec on code and tests. Session 4 showed what it costs when the spec is stale and nobody checks. Session 5 puts the whole thing in a loop and draws the line around what AI is allowed to own.
What I took away
"Quality gate = evidence is sufficient and trustworthy" is a better definition than any CI status. A green pipeline answers did the checks I wrote pass. A gate answers is there enough trustworthy evidence to justify a decision — and Session 4 already demonstrated the gap between those two.
Splitting spec PRs from code PRs is the cheapest structural change on this list. Different content, different reviewers, different question being asked. Merging them into one review is how "business intent changed" gets approved by someone reviewing indentation.
Put the IDs in the test names. [UC-042][AC-03] turns "is this covered?" from a discussion into a grep. It is a small convention that makes rule 4 enforceable rather than aspirational.
Ask where a decision came from, not just whether it looks right. The plusHours(24) example is the whole course in six lines of code: correct-looking, well-formed, passing, and traceable to nobody.