
Session 3 — AI-Driven Development and Quality Engineering
Session 3 — AI-Driven Development and Quality Engineering
One sentence on slide 3 carries the whole day:
AI is only strong when the specification is good enough.
Sessions 1 and 2 were an argument for building context. Session 3 is where you find out whether that argument paid. The AI-Ready Specification Package assembled on Day 2 stops being a deliverable and becomes an input — the only input — for writing code, writing tests, and deciding whether any of it ships.
About this write-up
These are my notes from Day 3 of the AI-Driven Software Development Life-Cycle course, taught by Dr. Bùi Thị Mai Anh, School of Information and Communication Technology (SOICT), Hanoi University of Science and Technology, July 2026. The framing and the workshop design are hers; the summary and wording here are mine. The prompts reproduced below are taken from the course workbook — they are the instructor's, not mine, and I have translated them into English.
It follows Session 1 and Session 2. A Vietnamese version of this article is also available.
Where this sits
The shape worth noticing: the spec is in the middle, not at the front. Everything on Day 2 exists to produce it. Everything on Day 3 spends it. And the two markers are gates in the literal sense — each one ends with a person issuing a verdict.
Part 1 — Development with AI
From specification context to development context
Day 2 ended with one packaged document. Day 3 does not hand that whole document to the model and ask for the feature. It repeats the context-engineering move one level down: for the coding task in front of you, retrieve only what that task needs.
The stated rule:
AI does not implement the whole feature in one go. Each pass focuses on one coding task with exactly the context that task requires.
Which is why the very first step of the workshop is decomposition, not coding.
The five moves
| Step | AI does | The developer owns |
|---|---|---|
| Plan & decompose | Splits the feature into independent coding tasks with dependencies and complexity | Which task to build, and whether the split is sane |
| Development context | Retrieves the stories, rules, components, files and conventions this task touches | Confirming the context and catching invented assumptions |
| Coding | Generates the implementation, explains its technical decisions | Business logic correctness, accepting or rejecting the code |
| Unit testing | Writes tests per scenario, names the rule each one covers, flags what it could not cover | Whether the tests actually verify the behaviour |
| Code review | Reports issues by severity with impact and a suggested fix | Which findings to act on — and the final version |
Two constraints in that table are easy to skim past and are the point of the exercise.
Unit tests target behaviour, not implementation details. A test that breaks when you rename a private method has verified nothing about the business rule.
In review, AI reports; it does not fix. The review prompt ends with "do not modify the source code until a human confirms" — and only after the team accepts specific findings is the model asked to address exactly those.
Part 2 — Quality assurance with AI
The second half repeats the whole shape one more time, with QA as the actor and the same rule about context:
Do not feed the entire project to AI. Provide exactly the information the task needs.
Four steps: build the QA context (scope, rules to verify, components under test, integration points, risks, and which kinds of testing suit them); design and execute the chosen test types — API, integration, end-to-end; analyse the results into defects with severity, affected rules, probable root causes and regression recommendations; then assess — the release-readiness call.
The detail I liked: step 3 explicitly forbids the model from concluding anything about release readiness. Analysis and verdict are kept apart on purpose, so the evidence is assembled before the judgement rather than in support of it.
Part 3 — The spec becomes a file in the repo
This is where Session 2's closing claim — spec is no longer documentation, it is the interface — gets tested against something concrete.
The Day 2 output ships as a self-contained markdown file. You drop it into the repo at doc/specs/waitlist-feature.md, and then the prompt is simply:
Read
doc/specs/waitlist-feature.mdand implement section 6.
Eight sections, each written to be consumed rather than admired:
- Context & scope — in/out of scope, mandatory constraints
- Frozen business rules — BR-01 → BR-08, with the technical detail and the reasoning; the AI developer may not change them
- User stories & acceptance criteria — US-01 → US-07, Given–When–Then across happy path, alternative, exception, timeout and conflict
- Data model — SQL DDL ready to drop into a migration
- API contract — method, path, role, request, response, error codes, with example payloads
- Component → file mapping — every component points at a real file in the repo
- Non-functional requirements — NFR-01 → NFR-08, each with a measurable threshold
- Open questions — OQ-01 → OQ-06, deliberately outside the implement scope
Plus a 12-point handoff checklist confirming the document is AI-ready.
Two things in there are worth stealing outright.
Section 1 records corrections against the real codebase. The spec states plainly that MedBook has no Notification Service and no reschedule API — that the design assumed components which do not exist. Most handoff documents quietly paper over that gap and let the developer discover it.
There is an explicit precedence rule. When the Specification Package contradicts anything said earlier, the Specification Package wins, because it records what was settled — no "needs confirmation" left in it. That single sentence is what turns a pile of documents into an interface.
The workshop: what the exercises actually ask you to do
Same system as Day 2 — MedBook, and the same feature, Dynamic Appointment Rescheduling & Waiting List Management. What changed is that you now have a frozen spec and you are expected to produce running code.
Activity 1 — AI-assisted development (35 min)
AI plays AI-Assisted Developer: analyse the spec context, build the development context, link requirement to architecture to source, propose an implementation strategy, generate the service, the API, the unit tests and the technical documentation, explain code, propose refactoring.
The developer owns: confirming the context, checking AI's assumptions, choosing the implementation strategy, judging the business logic, reviewing generated code, accepting or rejecting proposals, and finishing the package before it goes to testing.
Step 0 — Task planning and decomposition. AI acts as tech lead and splits the feature into independently deliverable coding tasks — objective, related story and criteria, components touched, dependencies, priority, complexity — then recommends one task suited to the workshop's time budget and explains why. It is told explicitly not to generate code.
Prompt — Task Planning
You are a Technical Lead planning the development of the Dynamic Appointment
Rescheduling & Waiting List Management feature for the MedBook system.
Based on:
- The AI-Ready Specification Package
- The Architecture Blueprint
Decompose the feature into coding tasks that can be delivered independently.
Do not generate source code.
For each task, identify:
- The objective of the task
- The related User Story and Acceptance Criteria
- The main components affected (UI, API, Service, Database, ...)
- Dependent tasks, if any
- Priority
- Estimated complexity (Low / Medium / High)
Finally, recommend the single coding task best suited to the scope of this
workshop and explain why you chose it.Step 1 — Build the development context. Given the chosen task, retrieve only what it needs: the story and criteria it belongs to, the business rules it must implement, the modules and UI components involved, the services and APIs to reuse or change, the entities and DTOs affected, the existing files to look at, the conventions and technical constraints, the dependencies and risks, and anything missing that a human must confirm first.
The human review for this step is a twelve-item checklist, and two items on it are the ones that matter: did AI propose anything beyond the scope of this task? and are there assumptions AI invented that nobody confirmed?
Prompt — Development Context
You are a Senior Full-Stack Software Engineer preparing to implement a coding task
of the Dynamic Appointment Rescheduling & Waiting List Management feature on the
MedBook system.
Coding task
- Task ID: [Task ID]
- Task name: [Task name]
- Task objective: [Objective]
- Related User Story: [User Story]
Based on:
- The AI-Ready Specification Package
- The Architecture Blueprint
- The existing source code
- The existing APIs
- The existing data model
- The coding standards
Build the Task Context for the coding task above.
Do not generate source code and do not plan the coding at this step.
Retrieve and present only the information directly related to the task:
- Related requirement, User Story and Acceptance Criteria
- Business Rules that must be implemented
- Related modules, screens or UI components
- Services, APIs or utilities to reuse or change
- Related data entities, DTOs and data states
- Existing source files that must be examined
- Coding conventions and technical constraints
- Dependencies and technical risks
- Assumptions, missing information, or questions a human must confirm
For each item, state:
- The information or component involved.
- Its relationship to the coding task.
- How you expect to use, reuse or change it.
- The requirement, Business Rule or Acceptance Criterion it rests on.
Rules:
- Do not include information not directly related to the task.
- Do not invent Business Rules, API contracts or new assumptions.
- Do not propose changes outside the scope of the coding task.
- Anything without sufficient grounding must be marked "Needs confirmation".
- If required information is not in the input documents, state that it is
missing rather than inferring it.Step 2 — AI-assisted coding. Before generating anything, the model must summarise what it will change or reuse, which business rules it will implement, and what it is assuming. Then it writes the code, summarises the changes, and lists what it did not implement. Output is source code plus coding-log.md.
Prompt — Coding
You are a Senior Full-Stack Software Engineer implementing a coding task of the
Dynamic Appointment Rescheduling & Waiting List Management feature on the
MedBook system.
Based on:
- The Development Context
- The existing source code
- The coding standards
Implement the selected coding task.
Before generating any source code, briefly summarise:
- The components that will be changed or reused.
- The Business Rules that will be implemented.
- Any assumptions or information a human needs to confirm.
Then:
- Generate the source code for the coding task.
- Summarise the changes you made.
- List anything that has not been implemented.
Do not expand the scope beyond the Development Context.Step 3 — AI-assisted unit testing. For each test: the scenario, the business rule or acceptance criterion behind it, the expected behaviour, and any dependency that needs mocking. Then the test code, a coverage summary, and — the useful part — an explicit list of what it could not cover or found hard to test.
Prompt — Unit Testing
You are a Senior Software Engineer and Test Engineer.
Based on:
- The Development Context
- The source code just implemented
- The Business Rules
- The Acceptance Criteria
- The coding standards
Write Unit Tests for the selected coding task.
For each test, state:
- The scenario under test
- The related Business Rule or Acceptance Criterion
- The expected behaviour
- Any dependency that needs mocking
Then:
- Generate the Unit Test source code.
- Summarise the test coverage.
- Point out the cases that are not covered or are hard to test.
Do not write Unit Tests outside the scope of the coding task.Step 4 — AI-assisted code review. Across functional correctness, business-rule compliance, consistency, maintainability, reliability, security and test quality. Every issue gets a location, a severity (Critical / Major / Minor), an impact and a suggested fix.
The workbook is direct that the human is not obliged to accept any of it: findings that do not fit the requirements, the architecture or the task's scope may be adjusted or rejected. Only after that does the model fix the accepted issues — then re-run the tests and confirm no obvious regression.
Prompt — Code Review
You are a Senior Full-Stack Software Engineer reviewing the source code and Unit
Tests of a coding task belonging to the Dynamic Appointment Rescheduling &
Waiting List Management feature on the MedBook system.
Based on:
- The Development Context
- The source code
- The Unit Tests
- The related Business Rules and Acceptance Criteria
- The coding standards
Review the code and focus on issues that materially affect:
- Functional correctness
- Business Rule compliance
- Maintainability
- Reliability
- Security
- Unit Test quality
For each issue, present:
- The issue and where it occurs
- Severity: Critical / Major / Minor
- Impact
- A suggested improvement
Do not modify the source code until a human confirms.Step 5 — Development quality gate. AI assembles the evidence and scores six criteria as met or not met, then proposes PASS – Ready for QA or FAIL – Not Ready for QA, listing what must be finished if it fails. The team makes the actual call, and may disagree with the model when the evidence is thin.
Prompt — Development Quality Gate
You are the Technical Lead of the MedBook project.
Based on:
- The Task Plan
- The Development Context
- The source code and Coding Log
- The Unit Tests and Unit Test Report
- The Code Review Report
Assess how ready this coding task is to be handed over to Quality Assurance.
For each criterion below, mark it Met / Not met and give the evidence:
- The coding task was implemented within the agreed scope.
- The related Business Rules and Acceptance Criteria are implemented.
- Unit Tests have been executed and pass.
- Code Review has been completed.
- No unresolved Critical Issue remains.
- Remaining limitations, assumptions and risks have been recorded.
Finally:
- Summarise what has been completed.
- Point out the remaining risks or issues.
- Propose one of two outcomes:
- PASS - Ready for QA
- FAIL - Not Ready for QA
If the outcome is FAIL, state the work that must be completed before handover.The activity produces a named set of artifacts: task-plan.md, development-context.md, coding-log.md, unit-test-report.md, review-report.md, api-specification.md, development-handoff.md.
Activity 2 — AI-assisted QA and release readiness
Here is the design decision that makes this workshop work: each team becomes the QA team for a Development Package handed over by another team. The handoff stops being hypothetical. If the package is thin, you feel it immediately, which is the entire lesson of Session 1 delivered as an experience rather than a slide.
Step 1 — Build the QA context. Test scope, components under test, business rules and acceptance criteria to verify, critical flows and integration points, high-risk areas, and a recommended testing strategy. No test cases yet; missing information is marked needs confirmation rather than assumed.
Prompt — QA Context
You are a QA Lead responsible for preparing the testing activity for a feature
before Quality Assurance begins.
Based on the Business Context and the Development Package, build the QA Context
for the feature.
Identify:
- The test scope.
- The components to be tested.
- The Business Rules and Acceptance Criteria to verify.
- The critical business flows and integration points.
- The risks to prioritise.
- The recommended kinds of testing (API Testing, Integration Testing and/or
End-to-End Testing).
Do not design test cases at this step.
If required information is missing, mark it "Needs confirmation" rather than
making an assumption.Step 2 — Test design and execution. For each scenario: objective, the rule or criterion behind it, test data and preconditions, expected result. Generate test code in the project's framework where it fits, then record pass/fail and any anomalies.
Prompt — Test Design & Execution
You are a Senior QA Engineer responsible for testing the feature.
Based on the QA Context, design and support the execution of the selected kind
of testing.
For each test scenario, identify:
- The objective of the test.
- The related Business Rule or Acceptance Criterion.
- Test data or preconditions.
- The expected result.
Where appropriate, generate test source code using the project's framework.
After execution (or after receiving the execution results from a human):
- Summarise the results.
- State which tests passed and which failed.
- Record the defects or anomalous behaviour found.
Do not design scenarios outside the QA Context.Step 3 — Test analysis. Classify defects by severity, identify the rule or criterion each one breaks, propose probable root causes, and recommend the regression tests to run after a fix. And — deliberately — do not draw any release-readiness conclusion at this step.
Prompt — Test Analysis
You are a QA Lead responsible for analysing the test results of the feature.
Based on:
- The Test Execution Report
- The Business Context
- The Development Package
Analyse the test results.
If defects are found:
- Classify severity (Critical, High, Medium or Low).
- Identify the affected Business Rule or Acceptance Criterion.
- Propose the probable root cause.
- Propose how to handle it and which Regression Tests to run after the fix.
If no defects are found, summarise the evidence showing the feature meets the
tested scope.
Do not draw any Release Readiness conclusion at this step.Step 4 — QA assessment. For every rule and criterion tested: verified or not, the evidence, the remaining risk. Then one of three verdicts — Ready for next stage, Ready with known risks, or Not ready — with the reasoning and the follow-up work.
Prompt — QA Assessment
You are a QA Lead responsible for assessing the Quality Assurance results of the
feature.
Based on:
- The QA Context
- The Test Execution Report
- The Test Analysis Report
- The Development Handoff
Summarise the test results and assess the quality of the feature.
For each Business Rule or Acceptance Criterion that was tested, state:
- Whether it has been verified.
- The related test evidence.
- The remaining risks or limitations.
Finally, give one of the following recommendations:
- Ready for Next Stage
- Ready with Known Risks
- Not Ready
Explain the basis for your recommendation and the issues that need to be
addressed next, if any.Both gates have the same shape, and it is worth naming: named artifacts in, an AI summary in the middle, a human verdict out. The middle is the only part AI owns.
What I took away
The spec is the product of the first two days, and the input to the third. If it is vague, Day 3 does not fail loudly — it produces plausible code against the wrong rules, which is worse.
Decompose before you generate. "Implement the feature" is the prompt that produces the review backlog. One task with exactly its context is the one that produces mergeable code.
Separate analysis from verdict. Forbidding the model to conclude "ready to release" during test analysis stops the evidence from being assembled to fit a conclusion already reached.
Write down what you could not do. Nearly every step demands a list of what was not implemented, not covered, not tested. That list is what makes the next gate an informed decision instead of a rubber stamp.