
Session 1 — From Traditional SDLC to AI-Driven SDLC: Where Are We?
Session 1 — From Traditional SDLC to AI-Driven SDLC: Where Are We?
Every phase of the software lifecycle now has an AI assistant pointed at it. Business analysts draft user stories with one. Architects sketch options with another. Developers generate code, testers generate cases, reviewers get automated comments.
And yet, at the level of the whole delivery pipeline, throughput has barely moved.
That gap is the entire subject of Session 1. Not "can AI write code" — that question is settled — but a harder one: if AI already accelerates nearly every individual activity, why isn't the overall process meaningfully faster?
About this write-up
These are my notes from Day 1 of the AI-Driven Software Development Life-Cycle course, taught by Dr. Bùi Thị Mai Anh, School of Information and Communication Technology (SOICT), Hanoi University of Science and Technology, July 2026. The framing and the workshop design are hers; the summary and wording here are mine.
A Vietnamese version of this article is also available.
What AI is actually good at
Before deciding where AI belongs in a process, the session grounds things in what AI contributes across any phase. Five capability groups:
| Capability | What it looks like in practice |
|---|---|
| Understand and analyse | Reading requirements, source code, logs, and existing artifacts to extract meaning |
| Generate | Turning a requirement into the artifacts you actually need — user stories, schemas, tests, docs |
| Transform and restructure | Moving information between formats and levels of abstraction (spec → design → code → test) |
| Evaluate and critique | Catching problems early, proposing improvements, raising artifact quality |
| Recommend and support decisions | Laying out options and trade-offs so a human can choose faster |
Read that list again and the paradox sharpens. AI can participate in nearly every activity in the lifecycle. So the bottleneck isn't capability.
Why local speed-ups don't add up
The course names three mechanisms that turn per-activity gains into no net gain.
1. Each AI interprets the problem differently. Every team optimises its own phase with its own assistant. Nothing forces those assistants to share an understanding of what is being built.
2. Information degrades at every handoff — context drift. What moves between phases is not just documents. It's knowledge and decisions: why this trade-off was accepted, which rule is non-negotiable, what was deliberately left out. Documents survive the handoff. The reasoning behind them often doesn't.
3. Local acceleration creates global loops. Going faster in one phase with a partial understanding just produces wrong output sooner. Rework is the consequence of lost context, and rework eats the speed-up.
The root cause underneath all three: each AI only ever sees a slice of the problem. It sees exactly what you gave it, and nothing else. The whole product problem is never in view.
The quality of each phase's output depends on the context it received from the phase before it. AI does not solve cross-phase coordination on its own, and it does not replace the conversation and confirmation between people.
AI across the six phases
The structural point of the session: the shape of the SDLC does not change. Requirements, design, coding, testing, deployment, operation — same six phases. What changes is how the work is done, what humans are for, and how phases feed each other.
For every phase, three questions get asked in the same order: what is AI doing, what must a human decide, and what context does AI need to be useful at all.
Requirements — Clarify
| AI does | Human decides | AI needs as context |
|---|---|---|
| Clarifies requirements, generates user stories, use cases, acceptance criteria; flags missing or contradictory requirements | Business goals, scope, priority, trade-offs, business rules, final approval | Business goals, current business process, business rules, stakeholders, constraints |
AI helps build the requirement. A human remains accountable for whether it is correct and worth building.
Design — Explore
| AI does | Human decides | AI needs as context |
|---|---|---|
| Analyses design requirements, proposes options, generates UML / API / database schemas, evaluates pros and cons, checks consistency | System architecture, trade-offs, technology choices, non-functional requirements, approval | Business requirements, existing architecture, non-functional requirements, technical constraints and standards |
AI's real value here is breadth — surfacing and evaluating more options than a person would have time to draft. Choosing the one that fits the system's context stays human.
Development — Implement
| AI does | Human decides | AI needs as context |
|---|---|---|
| Reads and generates code, refactors, writes unit tests and documentation, detects bugs and proposes fixes | Detailed design, business logic, PR acceptance, merge decisions, code quality, production accountability | Approved requirements, system design, coding standards, existing source, API contracts, project conventions |
Testing — Verify
| AI does | Human decides | AI needs as context |
|---|---|---|
| Analyses test requirements, generates cases, scripts and data, analyses failures and root causes, proposes regression tests | Test strategy, coverage targets, acceptable risk, fix priority, quality judgement | Acceptance criteria, business requirements, system design, source code, defect history, previous test results |
AI can generate and run a great deal of testing. The question it cannot answer is "is this good enough to release?"
Deployment — Deliver
| AI does | Human decides | AI needs as context |
|---|---|---|
| Generates release notes, builds CI/CD pipelines, validates deployment config, monitors releases, detects and analyses incidents, proposes rollback | Release timing, deployment strategy, acceptable risk, rollback calls, release approval | Deployment plan, target environments, CI/CD pipeline, system configuration, release policy |
Operation — Optimize
| AI does | Human decides | AI needs as context |
|---|---|---|
| Monitors systems, detects anomalies, analyses incident causes, proposes performance optimisations, analyses user feedback, suggests next-version improvements | Improvement priorities, incident response plans, feedback weighting, optimisation and release planning | Operational goals, logs and metrics, user feedback, incident history, performance data |
Operation is where the loop closes: AI turns operational data into knowledge that feeds the next version's requirements.
Read end to end, the six phases become six verbs:
Clarify → Explore → Implement → Verify → Deliver → Optimize
AI doesn't replace the human at any of these. It lets the human decide faster, on more information, at higher quality.
Three models of human–AI collaboration
The second half of the session is about how much autonomy to grant, and it introduces three models.
Human-in-the-loop (HITL) — a human must approve before AI acts, or before AI output gets used. Appropriate when: the decision is high-risk, customer-impacting, legally constrained, or hard to undo.
Human-on-the-loop (HOTL) — AI executes; the human supervises and intervenes when needed. Appropriate when: the work is repetitive and high-frequency, results are measurable, rollback exists, and there's a clear established process.
Human-out-of-the-loop (HOOTL) — AI decides and acts within pre-set policy. Appropriate when: risk is low, the process is standardised, auditing is in place, rollback is possible, and the decision recurs.
The sentence that carries the most weight in the whole session:
What determines the choice between HITL, HOTL and HOOTL is not how smart the AI is. It is the level of risk, accountability and control attached to the decision.
That reframes autonomy as a governance question rather than a capability question — which is why "the model got better" is not, by itself, a reason to remove a human approval gate.
The workshop: what the exercises actually ask you to do
The theory is delivered against a running case study, worked in teams. This is the part worth copying if you want to run the same exercise.
The case study
A general hospital runs an appointment and patient-service management system. Patients search for doctors and specialities, book, reschedule and cancel appointments, track status and receive notifications. Doctors and staff manage working schedules and coordinate consultations. When a slot frees up — a doctor's schedule changes, or a patient cancels — staff manually handle the knock-on changes. At peak times patients who can't get a slot go onto a waiting list.
The change request: build flexible appointment-change and waiting-list management. When a slot becomes available, the system should identify suitable waiting patients and offer them the slot, weighing priority, time spent waiting, and fit between the patient's need and the slot. Patients can accept or decline; if they decline or don't respond within a window, the slot passes to someone else. The system must also handle reschedules initiated by patients, doctors or the hospital, avoid scheduling conflicts, and notify affected parties.
The team setup is deliberately traditional: a BA takes the requirement, an architect designs, a developer implements, a tester writes cases, a tech lead reviews before deployment. Everyone may use their own AI assistant — ChatGPT, Claude, Copilot, Gemini — but those assistants share no context with each other.
Activity 1 — Map the current SDLC and find where information breaks (20 min)
The goal is to locate context drift in a process you recognise.
Step 1 — Pick what actually has to travel. You're given the artifacts each phase produces: requirements analysis yields business goals, functional requirements, business rules, user stories, acceptance criteria, priority, constraints, stakeholders; design yields architecture decisions, API contracts, database schema, component responsibilities, design constraints, technology stack, security design; development yields source code, logging, exception handling, configuration, known limitations, technical assumptions.
For each handoff — Requirements → Design, Design → Development, Development → Testing — choose only the three most important items to carry forward as AI input, and justify the choice. The constraint to three is the point of the exercise.
Step 2 — Test it against roles. For the developer, the tester and the architect in turn: which single piece of information do they most need, and what specifically goes wrong with AI output if it's missing?
Step 3 — Converge. Agree on the three pieces of context that must persist across the entire lifecycle, ranked, with reasons.
Step 4 — Discuss. Four questions: Which phase produces the most critical information, and why? Which information is most easily lost at handoff? What is the single biggest cause of rework? If each phase uses a different AI assistant, which one struggles most, and why?
Deliverable: one slide, five-minute presentation.
Activity 2 — Design a human–AI collaboration model (30 min)
Activity 1 establishes what context must survive. Activity 2 designs the process around it.
Step 1 — Decide where AI belongs. Across ten activities — requirements analysis, user-story generation, architecture proposal, API generation, code generation, unit-test generation, code review, security review, documentation, production release — classify each as AI should not participate / AI assists / AI can automate. Then answer: where does AI add the most value, and where does it carry the most risk?
Step 2 — Build a Human–AI Responsibility Matrix. For each phase (requirements, design, development, testing, code review, deployment): what is AI responsible for, what is the human responsible for, and why.
Step 3 — Separate deciding from suggesting. For decisions grouped by area — requirements engineering, system design, development, testing and QA, code review and security, deployment and operations — sort each into AI can decide / AI only suggests / human decides. Some are deliberately provocative: merging a pull request, committing directly to main, closing a bug, accepting a security risk, approving a hotfix, rolling back a release.
There is no single correct answer. Teams sort by risk, accountability and impact. Then rank the three decisions that must always stay human, and explain why.
Step 4 — Distinguish HITL from HOTL. Classify concrete scenarios: AI generates user stories, BA checks before use. AI generates code, developer reviews before merge. AI reviews the PR, developer only looks at warnings. AI generates unit tests, developer only looks when one fails.AI auto-deploys to production. AI auto-rolls-back on error.
Then the closing question, which is the one that matters: assume AI keeps getting more accurate — what conditions must hold before an activity may move from human-in-the-loop to human-on-the-loop? Candidate criteria include stable accuracy, explicit confidence scores, explainability, audit logging, rollback capability, sufficient historical data, a continuous human-feedback mechanism, an established QA process, low decision risk, and easy recovery from mistakes. Pick the three that matter most.
Deliverable: one slide covering the responsibility matrix, the three always-human decisions, three activities that need HITL, three ready to move to HOTL, and a set of completed AI governance statements — "AI may decide autonomously only when…", "decisions touching business rules must…", "every AI-made decision requires…", "when AI and human conclusions differ…".
What I took away
Adopting AI tools is not the same as adopting an AI-driven SDLC. Dropping an assistant into each existing activity leaves the process shape untouched — and the process shape is where the losses are.
Context is the real artifact. Documents move between phases easily; the decisions and reasoning behind them don't. Both humans and AI degrade when that reasoning is missing, but AI degrades silently — it produces fluent, confident output from insufficient context, which is worse than producing nothing.
Autonomy is a governance decision, not a capability decision. The question is never "is the model good enough yet". It's whether the decision is reversible, auditable, and low-stakes enough that supervision can be relaxed.
The rest of the course builds on this: if shared context is the bottleneck, the interesting work is constructing a context layer that every phase — and every assistant — can actually read from.