
Claude Fable 5 Is Back: A Demo That Can Actually Be Wrong
Claude Fable 5 Is Back: A Demo That Can Actually Be Wrong
Most model demos are unfalsifiable. Build a landing page, write a slide deck, summarise a report — the output looks impressive and nobody can point at a line and say that is factually incorrect. Which means the demo proves the model can produce plausible shapes, and not much else.
This one is different, and that is the only reason it is worth writing up.
The task: take five years of published university cutoff scores, take a student's exam result, and tell them which programmes are realistic. Every number in the output traces to a specific row in a specific file. If the model invents a cutoff, you can catch it by opening the file.
About this write-up
Written from the video on the channel. The observations below — including the bug the demo found in its own output — are what happened on screen. Where a claim comes from the vendor's own material rather than the run, it is labelled as such.
Why this task is a good test
Vietnamese university admission works off a published cutoff score per programme per year. A student finishes the national exam, gets three subject marks plus a regional bonus, and has to guess which programmes are within reach.
That makes it an unusually honest benchmark for an agent:
- There is a ground truth. The 2023 cutoff for a given programme is a fact in a file.
- The arithmetic is checkable. Three subjects plus a bonus, with rounding rules that matter.
- Being wrong is expensive. A student acting on a hallucinated cutoff makes a real decision badly.
The build was done in Claude Code with Fable 5 selected, in plan mode, against a folder holding the score data as markdown and a separate file of the 2026 regional-bonus regulations.
This is the written companion to the video, where the whole run happens on screen — the plan, the fifteen minutes of building, the self-test, and the filter that comes out wrong.
Prefer to watch it on YouTube?
Watch on YouTube — or subscribe to the channel for the next one.
The step worth copying
Steps one to four are what every agent demo does: read the input, plan, fan out to sub-agents, rank the results into safe, matched, and reach.
Step five is the one that matters. After producing the answer, the agent re-opens the same source files and checks every number it used against them. Not a test suite — a cross-check against the evidence the answer was built from.
That is the difference between an agent that sounds right and one whose output you can audit. It is the same argument the governance session makes at length: a result is worth exactly what its evidence is worth, and evidence you never re-read is not evidence.
What actually happened in the run
It read before it planned. Given the folder, it walked the data directory to understand the structure first, and only then produced a plan.
It caught an error in the input data. The 2023 file was not a table — it had been left as plain text while the other years were structured. The model flagged the inconsistency and re-checked the format rather than parsing it wrong and moving on. That is the more impressive moment in the whole demo, and it was not asked for.
Plan mode produced something reviewable — data structures, formats, university-name handling, core logic, files to create, implementation steps — before a line was written.
It tested itself, unprompted. After building, it launched the site and clicked through it, checking buttons and logic. Testing was never mentioned in the prompt.
Then a second pass. A follow-up prompt added subject-combination input (pick A1, enter maths, physics, English separately). Another ~15 minutes.
And it shipped a bug
The city filter did not work. Filtering for Hà Nội returned Bình Dương.
This is worth stating plainly because the video states it plainly. An agent that cross-checks its arithmetic against source files is not thereby correct about everything else. The verification step covered the numbers — the part with a ground truth — and said nothing about filter logic, which had none.
That is exactly the failure shape to expect: verification covers what you pointed it at. A green cross-check on cutoff scores is not a claim about the UI.
What it cost
| First build | ~15 minutes |
| Second pass (add subject combinations) | ~15 minutes |
| Tokens for the session | ~200,000 |
The creator's own read: it is noticeably slower to generate than other products, and that seems to be the trade — it plans, cross-checks, and re-reads rather than emitting code immediately. Token consumption is high. On the published rate for Fable 5 — $10 per million input, $50 per million output — a 200k-token session is real money, and worth knowing before you point it at something long-running.
On the "comeback" framing
The video opens on the model having previously been held back over security concerns and now being available again, and reviews the vendor's published benchmark table on the way in.
Two notes on that. The framing and the numbers both come from the vendor's own material, not from this run — a benchmark table published by the company that makes the model is marketing until someone independent reproduces it. And the creator's own verdict is more measured than the opening suggests: this release feels weaker than the first one, while still being very strong against what else is available.
I would rather point at the run than the table. The run showed a model that read before planning, caught a formatting error nobody flagged, tested itself without being asked, and still shipped a broken filter. That is a more useful picture of a model than any row of percentages.
What I take from this
Give the task a ground truth if you want to learn anything. The reason this demo is informative is that its output can be checked. If your evaluation task cannot produce a detectably wrong answer, you are measuring fluency.
Ask for the cross-check explicitly. The verification step here was designed in, and it is the only reason "no hallucination" is a statement rather than a hope.
Scope your confidence to what was verified. Numbers checked against files, filters not checked at all — and it was the filter that broke.
Related
- Fable Mode: Getting a Frontier Model to Write the Manual for Its Replacement — the same model, used to write operating rules instead of code.
- Claude Fable 5 vs Claude Opus 5: One Prompt, Two Models, One Excel File — the same two tiers on a single build task, with the bill at the end.
- Session 4 — AI-Engineering Governance — why a green check is only worth its evidence.