
OpenAI Codex: The Local Folder Is the Product
OpenAI Codex: The Local Folder Is the Product
The common assumption about Codex is that it codes and fixes bugs, and that is the extent of it. Watch the walkthrough and a different shape emerges: reading documents, analysing requirements, planning work, remembering context, automating recurring jobs, connecting to third-party tools. None of that is coding.
But listing features is the wrong way to understand this tool, because the features are not independent. Nearly all of them are consequences of a single design decision:
A Codex project is a real folder on your machine, and Codex operates on it directly.
ChatGPT works on a copy you uploaded to the cloud. Codex works on the original. Once you accept that one sentence, projects, the generated spreadsheet, the surprising behaviour of "delete", the memory files and the two settings screens all stop being a feature list and start being one idea with five faces.
About this write-up
Written from the video on the channel. The walkthrough, the Linear integration and the scheduled job are what happened on screen. Two things are labelled rather than asserted: the model name in the auto-captions was garbled beyond reconstruction, so the model picker is described by behaviour; and the subscription price quoted in the video is the creator's own figure, not something verified against a pricing page.
Getting in: three shapes, two billing models
Codex ships in three forms, and which one you pick changes the mental model more than the capability:
| Form | How you get it | Feels like |
|---|---|---|
| Desktop app | download for macOS or Windows | a ChatGPT window that can touch your disk |
| IDE extension | VS Code, Cursor, JetBrains, Vim | an assistant inside the editor |
| Terminal | an npm install | a CLI agent |
The walkthrough uses the desktop app, which is the version where the "local folder" idea is most visible.
Sign-in offers two paths, and they are two different cost structures rather than two doors to the same room:
- Continue with ChatGPT — a paid plan, fixed monthly cost regardless of how hard you push it.
- Enter API key — metered per token: you pay for exactly what you consume.
The video takes the subscription and quotes roughly $20 a month for it. Treat the number as the creator's own, and check the current plan tiers yourself before budgeting — the point that survives either way is the shape of the choice. A fixed monthly cost makes long, exploratory, agent-driven sessions psychologically free; metered billing makes you flinch every time the agent decides to re-read a folder. That difference will change how you use the tool more than any feature on this page.
The settings screen you should not skip
Most tool tours treat settings as throat-clearing before the demo. Here it is the opposite, and the reason is the sentence at the top of this article: an agent that operates on your real filesystem has a real blast radius.
Appearance is genuinely cosmetic — light, dark or custom themes, a font picker, imported themes. Configuration is not.
Approval policy — when does it stop and ask you?
Four settings, in descending order of how much they protect you:
| Behaviour | What it means | Worth using? |
|---|---|---|
| Agent judgement | runs what it considers safe, stops to ask on anything it thinks touches your system | the video's choice, and a reasonable default |
| Always ask | approval prompt on every action that touches data | safe, and slow enough that you will get click-fatigue |
| Ask on failure | auto-approves everything, comes back to you only when something errors | see below |
| Never ask | executes everything, silently | not recommended |
The video's objection to ask on failure is the sharpest observation in the whole settings tour, and it is worth restating because the option sounds prudent: by the time an action has failed, it has already run. An approval prompt that arrives after execution is not an approval prompt — it is a notification. The failure you wanted to prevent has already happened; you are just being told about it, and rolling back is now your problem.
Never ask is dismissed for the obvious reason. That leaves the top two as real choices, and the trade between them is entirely about how much you trust the agent's own risk judgement.
Sandbox — what can it reach?
Three settings, and this is the axis that actually bounds the damage:
- Read only — the agent reads files and cannot modify or delete anything. The safe setting.
- Workspace write — it can create, edit and delete, but only inside the folder you pointed it at. Outside that folder it cannot touch anything. This is the video's choice for the project.
- Full access — it can modify and delete outside the workspace, including sensitive system files. Dismissed as too risky.
Note how the two dials compose. Approval policy controls how often you are interrupted; sandbox controls how bad it can get if you approve the wrong thing. Read-only plus never-ask is safer than full-access plus always-ask, because the first pair makes destructive action impossible and the second only makes it slow. If you only have the attention to set one of them deliberately, set the sandbox.
All of this is written to a config file you can open and read — which is the right design, because a permission model you cannot inspect in plain text is a permission model you are trusting on faith.
Projects: the same idea, five faces
Create a project — from scratch or by pointing at an existing folder — and Codex binds the conversation to that directory. Whatever is in there is in scope: spreadsheets, PDFs, slide decks, an entire codebase. Whatever is outside it is not.
Inside a project you can run several conversations at once. Each shows a spinner while it works and a green dot when it finishes, so parallel work is legible rather than something you have to track in your head.
The demo makes the difference concrete in the least dramatic way possible. First a plain summarisation prompt, of the kind ChatGPT answers identically. Then: create an Excel file to save this information.
Codex writes a build script and produces the spreadsheet — overview, detail and sources — in the workspace, on the local disk. Not a download link. Not a file in a cloud sandbox that expires. A file in your folder, where your other files are, which your other tools can already open.
You can then open it in Excel or Google Sheets, or view it inside Codex directly, and ask questions about its contents in place. That last part is a nice convenience. The part that matters is the one before it: the artefact landed where artefacts live.
"Delete" does not delete
The search feature looks like housekeeping until the demo turns it into something more interesting.
A project is deleted from the Codex interface. Then search finds it again — because the project was never in Codex to begin with. It was a folder on disk that Codex had a pointer to. Re-point at the original directory and the conversation comes back intact.
This is worth sitting with, because it is genuinely two things at once and the video only says the happy half:
- As recovery, it is excellent. Removing something from the UI cannot cost you work. There is no "are you sure?" that can ruin your afternoon.
- As data hygiene, it is a trap. If your mental model is "I deleted that", you are wrong, and the material is still sitting in a folder — which matters the moment the project contained anything you would not want lying around.
Removing a project from Codex is closing a pointer, not deleting data. Both halves of that follow from the same design, and you should know which half you are relying on.
Plugins and skills: connection versus instruction
These two words get used interchangeably in most tooling, and here they mean genuinely different things. Getting the distinction right is the difference between an agent that can technically reach your tools and one that uses them correctly.
A plugin is integration — an MCP server or API connection to a third-party tool: GitHub, Linear, Google Drive, Gmail, Teams. It answers what can the agent reach?
A skill is a markdown instruction file. It defines rules and guidance for how the agent should work. It answers how should the agent behave once it can reach that thing? Connect GitHub via a plugin, then write a skill that encodes your branch-naming convention and your git flow, and the agent stops inventing its own conventions per project.
Codex ships a catalogue of plugins, and installing one brings default skills with it, which you can open and read before trusting them. You can also write your own — and you should, because the default skill encodes somebody else's house style, not yours.
The Linear demo, and the detail that makes it worth watching
Connecting Linear takes the shape you would expect: install, and get two components — the app (the MCP connection) and the skill (which teaches the agent how to interact with Linear). Enable both, click connect, authenticate through the web, approve, choose a workspace. Verify it in the Manage screen.
Then a prompt asking for a Linear issue to be created, with a title and an instruction to use the REST API — and the prompt never mentions Linear as a tool. Codex reads the request, recognises that it belongs to the Linear plugin and skill, routes it, and creates the issue. Reloading Linear confirms it landed.
That is the payoff of the distinction. The plugin made Linear reachable; the skill made the request recognisable as a Linear request. Without the skill you are back to naming your tools by hand in every prompt, which is only a small annoyance until you have eight of them connected.
Automation: three parts, one scheduled report
Automation here is standard in structure and worth stating explicitly because people skip the vocabulary:
| Part | In the demo |
|---|---|
| Trigger | 9am, every day |
| Condition | none — the schedule is the whole condition |
| Action | check Linear issues and review the tasks Codex handled |
You fill in a form, name the automation, set the schedule, and it registers as an active job with a next-run time. You can also fire it manually with Run now rather than waiting for the clock, which is the difference between a scheduler you can test and one you have to trust.
One small detail with a general lesson in it: the prompt never specified an output format, so Codex produced a markdown report. Not an error — a default. The same principle as anywhere else in agentic work: an unspecified requirement is not an absent one, it is one the agent will decide for you.
And this is where the settings screen earns its place at the top of this article. A job that fires at 9am, when you are not at the desk, with write access to a folder, is a different risk profile from a chat window you are watching. Sessions four and five of the SDLC notes make the same argument from the governance side, and the Antigravity 2.0 write-up hits it from the other direction: the moment a tool can run unattended, the permission model is the product surface.
Memory: one you write, one you should only read
Codex splits memory in two, and the useful advice is different for each.
Manual memory is yours. Markdown files you write, edit and update by hand, or have the agent draft for you — and you can have several, one per level of your folder structure, so a rule can be global or scoped to a single project. In the demo, Codex generates one from the project's local history, with sections for scope and current project state. It lives on your disk, in your workspace, next to the code it describes.
Auto memory is the agent's. Codex learns from your conversations — how you work, what you prefer, your style — and writes it down without being asked, by default under ~/.codex/memory.
The recommendation in the video is a good one and it is asymmetric: do not edit auto memory — let the agent own it — but do read it every week or two to see what it has concluded about you.
That habit is worth adopting for a reason the video does not spell out. Auto memory is the part of the system that silently changes the agent's behaviour over time. If it learns something slightly wrong about how you work, nothing announces it — you just get answers that are subtly off, for weeks, with no visible cause. A fortnightly read is not housekeeping. It is the only feedback loop you have on a component that is otherwise invisible.
And there is a robot pet
There is a small robot you place anywhere on screen. When you switch to another tab or app, it notifies you that your run has finished. You can pick which pet from Appearance settings.
This is a toy, and it solves a real problem. Agent runs take long enough that you leave, and long enough that you forget you left. An ambient indicator that survives you switching away is the difference between a thirty-second wait and a fifteen-minute one you did not notice.
What I take from this
Set the sandbox before you set anything else. Approval policy decides how often you are interrupted; the sandbox decides how bad it gets when you approve something without reading it.
"Ask me on failure" is not an approval mode. The action already ran. You are configuring a notification and calling it a safeguard.
Learn the plugin/skill split properly. Plugin is reach, skill is behaviour. The Linear demo works — with no tool named in the prompt — precisely because both were in place.
Anything you leave unspecified, the agent specifies. No format requested, so it chose markdown. That is benign here and it will not always be.
Read your auto memory on a schedule. It is the one part of the system that changes your results without telling you.
Related
- Google Antigravity 2.0: The Settings Screen Is the Story — the same argument about unattended agents, from a tool with a scheduler and shell access.
- Claude Fable 5 Is Back: A Demo That Can Actually Be Wrong — an agent that cross-checks its own output, and the bug that check did not cover.
- A Second Brain That Maintains Itself: Claude Fable 5 + Obsidian — another agent pointed at a real local folder, kept honest by an explicit contract.
- Session 4 — AI-Engineering Governance — automate the execution, keep the accountability.