
Gemini 3.7 Flash: Rank 23 Was Enough
Gemini 3.7 Flash: Rank 23 Was Enough
Two scoreboards, same model, same week.
On Google's launch page, Gemini 3.7 Flash sits at or near the top of every category that matters for agent work: professional code quality, long-horizon software engineering, web development, PDF handling, enterprise automation. Occasionally a competitor edges it out. Mostly it leads.
On arena.ai, an independent leaderboard built from real sessions rather than a vendor's own evaluation, the same model ranks like this:
| Category | Gemini 3.7 Flash |
|---|---|
| Overall | 21st |
| Code | 23rd |
| Work | 24th |
| Chat | 11th |
Anthropic's top model holds first place in every one of those categories.
Neither scoreboard is lying. The interesting question is what to do with a model that its maker calls best-in-class and the independent board calls twenty-third — and the answer only shows up when you actually build something with it.
About this write-up
Written from the video on the channel. The two builds, the timings, the leaderboard positions and the quota readings are what happened on screen. Prices and benchmark placements come from Google's own launch page and are labelled as such. A few product names in the auto-captions were too garbled to reproduce and are described rather than guessed at.
What the price actually is
The launch page puts Gemini 3.7 Flash at roughly $0.75 per million input tokens and $3.75 per million output. Those are vendor figures, not measured ones, but they set up the whole argument.
The video also walks a cost-per-task chart for one long-running agentic job across model generations — a task costing around $15 on the frontier tier, about $8 on Flash two generations back, about $4 on the last one, and well under a dollar here. Read it as a direction of travel rather than a number you can bank: it is the vendor's own chart, and the task is theirs too.
What survives from a vendor chart is only the shape: the interesting claim is not cheaper, it is cheaper by an order of magnitude while staying in the same conversation about capability.
Two builds, one prompt each
Both builds ran the same way: one prompt, three sub-agents, no hand-holding between steps.
The prompt asks for exactly three things — a backend agent that writes the API and logic, a frontend agent that builds the UI against it, and a QA agent that exercises the result and sends failures back — a test suite for the web app, an auto-playing bot for the game. The dispatch rule is the part worth copying: fire the first two together, hold the third until both have produced something.
Run a QA agent in parallel with the builders and it reviews an app that does not exist yet. Run it too late and you have merged two halves that never agreed on a contract. The gate has to sit exactly where the video puts it.
Two details in the prompts do more work than their length suggests.
Every agent is given a target path. /server for the backend, /client for the frontend, /game/core.js and /game/ui.js for the two halves of the game. Parallel agents that share a working directory and no territory will overwrite each other; naming the folder per agent is the whole collision-avoidance mechanism, and it costs one line.
The QA agent is told what to try, not to "test it". Send an empty array, send a string over 100 characters, check the checklist survives switching between dishes. Spawn 150 enemies at once and watch the frame rate, walk the player off the map edge, take damage continuously and check nothing goes invulnerable, open the level-up panel and check the game actually pauses.
That list is a set of predictions about how the build will fail. An agent asked to "verify the app" reports that the app works, because it has no idea what wrong looks like. You have to name the failure you are afraid of before anything will go looking for it.
Build one — a recipe app
SmartChef: pick what is in your fridge, get three dishes back.
Both prompts are reproduced in the locale they were written in — a prompt is an artifact, and translating one changes what was run.
Prompt — Smart Chef: three parallel sub-agents
ACT AS A LEAD SOFTWARE ARCHITECT & MULTI-AGENT ORCHESTRATOR.
PROJECT: "Smart Chef - AI Recipe & Meal Planner Web App"
MODEL ENGINE: Gemini 3.7 Flash
INSTRUCTIONS:
Phân tích yêu cầu dự án và TỰ ĐỘNG KHỞI TẠO (SPAWN) 3 SUB-AGENTS ĐỘC LẬP chạy song song để hoàn thành sản phẩm full-stack:
---
### SUB-AGENT 1: [BACKEND SPECIALIST]
- **Target Folder:** `/server` hoặc `/api`
- **Task:**
1. Viết REST API bằng Express/FastAPI có endpoint `POST /api/recipes` nhận danh sách nguyên liệu `{ ingredients: string[] }`.
2. Tạo logic trả về danh sách 3 món ăn chuẩn JSON schema (gồm: id, title, cooking_time, calories, difficulty, matched_ingredients, missing_ingredients, steps[]).
3. Xử lý an toàn: Validate dữ liệu rỗng và xử lý các nguyên liệu xung đột kỳ lạ.
---
### SUB-AGENT 2: [FRONTEND & UI SPECIALIST]
- **Target Folder:** `/client` hoặc `/src`
- **Task:**
1. Xây dựng giao diện React + TailwindCSS trực quan:
- Khung chọn tag nguyên liệu nhanh (Thịt bò, Trứng, Rau...) + Ô text input nhập tự do.
- 3 thẻ Card hiển thị món ăn kèm badge thời gian, calo và danh sách nguyên liệu thiếu màu cam.
- Danh sách các bước nấu dạng Interactive Checklist (tick hoàn thành có hiệu ứng gạch ngang chữ).
2. Đảm bảo trạng thái tick checklist không bị mất khi chuyển qua lại giữa các món ăn.
---
### SUB-AGENT 3: [QA & TEST AUTOMATION SPECIALIST]
- **Target Task:**
1. Đợi Sub-Agent 1 & 2 sinh code xong, tự động chạy test suite hoặc curl test.
2. Kiểm tra 3 trường hợp: Gửi mảng rỗng `[]`, nhập chuỗi dài >100 ký tự, và kiểm tra tính bền vững trạng thái (state persistence) của checklist.
3. Nếu phát hiện lỗi (FAIL), tự động tạo bug ticket và kích hoạt Sub-Agent tương ứng để sửa lại mã nguồn cho đến khi PASS 100%.
---
EXECUTION RULE:
- Tự động chia task và kích hoạt đồng thời Sub-Agent 1 & 2.
- Sub-Agent 3 sẽ tự động vào can thiệp ngay khi quá trình sinh mã nguồn hoàn tất để thực hiện vòng lặp Self-Healing.
- Bắt đầu thực thi ngay lập tức!| Frontend + backend, in parallel | ~4 minutes |
| QA pass | ~5 minutes |
| Total including the run | ~6 minutes |
| Quota consumed | under 1% of a five-hour limit |
The result: choose beef, tofu and mushrooms, and it returns three dishes — a seaweed and tofu soup, shaking beef with potatoes, tofu in tomato sauce — each split into ingredients you have and ingredients to buy, with numbered steps you tick off, and time, calories and difficulty on the card. Swap the inputs to chicken, pork, potato, carrot and egg and the suggestions change to match.
The creator scores it 8/10, and the reason he gives is worth more than the number: the suggestions are dishes people actually cook. A recipe generator that returns technically valid food nobody makes is a harder failure to notice than a crash.
Build two — a 2D survival game
Cyber Survivor: HTML5 canvas, plain JavaScript, no framework.
Prompt — Cyber Survivor 2D: three parallel sub-agents
ACT AS A LEAD GAME ARCHITECT & MULTI-AGENT ORCHESTRATOR.
PROJECT: "Cyber Survivor 2D - Browser Action Game"
TECH STACK: HTML5 Canvas + Vanilla JavaScript (hoặc Phaser 3) + TailwindCSS (Single-page web game, chạy ngay trên trình duyệt không cần setup phức tạp).
MODEL ENGINE: Gemini 3.7 Flash
INSTRUCTIONS:
Tự động khởi tạo (spawn) 3 Sub-Agents độc lập chạy song song để thiết kế và hoàn thiện toàn bộ trò chơi:
---
### SUB-AGENT 1: [GAME ENGINE & LOGIC SPECIALIST]
- **Target File:** `/game/core.js` hoặc file logic chính.
- **Nhiệm vụ:**
1. Xây dựng Game Loop (60 FPS) với các cơ chế chính:
- Player Movement: Di chuyển nhân vật mượt mà bằng phím WASD / Phím mũi tên.
- Auto-Shooting: Tự động bắn đạn về phía quái vật gần nhất theo chu kỳ.
- Enemy Spawner: Sinh quái vật theo từng đợt (Wave), quái tự động tìm đường đuổi theo Player.
- Collision Detection: Xử lý va chạm giữa đạn - quái và quái - player (mất máu, tính điểm XP, rơi vật phẩm tăng cấp).
2. Cân bằng chỉ số (Game Balancing): Tăng dần tốc độ và số lượng quái theo thời gian sống sót.
---
### SUB-AGENT 2: [UI/UX, SPRITES & SOUND SPECIALIST]
- **Target File:** `/game/ui.js` & `index.html`
- **Nhiệm vụ:**
1. Thiết kế giao diện phong cách Cyberpunk Retro:
- Màn hình Game HUD: Thanh máu (HP Bar), Thanh kinh nghiệm (XP Bar), Điểm số (Score), Đồng hồ sống sót (Survival Timer).
- Hiệu ứng hình ảnh Canvas (VFX): Hiệu ứng nổ hạt (Particle effects) khi quái bị tiêu diệt, màn hình nhấp nháy đỏ khi nhận sát thương.
- Tạo pop-up "Level Up - Chọn 1 trong 3 nâng cấp" (Tăng tốc bắn, Tăng máu, Đạn chùm).
2. Tích hợp âm thanh Web Audio API đơn giản (tiếng bắn, tiếng nổ, tiếng nhặt đồ bằng code âm tần tổng hợp, không cần tải file ngoài).
---
### SUB-AGENT 3: [GAMEPLAY TESTER & BALANCE QA]
- **Target File:** `/game/tester.js` hoặc kịch bản kiểm thử tự động.
- **Nhiệm vụ:**
1. Chạy giả lập Auto-Play (Bot tự chơi) để stress-test:
- Test Lag / Drop FPS: Sinh đồng thời 150 quái vật trên màn hình để kiểm tra hiện tượng tụt khung hình.
- Test Edge Cases: Nhân vật đi ra ngoài mép bản đồ (Out of bounds), quái bị kẹt góc không di chuyển, hoặc lỗi bất tử khi nhận sát thương liên tục.
- Test Level-up Freeze: Kiểm tra game có pause chính xác khi mở bảng chọn nâng cấp không.
2. Báo cáo lỗi và yêu cầu Sub-Agent 1 & 2 vá code ngay lập tức nếu game bị giật hoặc phát hiện bug logic.
---
EXECUTION RULE:
- Sub-Agent 1 & 2 làm việc song song để đồng bộ giữa Canvas Rendering và Game Logic.
- Sub-Agent 3 chạy ngay sau khi dựng xong game base để cân bằng chỉ số và fix bug.
- Đầu ra là một file `index.html` (hoặc project bundle) mở lên là chơi được ngay lập tức!This one produced an architecture plan first — UI layer, game logic layer, game loop — before writing anything. The plan is detailed enough to review, which is the difference between an agent you can supervise and one you can only watch.
Then it built the game, and the game ran on the first attempt: a top-down survival shooter with HP, survival timer, level and score, and power-ups that take you from one projectile to two to three. No crashes during play.
Both builds together came to about 8% of the five-hour quota.
The parts the demo does not flatter
Approvals are manual unless you turn them off. The run was not in auto mode, so the agents stopped and waited for npm install to be approved — twice. That is the right default, and it also means "six minutes" is six minutes of someone being present.
Sub-agents cannot be corrected mid-flight. Once dispatched, there is no channel to send them a message; you can watch what they are doing and that is all. If a sub-agent starts down the wrong path, the intervention available to you is to let it finish and re-prompt.
Neither build is hard. A recipe suggester and a canvas shooter are well-trodden shapes with enormous amounts of prior art. They demonstrate that the agent loop holds together; they do not probe the ceiling. A model ranked 23rd for code and a model ranked 1st would both be expected to clear this bar — which is exactly why the result is interesting rather than damning.
What rank 23 actually means
Here is the reconciliation, and it is the thing I would take away.
A leaderboard ranks models against each other on the hardest thing they can be asked. That is a useful measurement, and it is not the measurement most work needs. Most work has a threshold: either the model clears the task or it does not. Above that line, rank stops changing the outcome and only cost and latency keep moving.
Two builds, first try, no crashes, under 10% of a quota is evidence that for this class of task, rank 23 is comfortably above the line. It is not evidence that the model is secretly first — the leaderboard is measuring something real, and on a genuinely hard problem those twenty-two places would show up as failed runs and burned hours.
So the honest framing is narrow and useful: pick by threshold, not by rank. Work out whether the cheap model clears your particular bar, and if it does, the twenty-two models above it are selling you headroom you are not using.
The corollary is the uncomfortable half: you only learn where your bar sits by running the cheap model at a real task and watching it fail. That is a cost too, and it is the one nobody charts.
What I take from this
Dispatch two, gate on the third. Parallel builders plus a QA agent that starts only after both have output is the structure doing the work here, not the model.
Ask for the plan before the code. The game build produced a reviewable architecture first, which is the only point where correcting it is cheap — especially since sub-agents cannot be corrected once running.
Vendor benchmarks and independent boards answer different questions. One says how good it can be, the other how it places among peers. Neither says whether it clears your bar.
Cheap changes what you are willing to try. At under 1% of a quota per app, the calculation stops being "is this the best model" and becomes "is there any reason not to attempt this" — and that shift matters more than eight places on a leaderboard.
Related
- Google Antigravity 2.0: The Settings Screen Is the Story — the platform these builds ran on, and what a scheduled agent is allowed to touch.
- OpenAI Codex: The Local Folder Is the Product — another agent platform, and the permission model underneath it.
- Claude Fable 5 vs Claude Opus 5: One Prompt, Two Models, One Excel File — the same question from the other end: what you give up by paying less.
- Session 4 — AI-Engineering Governance — why a green check is only worth the evidence behind it.