Google's newest workhorse model landed yesterday with real numbers attached. Here is how it compares to GPT-5.6 on the things that matter at a desk, using only figures the vendors and coverage actually published.
Google shipped Gemini 3.6 Flash yesterday, and for once a model launch came with numbers you can hold onto. Per 9to5Google's launch coverage: 17% fewer output tokens than its predecessor, 83% on the OSWorld-Verified computer-use benchmark, and API pricing of $1.50 per million input tokens and $7.50 per million output. Its natural rival is GPT-5.6, which OpenAI released on June 26 in four variants: Sol, Sol Ultra, Terra, and Luna. So which one should be doing your work?
A ground rule before the table. This comparison uses only figures published by the vendors or in the launch coverage linked here. Where OpenAI has not put a comparable number next to Google's, we say so instead of guessing, because a spec sheet with invented cells is worse than no spec sheet. What the sources give us is lopsided: hard numbers on one side, a variant lineup on the other. That asymmetry is itself information, and we will use it.
| Gemini 3.6 Flash | GPT-5.6 | |
|---|---|---|
| Released | July 21, 2026 | June 26, 2026 |
| Lineup | 3.6 Flash, plus 3.5 Flash-Lite alongside | Four variants: Sol, Sol Ultra, Terra, Luna |
| API price, input | $1.50 per 1M tokens ($0.30 Flash-Lite) | No comparable figure in our sources |
| API price, output | $7.50 per 1M tokens ($2.50 Flash-Lite) | No comparable figure in our sources |
| Computer use (OSWorld-Verified) | 83% | Not published in our sources |
| Efficiency claim | 17% fewer output tokens than predecessor | None cited |
| How consumers pick it | Gemini app model menu | Effort levels: Instant, Medium, High, Extra High, Pro |
Those empty GPT-5.6 cells are not a verdict against OpenAI. They reflect how differently the two companies now talk. Google launched 3.6 Flash the way you launch a commodity: here is the price, here is the benchmark, do the math. OpenAI increasingly launches models the way you launch a product line, four named variants and, since June 10, a ChatGPT picker that hides model names entirely behind five effort levels: Instant, Medium, High, Extra High, and Pro. One company is selling you tokens. The other is selling you outcomes and would rather you stopped reading spec sheets altogether.
Token efficiency sounds like an accounting detail. It is closer to a personality trait. A model that needs 17% fewer output tokens to do the same job is a model that answers with less throat-clearing, which you experience as faster responses and shorter, denser output. If you have ever asked an AI a yes-or-no question and received four paragraphs and a bulleted summary, you already care about output token count. You just did not know its name.
For API users the meaning is blunter: output tokens are the expensive kind, at $7.50 per million on 3.6 Flash against $1.50 for input. A model that emits 17% fewer of them cuts the priciest line of your bill by roughly that share before you change a single prompt. Compounded across a product doing millions of calls, that is not a rounding error. It is a hiring decision.
OSWorld-Verified tests whether a model can operate a real computer: open apps, click through interfaces, finish multi-step tasks. Scoring 83% means the model completes the large majority of those tasks unassisted, and it matters because computer use is the substrate of every agent feature vendors are now shipping. A model that mishandles one screen in five is an agent you babysit. The gap between 83% and whatever your patience requires is the gap between a demo and a delegation.
Worth being precise about the caveat, though: 83% on a benchmark is not 83% on your work. Benchmarks use contained tasks with checkable answers. Your expense system has a popup the benchmark never met. Treat the number as evidence Gemini is serious about agent reliability, not as a promise about Thursday. Our computer-use agents comparison from this week covers where all three big ecosystems stand on letting AI actually drive.
If you pay for a chat app: the API prices above do not touch your bill, which is flat-rate either way. What the launch changes for you is speed and verbosity, and there the model-picker philosophies matter more than the models. ChatGPT since June wants you to choose an effort level and forget model names. Gemini still lets you pick the model. Whether you find that respectful or tedious is genuinely a matter of taste; we know readers in both camps.
If you pay per token: Google just handed you a very legible menu. Gemini 3.6 Flash at $1.50 in and $7.50 out for capable everyday work, and 3.5 Flash-Lite at $0.30 and $2.50 for high-volume jobs that do not need the top model, a fivefold price gap you can route between per request. Against GPT-5.6 you will have to check OpenAI's current pricing page yourself, since none of the launch coverage we cite puts a per-token figure beside it. That checking is worth an hour of someone's time before signing anything: at these volumes, pricing differences compound faster than quality differences.
Benchmarks in hand, here is our reasoned read on where each fits, stated as assessment rather than measurement, since neither vendor publishes head-to-head numbers for most everyday tasks.
One more piece of advice that beats any table: run your own bake-off, and keep it small. Take three real tasks from last week, the email you rewrote four times, the document you summarized, the spreadsheet question you gave up on, and put the same prompt through both models. Twenty minutes, no benchmark required, and the result is calibrated to your work instead of to a lab's task list. In our experience the winner of that little contest varies by person far more than the launch coverage would suggest, which is exactly why we resist declaring one model better at "writing" for everybody.
For everyday work in a chat app, pick the ecosystem you already live in and stop worrying; nothing in the published record says a Workspace user should defect to ChatGPT this week or vice versa. For builders and anyone buying tokens, Gemini 3.6 Flash is the easier recommendation today for one unfashionable reason: Google showed its numbers, and numbers you can see are numbers you can plan around. And hovering over all of it is the launch detail that matters most and got the least attention: Google confirmed Gemini 4 pre-training is already underway. Both of these models are the middle of a story, not the end of one. Buy for the next two quarters, keep your setup portable, and let the vendors keep racing on your behalf.
Gemini 3.6 Flash, launched July 21, 2026 alongside Gemini 3.5 Flash-Lite, uses 17% fewer output tokens than its predecessor and scores 83% on the OSWorld-Verified computer-use benchmark. API pricing is $1.50 per million input tokens and $7.50 per million output tokens, with Flash-Lite at $0.30 and $2.50. Google also confirmed Gemini 4 pre-training is underway.
GPT-5.6, released June 26, 2026, ships in four variants: Sol, Sol Ultra, Terra, and Luna. Most ChatGPT users never pick between them directly, because since June 10, 2026 the ChatGPT model picker has been simplified to five effort levels: Instant, Medium, High, Extra High, and Pro.
Not directly to your bill, since consumer subscriptions are flat-rate. It matters indirectly: a model that says the same thing in 17% fewer output tokens answers faster and pads less, and it lowers costs for the developers building the AI features inside the other apps you use. Where token efficiency directly saves money is API usage, where you pay per token.