Skip to content
← all posts
·3 min read·by Dru Edwards·#ai #models #craft #cost

The Effort Toggle Era

Frontier models are shipping with dials for how hard they think: a five-level effort toggle on one, tunable thinking on another. Reasoning is becoming a setting you pay for by the notch, so the skill that matters is knowing when to turn it up.

Reasoning is now a dial. You pay for it by the notch, so the skill that matters is knowing when to turn it up.

Two of the recent frontier releases shipped with the same idea from different labs: a dial for thinking. One has a five-level effort toggle. The other has tunable thinking levels. Same concept. Not every question deserves the full furnace, so now you get to say how much fire you want.

what shipped

The mechanics are simple. Low effort: fast, cheap, good enough for the easy stuff. High effort: the model thinks longer, checks its own work, burns more compute, and costs you more per answer. You pick the level per call.

This is overdue. For two years we have been paying full price for every query, the same price for "summarize this paragraph" as for "design the database schema for a multi-tenant billing system." That was always absurd. It was like paying for a charter flight to go to the grocery store. The toggle fixes the pricing to match the problem.

why this was inevitable

Most queries are easy. That is the open secret of everyone running models in production. The overwhelming majority of real-world calls are simple: classify this, extract that, rewrite this paragraph, answer this FAQ. A small model at low effort handles them fine. The expensive reasoning was always overkill for the bulk of the workload, and the labs knew it. They were just selling it bundled because bundled is simpler.

Unbundling it is better for everyone except the people who were collecting the margin on the bundle. You get cheaper answers for easy work. The lab gets to stop subsidizing your grocery runs with everyone else's charter flights.

what changes for builders

Here is the practical part. You now budget cognition the way you budget compute, because it is compute.

My rule, and I am keeping it simple on purpose: default to low effort. Step it up only when the answer fails. I have a small test set that looks like my real work, and a new effort level earns its place by clearing a bar the cheaper one missed. Not by feeling smarter. By measuring better.

The failure mode to watch is the obvious one. People will crank the dial to max for everything because max feels safer, and then wonder why the bill tripled. That is not a model problem. That is a discipline problem. The dial is a tool. Tools do not make decisions.

There is a second-order effect worth naming. When reasoning has a visible price, people start asking whether the thinking was worth it. That is a healthy question, and it is one the industry has been avoiding. "The model thought really hard about it" is not a unit of value. The answer being right is the unit of value.

the honest read

This is one of those changes that looks small and compounds. It will not get keynote slides. Nobody is going to make a hype video about a dropdown menu. But a year from now, every serious deployment will be routing queries by difficulty without thinking about it, the same way we route storage by access pattern today. Hot data, cold data. Hard questions, easy questions. Pay accordingly.

The models keep getting smarter. Now they are also getting cheaper to use wisely. That second trend matters more than the first for anyone actually shipping something.