Claude Code effort settings: faster drafts or deeper verification

Thariq’s tests suggest Claude Code’s effort slider mostly changes verification, edge-case testing, and independent judgment—not just response length. Low/medium shine for fast iteration, while high/max can pay off for security, code review, and tricky benchmarks.

claude cover

TL;DR

  • Effort levels (low/medium/high/max): Primarily change verification, edge-case testing, and independent judgment; prompt cache remains intact
  • Test basis: Opus 5.5, Fable 5.1, and Terminal Bench 3.0; scores and token use generally rose with effort
  • Workflow recommendation: Spec + interview → build at low/medium → iterate with developer involvement → switch to high for verification/testing
  • Security/edge-case gains: html-js-filter improved 1/5→5/5; higher effort added adversarial review, XSS suite, fuzzing, ~33 minutes runtime
  • Terminal Bench improvements: Opus 5.5 mvcc-lsm-compaction 0/5→4/5; cli-2ph-simple 0/5→5/5 via deeper testing and strategy revision
  • Limits noted: Higher effort reduced missed edge cases, but did not reliably fix incorrect overall strategies

Thariq, writing about Claude Code’s effort settings, argues that changing effort primarily changes how much verification, edge-case testing, and independent judgment Claude applies to a task. The system offers four levels—low, medium, high, and max—without breaking the prompt cache, according to the post.

The analysis draws on tests of Opus 5.5 and Fable 5.1, alongside results from Terminal Bench 3.0. Thariq reports that benchmark scores and token use generally rose as effort increased, although the benefits varied by task.

Low and medium effort were most useful for fast, collaborative work, while higher settings produced better results on tasks involving security, hardware, code review, and other areas with numerous edge cases.

Effort changes how Claude works

Thariq describes effort as an approximation of how much compute Claude should spend on a task. Higher effort does not simply produce a longer response; it gives the model more latitude to make independent decisions, test its work, and reconsider an initial approach.

That distinction was evident in a series of software-building experiments. Asked to create a personal fitness and workout tracker, Claude produced a simple log and graph at low effort. Higher settings resulted in a more detailed application, while max effort added a heat chart.

A similar pattern appeared in a redesign of Claude Code’s /config menu. A low-effort run took about one minute and produced an interactive sketch that conveyed the basic idea. A max-effort run took 28 minutes and generated a more polished mockup resembling Claude Code, along with walkthroughs for several user flows.

The faster version was better suited to an iterative process, according to Thariq. The longer run made more design decisions independently and delivered a more complete result without as much intervention.

When the fitness app was fully specified through an interview beforehand, the differences between effort levels narrowed. The resulting designs and implementations were broadly similar, although the highest-effort run spent additional time simplifying some details.

A proposed workflow for software development

For regular feature work, Thariq recommends combining effort levels rather than relying on the highest setting throughout a project:

  1. Provide a specification and have Claude identify unresolved details through an interview.
  2. Implement the feature using low or medium effort.
  3. Review the result and iterate while remaining involved.
  4. Switch to high effort for verification and testing.

The approach reflects a trade-off identified in the experiments. Lower effort allows Claude to respond quickly and leaves more decisions with the developer. Higher effort can complete more work independently, but also gives Claude more freedom to make assumptions.

Thariq later characterized low effort as the preferred setting for staying involved, while max effort was reserved largely for autonomous work or searching for security vulnerabilities.

Higher effort helped with hidden edge cases

To examine more difficult tasks, Thariq evaluated results from Terminal Bench 3.0, a community-sourced benchmark covering areas including hardware, ML, science, software, operations, and media.

The benchmark includes tasks such as building an 8-bit game console in Verilog for a small FPGA, formally proving Takens’ embedding theorem in Lean 4, consolidating shards of a mixture-of-experts checkpoint, completing EU trade-statistics filing, and recreating a poster as an editable layout file.

According to the analysis, increased effort was especially useful when a task could appear complete while still failing on unusual inputs.

One example involved html-js-filter, which required an HTML sanitizer capable of blocking JavaScript injection techniques. Fable 5.1 reportedly improved from a score of 1/5 at low effort to 5/5 at the highest tested setting.

The low-effort attempts generally wrote a filter in one pass and tested it against a single hand-written page. In a higher-effort run, Claude reviewed its initial implementation adversarially, examined the installed parser’s source, ran clean-input tests, used a standard XSS test suite, and built a random-document fuzzer.

That additional testing took about 33 minutes, compared with roughly two minutes for a typical low-effort attempt. Thariq argues that the extra time is more justified for security reviews, performance work, and other production tasks where overlooked edge cases can have serious consequences.

The benchmark results also suggested a limit to the approach. Higher effort reduced failures attributed to missing edge cases, but did not reliably correct cases where the model had chosen the wrong overall strategy.

Where the trade-off was most visible

Several Terminal Bench tasks showed large differences between lower and higher settings for Opus 5.5.

In mvcc-lsm-compaction, which involved fixing a storage-engine bug without breaking compaction, Opus 5.5 reportedly moved from 0/5 at low effort to 4/5 at the highest tested setting. Low-effort attempts edited the code before building it or reproducing the crash. Higher-effort attempts reproduced the failure first, created a randomized test against a non-compacting reference, and checked whether incomplete fixes failed as expected.

For cli-2ph-simple, a task involving a Python CLI linear-program solver, the reported score rose from 0/5 at low effort to 5/5 at high effort. Lower-effort runs tested a few small problems and stopped without checking performance on larger inputs. Higher-effort runs compared results with a brute-force solver, timed larger cases, and revised the search after encountering crashes and excessive runtimes.

A proteomics analysis task, gsea-proteomics, produced another example. At low effort, Claude selected one plausible data-preparation method and reported the resulting analysis. At high effort, it tried two preparation methods, noticed that the list of significant treatments changed, and investigated the discrepancy before choosing an approach. Thariq notes that a human collaborator might have clarified the setup earlier, while a higher-effort run helped compensate when no such interaction occurred.

Suggested settings

Thariq’s recommendations divide Claude Code’s settings by the desired balance between speed, autonomy, and verification:

  • Low: Brainstorming, sketching, easy changes, and tasks where frequent interaction matters.
  • Medium: Feature development and regular day-to-day software engineering.
  • High: Bug fixing, code review, and work with significant edge-case or verification requirements.
  • Max: Difficult end-to-end projects where Claude is expected to work with minimal input, including security investigations.

The findings come from Thariq’s own experiments and benchmark runs rather than an independent evaluation. They nevertheless suggest that “more effort” is most useful when the task rewards repeated testing and reconsideration—not merely when a longer or more elaborate first draft is preferred.

Source: Thariq on X

Continue the conversation on Slack

Did this article spark your interest? Join our community of experts and enthusiasts to dive deeper, ask questions, and share your ideas.

Join our community