Paweł Huryn’s model-routing comparison assigns different AI models to different types of work after testing them against 105 hidden bugs. His recommendation favors Fable 5 for judgment, strategy, and complex cases; Opus 5 for writing and frontend work; and Luna, Sol, or Grok 4.5 for coding and debugging.
The Bug Hunt Bench v6 covered a VS Code extension and an LMS. Nine frontier models were evaluated across 14 runs, with wall-clock time and costs totaled across both repositories. Fifty-four of the 105 bugs survived every model, while the light areas in the benchmark represent bugs that nobody planted.
GPT-5.6 Sol at maximum effort produced the highest raw result, fixing 42 bugs in one repository and 40 in the other. The two runs took 163 minutes and 48 seconds and cost $69.61. A July 31 rerun at high effort fixed 34 and 28 bugs in 66 minutes and 40 seconds for $33.92.
GPT-5.6 Luna at maximum effort fixed 33 and 31 bugs in 85 minutes and 44 seconds, at a reported cost of $1.80. Luna’s pricing used new rates introduced after OpenAI’s 80% price cut on July 30. At high effort, Luna fixed 13 and 22 bugs in 64 minutes and 5 seconds for $0.57.
The contrast with Fable 5 was substantial. At maximum effort, Fable fixed 29 bugs in one repository and five in the other, taking 57 minutes and 15 seconds at a cost of $104.49. Opus 5 at maximum effort fixed 27 and two bugs for $51.33, while its high-effort run fixed 21 and six for $38.77.
Other results included 21 and six bugs for Kimi K3 at high effort, 16 and seven for Grok 4.5 at high effort, and 13 and five for Grok 4.5 at maximum effort. Sonnet 5, Opus 4.8, and DeepSeek V4-Flash produced lower totals in the reported runs.
Huryn’s conclusion is that the most expensive model in the test “never touches code.” A reply from Kris Patel proposed a funnel approach that would run lower-cost models first, remove bugs they solve, and pass the remainder to more expensive models. The benchmark data does not establish how that blended sequence would perform, but it points to a potential alternative to using a single model across every task.
Source: Paweł Huryn on X
