Bug Hunt Bench ranks coding AIs by real repo fixes
Paweł Huryn has published Bug Hunt Bench, a new benchmark testing whether coding models can find and fix planted bugs in real repositories. The early leaderboard shows GPT-6 Astra Max leading with 48 of 105 verified fixes, while runtime, cost, and unplanted defects vary widely.