A small, real-repository coding benchmark

Notes from testing local coding models against private repos, imperfect prompts, and real implementations.

Leaderboard

04 models ranked

Mean blinded-judge scores on a 0–4 rubric. Each model links to its full benchmark writeup.

RankModelOpen bookClosed bookOverallTokens / run
01Qwen3.6-27B4 bit2.873.603.234,851
02Qwen3.6-27B Fable-Fusion4 bit2.833.373.103,397
03Tess-4-27B4 bit2.503.503.007,585
04Qwopus3.6-27B-Coder4 bit2.532.932.73thinking off