A small, real-repository coding benchmark
Notes from testing local coding models against private repos, imperfect prompts, and real implementations.
Leaderboard
04 models rankedMean blinded-judge scores on a 0–4 rubric. Each model links to its full benchmark writeup.
| Rank | Model | Open book | Closed book | Overall | Tokens / run |
|---|---|---|---|---|---|
| 01 | Qwen3.6-27B4 bit | 2.87 | 3.60 | 3.23 | 4,851 |
| 02 | Qwen3.6-27B Fable-Fusion4 bit | 2.83 | 3.37 | 3.10 | 3,397 |
| 03 | Tess-4-27B4 bit | 2.50 | 3.50 | 3.00 | 7,585 |
| 04 | Qwopus3.6-27B-Coder4 bit | 2.53 | 2.93 | 2.73 | thinking off |
