I built GitBench because I was curious about how well models did git things. It turns out that it was a good idea because some models will surprise you. Some of the frontier models were remarkably bad and some cheap open models punched above their weight class. But, as they say, “the times they are a changin’”.
Before I go too deep into the latest results, you may want to read up on GitBench and what it is and why I built it in the previous blog post.
The Latest Batch
I needed to run a bunch of newer models. The frontier labs just released brand new models on the 22nd. I missed Fable 5.1 when it came out. Deepseek dropped their latest iteration of 4.1 Flash a little while back. And more. Spoiler alert: Some of the models massively improved on their predecessors.
The Anthropic in the Room
The biggest improvement of any of the newly tested models was Fable. It went from an ~82% pass rate to ~95% when delivering text output. That’s crazy. The JSON output improved too but, even better, it no longer performs worse with more reasoning effort. That behavior was baffling.
Opus also improved on the models that came before it in a massive way. The Opus family of models have been on a downward trend since 4.7. Each new model seemed to get worse at git. 5.5 turns it all around and scores at the top of the range.
OpenAI’s New Models Continue to Excel
The models that OpenAI has been releasing have consistently been in the top 10 performers. Sol regressed in its score. But, Luna improved its score and got cheaper. If you are keeping up with social media and the consensus around Luna, it’s probably your subagent rockstar. Well, now it’s cheaper and better at the micro tasks like git operations. Something I find surprising is that Luna takes longer and burns more tokens but, somehow, costs less over the whole suite.
Muse 1.3 Brings Meta Back Into the Conversation
Meta has been a big part of what has made AI what it is today. They created PyTorch back in 2017 which is still the prominent way of training AI models due to its Nvidia focus and optimizations. Llama was the first big open LLM that seeded the entire ecosystem. So, it’s surprising that their latest Llama releases have been lackluster. Well, the latest iteration brings Meta back into the conversation on artificialanalysis.ai.
While Muse has been pretty good at git since 1.1, Muse 1.3 improves on that score still.
Open Weight Models Are In On It Too
The frontier labs can’t have all the fun. If you’ve seen GitBench before, you would notice that open weight models have been good at git for a while. Open models like Gemma and GPT OSS 20b/120b have been leading the pack in the cost/intelligence quadrant. But some of the darling coding models like Deepseek and GLM have lagged behind. That changes with their latest iterations. GLM 5.3 Flash improved massively from GLM 5.2.
Deepseek 4.1 Flash and GLM 5.3 Flash are both much better than previous much better and the fact that they are flash models means they are cheap, too. Deepseek also did something other providers haven’t done which is optimize their existing model v4 Flash on July 31 after its initial release. But, 4.1 goes even further. Side note, Deepseek 4.1 Flash is my current daily driver on my personal projects.
The Writing Is On The Wall
Looking back at the cover image, with each new batch of models I run, it’s becoming obvious that I’m going to need to create a new more difficult set of fixtures to benchmark models on git operations. I’m not certain what that will even look like. Maybe it’s just more complicated fixtures with larger input contexts. Maybe I need to bring a harness into the equation. If you have suggestions, I’m all ears. But, rest assured that I’ll be posting about it as soon as I have something to report.
GitKraken MCP
GitKraken Insights
