Git Blog

Releasing the Power of Git

Models Are Getting Really Good At Git

I built GitBench because I was curious about how well models did git things. It turns out that it was a good idea because some models will surprise you. Some of the frontier models were remarkably bad and some cheap open models punched above their weight class. But, as they say, “the times they are a changin’”.

Before I go too deep into the latest results, you may want to read up on GitBench and what it is and why I built it in the previous blog post.

The Latest Batch

I needed to run a bunch of newer models. The frontier labs just released brand new models on the 22nd. I missed Fable 5.1 when it came out. Deepseek dropped their latest iteration of 4.1 Flash a little while back. And more. Spoiler alert: Some of the models massively improved on their predecessors.

The Anthropic in the Room

The biggest improvement of any of the newly tested models was Fable. It went from an ~82% pass rate to ~95% when delivering text output. That’s crazy. The JSON output improved too but, even better, it no longer performs worse with more reasoning effort. That behavior was baffling.

Opus also improved on the models that came before it in a massive way. The Opus family of models have been on a downward trend since 4.7. Each new model seemed to get worse at git. 5.5 turns it all around and scores at the top of the range.

OpenAI’s New Models Continue to Excel

The models that OpenAI has been releasing have consistently been in the top 10 performers. Sol regressed in its score. But, Luna improved its score and got cheaper. If you are keeping up with social media and the consensus around Luna, it’s probably your subagent rockstar. Well, now it’s cheaper and better at the micro tasks like git operations. Something I find surprising is that Luna takes longer and burns more tokens but, somehow, costs less over the whole suite.

Muse 1.3 Brings Meta Back Into the Conversation

Meta has been a big part of what has made AI what it is today. They created PyTorch back in 2017 which is still the prominent way of training AI models due to its Nvidia focus and optimizations. Llama was the first big open LLM that seeded the entire ecosystem. So, it’s surprising that their latest Llama releases have been lackluster. Well, the latest iteration brings Meta back into the conversation on artificialanalysis.ai.

While Muse has been pretty good at git since 1.1, Muse 1.3 improves on that score still.

Open Weight Models Are In On It Too

The frontier labs can’t have all the fun. If you’ve seen GitBench before, you would notice that open weight models have been good at git for a while. Open models like Gemma and GPT OSS 20b/120b have been leading the pack in the cost/intelligence quadrant. But some of the darling coding models like Deepseek and GLM have lagged behind. That changes with their latest iterations. GLM 5.3 Flash improved massively from GLM 5.2.

Deepseek 4.1 Flash and GLM 5.3 Flash are both much better than previous much better and the fact that they are flash models means they are cheap, too. Deepseek also did something other providers haven’t done which is optimize their existing model v4 Flash on July 31 after its initial release. But, 4.1 goes even further. Side note, Deepseek 4.1 Flash is my current daily driver on my personal projects.

The Writing Is On The Wall

Looking back at the cover image, with each new batch of models I run, it’s becoming obvious that I’m going to need to create a new more difficult set of fixtures to benchmark models on git operations. I’m not certain what that will even look like. Maybe it’s just more complicated fixtures with larger input contexts. Maybe I need to bring a harness into the equation. If you have suggestions, I’m all ears. But, rest assured that I’ll be posting about it as soon as I have something to report.

Like this post? Share it!

Read More Articles

Visual Studio Code is required to install GitLens.

Don’t have Visual Studio Code? Get it now.

Team Collaboration Services

Secure cloud-backed services that span across all products in the DevEx platform to keep your workflows connected across projects, repos, and team members
Launchpad – All your PRs, issues, & tasks in one spot to kick off a focused, unblocked day. Code Suggest – Real code suggestions anywhere in your project, as simple as in Google Docs. Cloud Patches – Speed up PR reviews by enabling early collaboration on work-in-progress. Workspaces – Group & sync repos to simplify multi-repo actions, & get new devs coding faster. DORA Insights – Data-driven code insights to track & improve development velocity. Security & Admin – Easily set up SSO, manage access, & streamline IdP integrations.
winget install gitkraken.cli