I Benchmarked the Two Cheapest Coding Models My Company Allows. One of Them Lies With Confidence.
Like a lot of companies, mine doesn’t let developers use whatever AI model they want. There’s a list. Admins enable models one at a time, and the expensive frontier models are rationed carefully. So I did what most engineers in that position do: I picked two cheap models and decided to hand them all the routine work. The boilerplate. The “what does this method do” questions. The small refactors. Save the expensive model for the hard stuff. The two I picked were Kimi K2.7 Code and MAI-Code-1.1-Flash , both selectable in GitHub Copilot. And then curiosity got the better of me. Before I trusted either of them with my daily work, I wanted to know which one was actually smarter — not according to the vendors’ own benchmarks, but on tasks I care about. So I wrote five questions and asked both. The result surprised me more than I expected. The two models They are not really competitors. They sit in different weight classes. Kimi K2.7 Code is made by Moonshot AI, a lab based in Beijing....