Home · Sports · Oct 7 archive
Chinese AI Model Kimi-K3 Surpasses US Competitors in Programming Challenge
Confirmed
In Short: An open weights Chinese model, Kimi-K3, developed by Moonshot AI, has surpassed US competitors Claude and Gemini in a programming challenge, according to a report by Futurism.

In the 'Frontend Code Arena,' a measure of an LLM's ability to perform multi-step web development tasks, Kimi-K3 went from number 17 to number one, outperforming Claude Fable 5 and GPT-5.6 Sol.
The model's performance in the 'Text Arena,' which evaluates an LLM's ability in text-to-text tasks like creative writing, also improved significantly, earning the number nine spot.
Kimi-K3's success is notable because it is an open-weight model, meaning its inner workings are viewable to the public, unlike proprietary models like GPT-5.6 that are kept under lock and key.
According to Futurism, the model's performance raises questions about the justification for the high pricing of closed-source models like those from Anthropic.
Futurism quoted an unnamed source as saying, 'When the best open weight model exceeds the best closed-source model, how does [Anthropic] justify its Fable pricing? Why would anyone pay for that?'
The benchmark results come as Chinese AI models are increasingly closing the gap with Western counterparts, with some offering open weights under different licenses.
Despite Kimi-K3's success, it does not beat ChatGPT, Gemini, or Claude outright, trailing them by just a few points on the Intelligence Index.
The performance gap is closing fast, making it less necessary for companies or individuals to pay exorbitant prices for Silicon Valley's frontier models.
This development highlights the growing competitiveness of Chinese AI models, which are often more affordable and accessible due to their open-weight nature.
In a separate development, Mistral AI's CEO Arthur Mensch claimed at the Ai Everything conference in Abu Dhabi that the company's newest model outperforms Chinese models on certain aspects, including cybersecurity.
However, Mensch did not provide specific model names, benchmarks, or scores to substantiate his claim, leaving the assertion unverifiable.
The rise of open-weight models like Kimi-K3 underscores the changing landscape of AI development, where transparency and accessibility are becoming key factors in model adoption.
Background
Chinese AI developer Moonshot is conducting an internal review after researchers were able to persuade two of its popular Kimi models to provide information on how to make biological weapons and carry out assassinations.
What's confirmed
- In the 'Frontend Code Arena,' a measure of an LLM's ability to perform multi-step web development tasks, Kimi-K3 went from number 17 to number one, outperforming Claude Fable 5 and GPT-5.6 Sol.
- The model's performance in the 'Text Arena,' which evaluates an LLM's ability in text-to-text tasks like creative writing, also improved significantly, earning the number nine spot.
- Kimi-K3's success is notable because it is an open-weight model, meaning its inner workings are viewable to the public, unlike proprietary models like GPT-5.6 that are kept under lock and key.
- According to Futurism, the model's performance raises questions about the justification for the high pricing of closed-source models like those from Anthropic.
- The benchmark results come as Chinese AI models are increasingly closing the gap with Western counterparts, with some offering open weights under different licenses.
- Despite Kimi-K3's success, it does not beat ChatGPT, Gemini, or Claude outright, trailing them by just a few points on the Intelligence Index.
- The performance gap is closing fast, making it less necessary for companies or individuals to pay exorbitant prices for Silicon Valley's frontier models.
- This development highlights the growing competitiveness of Chinese AI models, which are often more affordable and accessible due to their open-weight nature.
- In a separate development, Mistral AI's CEO Arthur Mensch claimed at the Ai Everything conference in Abu Dhabi that the company's newest model outperforms Chinese models on certain aspects, including cybersecurity.
- However, Mensch did not provide specific model names, benchmarks, or scores to substantiate his claim, leaving the assertion unverifiable.
- The rise of open-weight models like Kimi-K3 underscores the changing landscape of AI development, where transparency and accessibility are becoming key factors in model adoption.
What's still developing
- Anthropic is bringing Claude directly into Google Docs, Sheets, and Slides through a new Google Workspace add-on, which is now available in beta for Claude Pro, Max, Team, and Enterprise users.
- The default “Ask before edits” mode shows an approval card before applying an edit, while “Accept all edits” lets Claude work through a task without stopping for confirmation each time.
- Anthropic is launching Google Docs, Sheets, and Slides connectors for Claude, which means you can start the job from Claude instead of from Workspace.
- Claude is starting to blur the line between asking an AI for help and actually getting the work done inside the file you need.
- Using Claude alongside Google Docs has always meant too much jumping back and forth.
- You ask Claude to rewrite a paragraph, copy the result into your document, fix the formatting if something breaks, and then repeat the whole process the next time you need help.
- And this goes well beyond asking Claude questions about what’s on screen — it can actually edit the file for you.
- In Google Docs, you could highlight a paragraph and ask Claude to shorten it, clean up the wording, or change the tone.
- Claude can write formulas, build pivot tables, create charts, clean up data, and add new tabs.
- You can just explain what you want and let Claude handle the setup.
- Anthropic is also giving you control over how freely Claude can make changes.
- When I fed it a complex Docker Compose file for my home lab that was throwing errors, Claude didn’t just suggest a fix.
