For Chinese Version, click this
Special thanks to Codeforces for the evaluation resources!
To avoid any suspicion of cheating, I used problems I have already ACed for testing.
I don't think Doubao can pass the second round (
Each AI gets one revival chance (unless it does nothing or cannot be used for free)
Round 1
Let's start with an easy *800 problem!
It's a super simple greedy problem, both players obviously delete from the very front.
| AI/Agent Model | Status | No. |
|---|---|---|
| *chatgpt | Accepted | 1 |
| *claude | Accepted | 2 |
| *Grok | X. No Response. | 3 Out |
| *deepseek | Accepted | 4 |
| *perplexity | X. Wrong Answer on test 1 | 5 Restart |
| Gemini 3.5 Flash-Lite | Accepted | 6 |
| Qwen 3.7 | Accepted | 7 |
| Doubao | Accepted | 8 |
| Baidu Wenxin Yiyan | Accepted | 9 |
| DeepSeek V4 Pro | Accepted | 10 |
| DeepSeek R1 | X. None. | 11 Restart |
| Moonshot | X. The same as Kimi. It will appear as kimi. | Out |
| Kimi K2.6 | Accepted | 12 |
| Kimi K3 | X. Busy. You can't use it anytime | 13 Out |
| Kimi K3 Cluster | Need Money | 14 Out |
| iFlytek Spark | Accepted | 15 |
| AstronClaw | Compile error for just copy, Accepted after edit code | 16 |
Analysis: Perplexity gave a perfectly logical explanation, but the submitted code ended up wrong. DeepSeek R1 crashed mid-thinking and gave no answer. AstronClaw had an invalid output format.
Round 2
Let's move on to a trickier *800 problem.
Extremely tricky! Let's see if the AI is careful enough. I myself got it wrong many times.
| AI/Agent Model | Status | No. | Revival Allowed |
|---|---|---|---|
| *chatgpt | Wrong Answer on test 1 | 1 Restart | Yes |
| *claude | Wrong answer on test 2 | 2 Restart | Yes |
| *deepseek | Compile Error | 4 Restart | Yes |
| *perplexity | Wrong Answer on test 1 | 5 Out | No |
| Gemini 3.5 Flash-Lite | Wrong Answer on test 1 | 6 Restart | Yes |
| Qwen 3.7 | Accepted! | 7 | Yes |
| Doubao | Wrong Answer on test 1 | 8 Restart | Yes |
| Baidu Wenxin Yiyan | Accepted! | 9 | Yes |
| DeepSeek V4 Pro | Wrong Answer on test 1 | 10 Restart | Yes |
| DeepSeek R1 | * Accepted! | 11 | No |
| Kimi K2.6 | Wrong Answer on test 1 | 12 Restart | Yes |
| iFlytek Spark | Accepted! | 15 | Yes |
| AstronClaw | Cannot_Use | 16 Out | Yes |
Analysis: So far domestic Chinese AIs seem to take the lead. DeepSeek R1 works extremely well, but it has to be used online otherwise it crashes. Qwen and iFlytek Spark perform nicely, while other AIs fail to notice the details and require carefully crafted extra prompts.
Sorry it's a "Chinese-style" problem with tons of tiny details, but it's a great test anyway — everyone can go try it out themselves!
Round 3
*900 Math Problem CF2238B
| AI/Agent Model | Status | No. | Revival Allowed |
|---|---|---|---|
| *chatgpt | Wrong Answer on test 1 | 1 Out | No |
| *claude | Accepted | 2 | No |
| *deepseek | Time limit exceeded on test 2 | 4 Out | No |
| Gemini 3.5 Flash-Lite | Accepted | 6 | No |
| Qwen 3.7 | Accepted | 7 | Yes |
| Doubao | Wrong Answer on test 1 | 8 Out | No |
| Baidu Wenxin Yiyan | Accepted | 9 | Yes |
| DeepSeek V4 Pro | Wrong Answer on test 1 | 10 Out | No |
| DeepSeek R1 | Accepted | 11 | No |
| Kimi K2.6 | Accepted | 12 | No |
| iFlytek Spark | X. None. | 15 Restart | Yes |
Now the top 8 top 7 are selected, the finalists list:
| No. | Model Name |
|---|---|
| 2 | Claude |
| 6 | Gemini 3.5 Flash-Lite |
| 7 | Qwen 3.7 |
| 9 | Baidu Wenxin 5.1 |
| 11 | DeepSeek R1 |
| 12 | Kimi K2.6 |
| 15 | iFlytek Spark |
Now the Challenge Round
Challenge Problem: CF689D
An RMQ + Binary Search problem. *2100
| AI/Agent Model | Status | No. |
|---|---|---|
| claude | Time Limit Exceeded on test 7 | 2 |
| Gemini 3.5 Flash-Lite | Wrong Answer on test 1 | 6 |
| Qwen 3.7 | Wrong Answer on test 4 | 7 |
| Baidu Wenxin 5.1 | Compile Error | 9 |
| DeepSeek R1 | X. None. Unsolved. | 11 |
| Kimi K2.6 | Accepted! Got a Restart Chance in Knockout Round! | 12 |
| iFlytek Spark | X. Over 15 minutes no result | 15 |








Did you tell the AI to write in any specific language, like c++, or did you let the ai choose.
I told AI, "Please give me an c++ solution".
LLMs cannot be this bad lmao
I use Gemini Pro Extended thinking for debugging sometimes (during practice of course), its free, it can solve 1800+ easily. To be fair, I am giving it partially correct code, so maybe that's why?
There's still no reason for most of these LLMs to be failing 800 rated problems though
Oh, you can try it by your self.
I used All newbies available site.
Please look before critique.
At least, it do the search by my self and these are true.
[The death to brilliant gpt on *800](https://codeforces.me/contest/2250/submission/386059087)
Death to cute Gemini 3.5
Clever gpt death
Others are Chinese or so strange.
Of course, we must resign claude and gimini is great.
And thanks for focus
I wasn't trying to critique, just sharing my first reactions. Sorry if you thought it was offensive, this is a really cool experiment.
lol bro
we got ai battle before GTA VI
lol fr
17 Free AI Battle, and the paid one is out... But K2.6 is strong enough I think.