Блог пользователя zap2727

Автор zap2727, история, 2 месяца назад, По-английски

For Chinese Version, click this

Special thanks to Codeforces for the evaluation resources!

To avoid any suspicion of cheating, I used problems I have already ACed for testing.

I don't think Doubao can pass the second round (

Each AI gets one revival chance (unless it does nothing or cannot be used for free)

Round 1

Let's start with an easy *800 problem!

CF2248A

It's a super simple greedy problem, both players obviously delete from the very front.

AI/Agent Model Status No.
*chatgpt Accepted 1
*claude Accepted 2
*Grok X. No Response. 3 ‌Out
*deepseek Accepted 4
*perplexity X. Wrong Answer on test 1 5 ‌Restart
Gemini 3.5 Flash-Lite Accepted 6
Qwen 3.7 Accepted 7
Doubao Accepted 8
Baidu Wenxin Yiyan Accepted 9
DeepSeek V4 Pro Accepted 10
DeepSeek R1 X. None. 11 ‌Restart
Moonshot X. The same as Kimi. It will appear as kimi. ‌Out
Kimi K2.6 Accepted 12
Kimi K3 X. Busy. You can't use it anytime 13 ‌Out
Kimi K3 Cluster Need Money 14 ‌Out
iFlytek Spark Accepted 15
AstronClaw Compile error for just copy, Accepted after edit code 16

Analysis: Perplexity gave a perfectly logical explanation, but the submitted code ended up wrong. DeepSeek R1 crashed mid-thinking and gave no answer. AstronClaw had an invalid output format.

Round 2

Let's move on to a trickier *800 problem.

CF2250A

Extremely tricky! Let's see if the AI is careful enough. I myself got it wrong many times.

AI/Agent Model Status No. Revival Allowed
*chatgpt Wrong Answer on test 1 1 ‌Restart Yes
*claude Wrong answer on test 2 2 ‌Restart Yes
*deepseek Compile Error 4 ‌Restart Yes
*perplexity Wrong Answer on test 1 5 ‌Out No
Gemini 3.5 Flash-Lite Wrong Answer on test 1 6 ‌Restart Yes
Qwen 3.7 Accepted! 7 Yes
Doubao Wrong Answer on test 1 8 ‌Restart Yes
Baidu Wenxin Yiyan Accepted! 9 Yes
DeepSeek V4 Pro Wrong Answer on test 1 10 ‌Restart Yes
DeepSeek R1 * Accepted! 11 No
Kimi K2.6 Wrong Answer on test 1 12 ‌Restart Yes
iFlytek Spark Accepted! 15 Yes
AstronClaw Cannot_Use 16 ‌Out Yes

Analysis: So far domestic Chinese AIs seem to take the lead. DeepSeek R1 works extremely well, but it has to be used online otherwise it crashes. Qwen and iFlytek Spark perform nicely, while other AIs fail to notice the details and require carefully crafted extra prompts.

Sorry it's a "Chinese-style" problem with tons of tiny details, but it's a great test anyway — everyone can go try it out themselves!

Round 3

*900 Math Problem CF2238B

AI/Agent Model Status No. Revival Allowed
*chatgpt Wrong Answer on test 1 1 ‌Out No
*claude Accepted 2 No
*deepseek Time limit exceeded on test 2 4 ‌Out No
Gemini 3.5 Flash-Lite Accepted 6 No
Qwen 3.7 Accepted 7 Yes
Doubao Wrong Answer on test 1 8 ‌Out No
Baidu Wenxin Yiyan Accepted 9 Yes
DeepSeek V4 Pro Wrong Answer on test 1 10 ‌Out No
DeepSeek R1 Accepted 11 No
Kimi K2.6 Accepted 12 No
iFlytek Spark X. None. 15 ‌Restart Yes

Now the top 8 top 7 are selected, the finalists list:

No. Model Name
2 Claude
6 Gemini 3.5 Flash-Lite
7 Qwen 3.7
9 Baidu Wenxin 5.1
11 DeepSeek R1
12 Kimi K2.6
15 iFlytek Spark

Now the Challenge Round

Challenge Problem: CF689D

An RMQ + Binary Search problem. *2100

AI/Agent Model Status No.
claude Time Limit Exceeded on test 7 2
Gemini 3.5 Flash-Lite Wrong Answer on test 1 6
Qwen 3.7 Wrong Answer on test 4 7
Baidu Wenxin 5.1 Compile Error 9
DeepSeek R1 X. None. Unsolved. 11
Kimi K2.6 Accepted! Got a Restart Chance in Knockout Round! 12
iFlytek Spark X. Over 15 minutes no result 15
  • Проголосовать: нравится
  • -15
  • Проголосовать: не нравится

»
2 месяца назад, скрыть # |
 
Проголосовать: нравится +1 Проголосовать: не нравится

Did you tell the AI to write in any specific language, like c++, or did you let the ai choose.

»
2 месяца назад, скрыть # |
 
Проголосовать: нравится 0 Проголосовать: не нравится

LLMs cannot be this bad lmao

I use Gemini Pro Extended thinking for debugging sometimes (during practice of course), its free, it can solve 1800+ easily. To be fair, I am giving it partially correct code, so maybe that's why?

There's still no reason for most of these LLMs to be failing 800 rated problems though

»
2 месяца назад, скрыть # |
 
Проголосовать: нравится 0 Проголосовать: не нравится

we got ai battle before GTA VI

»
2 месяца назад, скрыть # |
 
Проголосовать: нравится 0 Проголосовать: не нравится

17 Free AI Battle, and the paid one is out... But K2.6 is strong enough I think.