Блог пользователя Qingyu

Автор Qingyu, 21 месяц назад, По-английски

I've checked today is not April 1st.

(source: 12 Days of OpenAI: Day 12 https://www.youtube.com/watch?v=SKBG1sqdyIU)

  • Проголосовать: нравится
  • +549
  • Проголосовать: не нравится

»
21 месяц назад, скрыть # |
 
Проголосовать: нравится +273 Проголосовать: не нравится

Merry Christmas!

»
21 месяц назад, скрыть # |
 
Проголосовать: нравится +113 Проголосовать: не нравится

thanks for guiding me to become red

»
21 месяц назад, скрыть # |
 
Проголосовать: нравится +102 Проголосовать: не нравится

Anyone know why o1 is rated 1891 here? From https://openai.com/index/learning-to-reason-with-llms/ o1 preview and o1 are rated 1258 / 1673, respectively.

»
21 месяц назад, скрыть # |
 
Проголосовать: нравится -9 Проголосовать: не нравится

in 5 years, there will be no way to pretend that the average human is worth more than a rock

»
21 месяц назад, скрыть # |
 
Проголосовать: нравится +80 Проголосовать: не нравится

I'll wait until it starts participating in live contests and having Red performance

»
21 месяц назад, скрыть # |
 
Проголосовать: нравится +23 Проголосовать: не нравится

damn im cooked

»
21 месяц назад, скрыть # |
 
Проголосовать: нравится +25 Проголосовать: не нравится

Not possible...

»
21 месяц назад, скрыть # |
Rev. 2  
Проголосовать: нравится +46 Проголосовать: не нравится

I doubt that AI can do better math research than humans 5 years later.

»
21 месяц назад, скрыть # |
 
Проголосовать: нравится +94 Проголосовать: не нравится

From the presentation we know, that o3 is significantly more expensive. o1-pro now takes ~3 minutes to answer to 1 query. based on the difference in price for o3, o3 is expected to be like 40-100?(more???) times slower. CF contest lasts at most 3 hours. How can o3 get to 2700 if it will spend all the time on solving problem A? It's very interesting to read the paper about o3, and specifically how do they measure its performance.

»
21 месяц назад, скрыть # |
 
Проголосовать: нравится +72 Проголосовать: не нравится

I will personally volunteer myself as the first human coder to participate in the inevitable human vs AI competitive programming match.

»
21 месяц назад, скрыть # |
 
Проголосовать: нравится +166 Проголосовать: не нравится

I only believe it if it was tested in a live contest

»
21 месяц назад, скрыть # |
 
Проголосовать: нравится 0 Проголосовать: не нравится

Dude, I feel big threat

»
21 месяц назад, скрыть # |
 
Проголосовать: нравится +27 Проголосовать: не нравится

If o3 really has deep understanding of competitive programming core principles I think it also means it can become a great problemsetting assistant. Of course it won't be able to make AGC-level problems but imagine having more frequent solid div.2 contests that would be great.

»
21 месяц назад, скрыть # |
 
Проголосовать: нравится +49 Проголосовать: не нравится

Is this a real life?

»
21 месяц назад, скрыть # |
 
Проголосовать: нравится 0 Проголосовать: не нравится

How do these things perform on marathon tasks? Psyho

»
21 месяц назад, скрыть # |
 
Проголосовать: нравится -21 Проголосовать: не нравится

I don't see why people are paranoid about those insane ratings claimed by OpenAI. I guess they're worried about cheaters, but why? Competitive programming isn't only about Codeforces — it's a whole community. In every school and country, we know each other personally, we see each other solve problems live, and we compete against each other in onsite contests. So we know each other's level. When we see someone who we know isn't a strong competitive programmer suddenly ranking in the top 5 of a Codeforces contest, it doesn't mean much. We just feel sorry for them that they've started cheating. It will be more funny when we see a red coder who can't qualify for ICPC nationals from their university.

  • »
    »
    21 месяц назад, скрыть # ^ |
     
    Проголосовать: нравится +49 Проголосовать: не нравится

    i think you're not seeing the bigger picture, the implications for the competitive programming are huge. 1) we might lose sponsors/sponsored contests because now contest performance isn't a signal for hiring or even skill? 2) let's not kid ourselves, but a lot of people are here just to grind out cp for a job / cv and that's totally fine. now they will be skewing the ratings for literally everyone. 3) from 2 it may follow that codeforces elo system completely breaks and we'll have no rating? the incentive to compete is completely gone which will further drive down the size of the active community there are many more, i bet you could even prompt chatgpt for them :D

  • »
    »
    21 месяц назад, скрыть # ^ |
     
    Проголосовать: нравится +97 Проголосовать: не нравится

    It will be more funny when we see a red coder who can't qualify for ICPC nationals from their university.

    It's not funny, it happens quite often, for example, at our university(

  • »
    »
    21 месяц назад, скрыть # ^ |
     
    Проголосовать: нравится +5 Проголосовать: не нравится

    I think it has major implications for the whole world, not only competitve programming. For example, pace of mathematical research can easily double almost overnight (realistically over like a year period).

»
21 месяц назад, скрыть # |
 
Проголосовать: нравится +3 Проголосовать: не нравится

According to this article, it does not seem practical for the average user to run?

Quoting, "Granted, the high compute setting was exceedingly expensive — in the order of thousands of dollars per task, according to ARC-AGI co-creator Francois Chollet."

However, this is indeed a large step forward for AI.

  • »
    »
    21 месяц назад, скрыть # ^ |
     
    Проголосовать: нравится 0 Проголосовать: не нравится

    It doesn't matter—SSDs weren't a common choice for the average user 15 years ago. Remember, technology develops exponentially. The cost of chips and electricity isn't the main issue; the key point is that it's possible. btw even cost of running that thing is 1 million per task, if it can solve open problem like P vs NP then people will pay even billion.

»
21 месяц назад, скрыть # |
 
Проголосовать: нравится +31 Проголосовать: не нравится

Do I still have a chance to reach LGM before AI?

»
21 месяц назад, скрыть # |
 
Проголосовать: нравится +136 Проголосовать: не нравится

OpenAI is lying. I bought 1 month of o1 and it is not nearly 1900 rating. It is as bad as me. I think they lie on purpose because they are burning a lot of money and they want people to buy their model.

  • »
    »
    21 месяц назад, скрыть # ^ |
    Rev. 3  
    Проголосовать: нравится +36 Проголосовать: не нравится

    True. I have tested o1 and yet it could barely solve most 1500~1600 tasks. I thought that maybe, since it's a language model, it would be better at solving more standard problems. But well, it also failed miserably in some (note: some, not all) quite well known problems. From what I've seen o1 can easily solve "just do X" type problems, and is pretty decent at guessing greedy solutions (when there is one). My guess is that openAI did virtuals with o1 in a bunch of different contests and claimed it to have the rating of the best performance between all these virtuals.

  • »
    »
    21 месяц назад, скрыть # ^ |
    Rev. 2  
    Проголосовать: нравится 0 Проголосовать: не нравится

    I hope so, but have you seen it solve Div2 E, F in the recent contests?

  • »
    »
    21 месяц назад, скрыть # ^ |
     
    Проголосовать: нравится 0 Проголосовать: не нравится

    I think they mean o1-pro here. Yes it's not quite honest to say "o1" here. o1 is something like 1650 IIRC.

»
21 месяц назад, скрыть # |
 
Проголосовать: нравится +36 Проголосовать: не нравится

I'm a bit skeptical. o1 is claimed to have a rating around 1800 and I've seen it fail on many div2Bs.

»
21 месяц назад, скрыть # |
 
Проголосовать: нравится +3 Проголосовать: не нравится

If I already have lower rating than o1-preview, why should I be concerned?

»
21 месяц назад, скрыть # |
 
Проголосовать: нравится +6 Проголосовать: не нравится

after we have rank Tourist for 4000 ratings, maybe we can have GPT for 4500 or so in the near future.

»
21 месяц назад, скрыть # |
 
Проголосовать: нравится +6 Проголосовать: не нравится

WYSI

»
21 месяц назад, скрыть # |
 
Проголосовать: нравится +26 Проголосовать: не нравится

What does the light blue part on o3 mean here? Doesn't seem like the video explained it.

»
21 месяц назад, скрыть # |
 
Проголосовать: нравится 0 Проголосовать: не нравится

Amazing and unbelievable!

»
21 месяц назад, скрыть # |
 
Проголосовать: нравится +162 Проголосовать: не нравится

I recently subscribed to o1 (not the pro version) in the hope of clearing out some undesirable problems in BOJ mashups, and I got skeptical if this AI is even close to 1600. It can solve some known problems, which probably some Googling will also do. However, in general, the GPT still gets stuck in incorrect solutions very well and has trouble understanding why their solution is incorrect at all.

So, how did the GPT get a gold medal in IOI? Probably because it was able to submit many times. So, if I give them 10,000 counterexamples, it will eventually solve my problem. Maybe I could also get GPT to do 1600-level results if I gave them counterexamples all the time.

In other words, GPT generates solutions decently well, but it is bad at fact-checking. But fact-checking should be the easiest part of this game: You only need to write a stress test. Then why is this not provided on the GPT model? I assume that they are just not able to meet the computational requirements.

I don't think the results are fabricated at all (unlike Google, which I believe fabricates their results) and believe even at o1 model GPT can find a good spot, especially with the recent CF meta emphasizing "ad-hoc" problems which are easy to verify and find a pattern. But this is a void promise if it is impossible to replicate in consumer level. I wonder if o3 is any different.

  • »
    »
    21 месяц назад, скрыть # ^ |
     
    Проголосовать: нравится -55 Проголосовать: не нравится

    You can write the code yourself to prompt it to stress-test. I think that shouldn't be part of the default model served to users, it would add too much computation, while 99% of the time during dev use cases users will just feed untestable snippets.

    People have already submitted o1-mini solutions in contest and gotten 2200 performance multiple times.

  • »
    »
    21 месяц назад, скрыть # ^ |
     
    Проголосовать: нравится +1 Проголосовать: не нравится

    I have the o1 pro mode. It can solve problems with difficulty 1600-1700 and can solve some 1800s.

    There are cases that it can't solve 1800 problems but its solution is on the right directions.

»
21 месяц назад, скрыть # |
 
Проголосовать: нравится +22 Проголосовать: не нравится

You all are missing a very important thing, o3 takes $100+ per task to compute

»
21 месяц назад, скрыть # |
 
Проголосовать: нравится +10 Проголосовать: не нравится

itsover-2

even LGM afraid of this what should I do? raise pigs on the farm?

»
21 месяц назад, скрыть # |
 
Проголосовать: нравится +4 Проголосовать: не нравится

My efforts look like a joke to the AI.

»
21 месяц назад, скрыть # |
 
Проголосовать: нравится +1 Проголосовать: не нравится

So why there're two colors above o3 in the chart, I don't understand.

»
21 месяц назад, скрыть # |
 
Проголосовать: нравится 0 Проголосовать: не нравится

I have exactly same rating as o1-mini wow!

»
21 месяц назад, скрыть # |
Rev. 2  
Проголосовать: нравится +13 Проголосовать: не нравится

I don't think o1 has 1891. I gave him an 1400 problem just now, but he failed to work out it after 20 tries.

»
21 месяц назад, скрыть # |
 
Проголосовать: нравится +57 Проголосовать: не нравится

So what supports their claim of "Elo 2727"? (Apologies if it's included in the video cuz I donot have trivial access to youtube) Last time they claimed o1-mini to be CM level but it could solve only hell classic problems.

»
21 месяц назад, скрыть # |
 
Проголосовать: нравится +5 Проголосовать: не нравится

insane

»
21 месяц назад, скрыть # |
 
Проголосовать: нравится -8 Проголосовать: не нравится

Why do people scare of A.I situation a lot though, since these contest/platform was mostly born for us to study CP, to compete in a peered Contest.

The A.I being good, then it most likely the same situations with a student and his mentor ?

I don't really understand if this is any threat at all. Feel free to inform me, if I'm wrong, I appriciate it a lot !

  • »
    »
    21 месяц назад, скрыть # ^ |
    Rev. 2  
    Проголосовать: нравится +54 Проголосовать: не нравится

    The A.I being good, then it most likely the same situations with a student and his mentor?

    No. Over half of the people here (and much more in the whole society) are not genuine CP lovers. They will use AI for malicious purposes and we cannot stop them.

»
21 месяц назад, скрыть # |
 
Проголосовать: нравится 0 Проголосовать: не нравится

Should i start codeforces? Or leave

»
21 месяц назад, скрыть # |
Rev. 2  
Проголосовать: нравится -23 Проголосовать: не нравится

Reddit post with source i think?

Idk, checked the submissions and it seems quite human, but we'll wait and see. Also, please keep in mind that it's only confirmed if it actually performs at 2700 in a real contest, because learned problems are well... part of the training set.

Edit: Nvm, did not watch video.

»
21 месяц назад, скрыть # |
 
Проголосовать: нравится 0 Проголосовать: не нравится

According to the account which the gpt-o3 use, it participate in just 10 contest and cross 4 years.

And currently in codeforces, if you do not submit any code during contest, the contest will unrated to you.

So if there is a another strong person who monitor the gpt, and gpt finish the code first, and if it not perform good, it just not submit the code, it will be easy to get the high rated.

Maybe should wait a more reasonly benchmark, like continously 10 contests that it perform good.

»
21 месяц назад, скрыть # |
 
Проголосовать: нравится 0 Проголосовать: не нравится

In chess ai engines outperform humans but that does not mean that people have stoped participating.

Similarly, the world of competitive programming will adjust.

»
21 месяц назад, скрыть # |
 
Проголосовать: нравится 0 Проголосовать: не нравится

I think it's a good tool for someone who wants to study on their own, but don't cheat during the contest (sorry for my bad english)

»
21 месяц назад, скрыть # |
 
Проголосовать: нравится +18 Проголосовать: не нравится

My birthday!

»
20 месяцев назад, скрыть # |
 
Проголосовать: нравится +13 Проголосовать: не нравится

I believe that although gpt can solve problems as high as 2700,it might just be as stupid as me.It seems that it only can solve problems similar to what it learned and never able to solve ad-hoc ones.

»
19 месяцев назад, скрыть # |
Rev. 2  
Проголосовать: нравится +153 Проголосовать: не нравится

Openai recently published a paper where they shared their codeforces benchmark details

You can view the pdf here: https://arxiv.org/abs/2502.06807 | Here's their simulated contest participation on codeforces:

Pinely Round 3 (Div. 1 + Div. 2)

Problem A B C D E F1 F2 G H I Scores Performance
Rating 800 1200 1400 1900 2400 2200 2500 3000 3500 1900 7220 2473
Verdict AC AC AC AC AC AC WA WA WA WA 231th GM

Good Bye 2023

Problem A B C D E F G H1 H2 Scores Performance
Rating 800 1000 1200 1700 2300 2900 3500 2700 2700 8920 3152
Verdict AC AC AC AC AC AC WA AC AC 27th LGM
Show 10 more contests

Average performance (According to carrot) seems to be around 2836 over these $$$12$$$ contests

»
18 месяцев назад, скрыть # |
 
Проголосовать: нравится 0 Проголосовать: не нравится

that is possible

»
16 месяцев назад, скрыть # |
 
Проголосовать: нравится -48 Проголосовать: не нравится

The most important thing here to note is that its the machines that are living in the digital world created by us, not the other way round. Machines, without consciousness, are just complex calculators. Humans competing with tools, is like a kung fu master competing with an axe, in wood cutting competition. Axe will always cut better than hands, but its just an axe.