AmShZ's blog

By AmShZ, 89 minutes ago, In English

Hello everyone!

We recently held the first Math Contest on Repovive, and it was a really interesting experience. You could genuinely feel people's excitement about trying something new, and we also had many very strong participants from Japan. I really love the Japanese competitive programming community!

You can check out the contest here: https://repovive.com/contests/math-round-1

Also, Um_nik, Golovanov399, and maspy made videos about their participation in the contest, which was really cool to see!

For those who did not follow the original announcement, each written solution was evaluated by 5 GPT 6 judge agents. The judging was binary: your solution was either Accepted or Rejected. Rejected submissions had no penalty, but you had at most 15 submissions per problem.

Since the contest ended, we have received a lot of feedback. Based on what people told us, the judge was probably a little too strict about details. A solution with the correct main idea could still be rejected if an important step was not explained sufficiently. I still think having no penalty for rejected submissions was the right decision, because it would not be very fair to penalize someone simply because they did not know how much detail the judge expected.

I want to discuss two of the most important points from the feedback and hear what you think.

1. What should the verdict system look like?

If we keep binary judging, I think rejected submissions should continue to have no penalty. Participants do not initially know how detailed their solution needs to be, and writing and resubmitting a proof can already take a significant amount of contest time. More importantly, when a submission is rejected, we currently give no feedback about why it was rejected. So I do not think adding another penalty on top of that would improve the contest.

However, one of the most common pieces of feedback was: why use binary judging at all? Why not use something similar to IMO grading, such as a score from 0 to 7?

There are two main reasons we did not do this. First, creating a good partial scoring rubric is difficult. A mathematical problem may have many completely different solutions and many different partially correct approaches, and it is nearly impossible to predict all of them before the contest and decide exactly how many points every possible amount of progress should receive. In a traditional olympiad, graders can read the submitted solutions, discover new approaches, discuss them, and adapt the marking scheme. In a live contest, the score has to be produced immediately.

The second problem is that partial scores can leak information about the solution. Suppose you have a conjecture but are not sure whether it is correct. You submit an incomplete argument based on it and receive a high partial score. Now you have learned that your direction is probably correct. Even if the grading rubric is hidden, the score itself can act as a hint.

physics0523 suggested an interesting middle ground: instead of only Accepted and Rejected, we could distinguish between a solution being rejected because it is incorrect and being rejected because it is not detailed enough. This would effectively give us three verdicts. I like the idea, but even this can leak information. If you have an uncertain conjecture, you could submit an incomplete proof and see whether the judge calls it incorrect or merely incomplete. Maybe there are ways to limit this, but it is still something we have to consider.

So I am curious: what kind of verdict system would you prefer in a live math contest? Binary judging with no penalty? Partial scores? A third verdict for insufficient detail? Or something else?

2. How much detail should a solution require?

This might actually be the harder question. How detailed should a written solution be before we accept it? Even as a human grader, I am not sure I can describe the exact boundary perfectly.

Obviously, participants should not need to explain every trivial algebraic manipulation or completely standard step. On the other hand, simply stating the main observation without explaining why it proves the result should probably not be enough. The difficult part is deciding where exactly the boundary lies. When is a missing step routine enough that a grader can safely infer it, and when is it an essential part of the proof that must be written explicitly?

This becomes especially important in a live contest. If the judge is too strict, participants may spend too much of the contest writing formal details instead of solving problems. If it is too lenient, incomplete arguments may get accepted.

So this is the second question I would really like to discuss: if you were grading these solutions, how would you define the minimum level of detail required for an Accepted verdict?

This was our first attempt at this format, so we definitely expect the judging system and rules to evolve. The experiment was really interesting for us, and the feedback we received has already been extremely useful.

I would love to hear your thoughts, especially from people who participated in the round.

  • Vote: I like it
  • +7
  • Vote: I do not like it