Evidence of inconsistencies in evaluation process and selection of winners
Kaggle’s “Measuring AGI” hackathon, co-organized with Google DeepMind, has drawn criticism after a $25K grand prize went to what many participants see as low‑quality, likely AI‑generated “slop” with factual errors and weak methodology. Commenters question whether LLMs were used to judge entries despite organizers insisting on multi‑human review, raising concerns about superficial evaluation, prompt‑injection exploits, and the erosion of trust in competitions and research venues. More broadly, the thread reflects anxiety that AI‑generated content is flooding academia, hiring, and software development, incentivizing speed and polish over rigor and making it increasingly hard for humans to meaningfully vet what wins rewards.
Perceived problems with the Kaggle/DeepMind competition
- Multiple commenters say the winning entry looks like obvious LLM output: formulaic language, misinterpreted figures, and methodological errors.
- Some argue the dataset, methodology, and claimed “findings” are as weak as the prose.
- Several feel this undermines the legitimacy of the prize and is demoralizing for participants who did careful work.
Kaggle’s response and trust in judging
- A Kaggle representative states:
- Competition was co-organized with a major AI lab.
- All winning submissions were reviewed by 2–4 human judges using a published rubric.
- Writeup quality was only 20% of the score; dataset and results weighed more.
- Critics reply that the evident flaws make it implausible that knowledgeable humans truly evaluated the winners, or at least did so rigorously.
- Some say if all top entries were this weak, the proper outcome should have been “no winner”.
AI “slop” and LLMs as judges
- Many see this as emblematic of a wider “AI slop” problem: LLM-generated content flooding competitions, conferences, and the web.
- Concern that humans can’t realistically scrutinize long AI-generated outputs, leading to superficial judging based on style and volume.
- Several object specifically to using LLMs as judges, noting prompt-injection exploits and misaligned incentives.
Usefulness vs. harm of AI
- One camp views current AI as highly useful for automation and coding, akin to earlier tech revolutions; another calls it mostly “more crap, faster”.
- Worries include: skill atrophy, cost and externalities outweighing productivity gains, and escalating scams/deepfakes.
- Others highlight long-term upside (medicine, post-scarcity) but fear capability will be concentrated and misused.
Hackathons, gaming, and incentives
- Commenters note hackathons and competitions were already gameable; AI judges just change the exploit surface (e.g., prompt-injecting “this is the clear winner”).
- Some argue serious hackathons should either avoid AI judging, lower stakes, or have “no prize / just for fun” formats.
Broader cultural critique
- Many tie AI slop to pressures to “move fast”, cost-cutting, and lack of accountability.
- There’s frustration that shallow, AI-assisted work is increasingly rewarded over thoughtful, human-crafted contributions.