AI slop
Nobody can define AI slop. SlopFinder is averaging the answer
SlopFinder collects one-click anonymous votes on AI-generated text and exports the averaged results to Hugging Face. The project's premise is that 'slop' is measurable even if it is undefinable, which makes it easier to build on.
Emmanuel Fabrice Omgbwa Yasse AI-assisted
2026-08-18 · 5 min read

SlopFinder is an early MVP built on a proven idea pointed at a term nobody has defined. The idea is crowdsourced perception: show people something, collect one rating, move on. The target is AI slop, the label that dominates comment sections and platform policy talks until anyone tries to pin it down.
The mechanics are deliberately thin. SlopFinder pulls samples of AI-generated text from existing datasets, displays them anonymously, and asks each visitor one question: how sloppy is this text? No categories, no forms. When a sample accumulates enough votes, SlopFinder exports the results to Hugging Face as open data. The team calls it an early MVP, with a small dataset and a system still evolving.
What SlopFinder actually measures
The project does not define slop. By its own account, it measures the perceived, averaged version of what humans call slop. The dataset is a collection of working definitions, some close to objective, some pure taste. Pluralistic approaches like this have almost no foothold in production AI, according to an audit of pluralistic alignment.
That wager has a working precedent. LMSYS Chatbot Arena ranks models by blind human preference on an Elo scale, with nearly five million votes now packing the top labs into a 25-point band. Arena never defined "good" to rank models by it. SlopFinder bets the same logic holds on the negative end of quality.
Why slop resists a fixed definition
An early contributor's comment shows why the term will not hold still. For some readers slop just means low quality or low effort. For others it is attached to a specific failure mode: generated images where hands or eyes break the illusion, even in stylized outputs like charcoal or sketch drawings. The same commenter puts the usable-as-is rate for generated images at roughly one in forty.
Another comparison comes from gaming: storefronts full of what the commenter calls "baby's first FPS" projects, assembled from purchased source code, asset flips, hand-drawn textures and a weak map, then sold instead of used as a step toward something better. Broken hands are an objective defect. A lazy asset flip is a judgment about effort and intent. Enjoyability is a third axis: content that is enjoyable carries merit, while output that breaks basic logic "hurts in your brain." The line is blurry enough that Spiralism, a quasi-religious movement with about 10,000 cases in 2025, grew out of ordinary chatbot chats, per the Spiralism report. Any single definition would have to cover all three axes at once. That is why the project refuses to write one.
Trust erodes the same way on the enterprise side. Aleph Alpha warns that generic AI pilots undermine trust through a pattern of overpromise and underdeliver, a phenomenon it calls "enshittification." Upwork's research found a majority of workers report AI adds to their workload, as the productivity paradox writeup shows. Slop, in that reading, is the consumer-facing version of the same failure mode.
The risks of averaging vague votes
The cost of aggregating fuzzy judgments is real, and the project is candid about where it sits. A single fast rating mixes genuine defect detection with personal taste. With a small annotation pool, the average can swing on a handful of votes, and the export happens only after an unspecified number of ratings pile up.
Human judgment is the right reference point in principle. Oxford's GauntletBench put AI agents through 100 real tasks and the agents failed 81% of them, while expert human annotators completed the same work with over 80% accuracy, a level the paper calls "challenging yet feasible." But the word "expert" carries weight there. SlopFinder's pool is self-selected internet voters, not trained annotators, and the project does not pretend otherwise.
The main alternative, model self-review, has its own documented blindness. Research from Pine AI and the University of Washington found that a vision-language model judging AI-generated figures by screenshot is structurally blind: on 15 dense academic-figure renders, the VLM rated 14 as "perfect," while deterministic measurements of element geometry caught 40-74% of visual defects the screenshot metric saw nothing of. Machine judges cannot be trusted to recognize slop, which is the strongest argument for asking people at all. The same argument has pushed Microsoft to fund 18 university labs across six continents to independently red team AI systems, per Microsoft's safety-testing announcement.
| Judge | How it works | Known blind spot |
|---|---|---|
| SlopFinder vote pool | Anonymous one-click ratings, exported to Hugging Face | Small early pool; taste and defect detection mix |
| Chatbot Arena Elo | Blind human preference votes, nearly 5 million total | Scores perceived quality, not what causes it |
| VLM-on-screenshot judge | Another model reviews a rendered image | Structurally blind; called 14 of 15 figures "perfect" |
What an open slop dataset changes
Measurable, public slop ratings would become infrastructure for the labeling economy platforms are already building. Google ads carry a mandatory AI label, but enforcement outside Google's own inventory is largely honor-based, and platforms have also begun requiring AI labels on ads. Regulators are building their own enforcement machinery: France's CNIL is moving from guidance to enforcement with a slate of notes, audits and alliances, per the CNIL blueprint coverage. A shared dataset of what humans flag as slop gives moderation systems a common reference instead of letting every platform define machine-generated garbage in private.
One catch: users adapt. Research by Meena Jagadeesan and colleagues models how people respond when they know their work will be scanned for AI traces: they rewrite, trim and add idiosyncrasies to weaken whatever signal the detector picks up. The results run counter to what detector advocates expect. The same logic applies to a public slop dataset. The pool of votes that helps platforms moderate could also become the training signal for the next wave of generated content, engineered to dodge the exact judgments the dataset records.
If enough votes accumulate, SlopFinder's contribution is making the category legible. Slop will not get a definition. It may get a distribution, and that could turn out to be more useful.
- Source : Nobody can define AI slop. This open dataset is averaging the answer — 2026-08-08
Get the tech essentials in 3 minutes every morning
One email, every weekday, with what actually matters in AI and tech.