De-slopping in practice
Applying slop research to human-AI collaborative writing — what we found and what we built
February 2026
You know AI-generated text when you see it. You might not be able to say why. Something about the rhythm, the word choices, the way every paragraph lands with the same weight. It sounds competent and says nothing. The literary equivalent of a stock photo.
Turns out this is statistically measurable.
In October 2025, a research team published the Antislop paper, a framework for identifying and eliminating overused patterns in language model output. They analyzed creative writing from 67 models against human baselines and found that certain words and phrases appear over a thousand times more frequently in LLM text than in human writing. Not just a little more. Orders of magnitude more. The word "flickered" shows up in the overuse list of 98.5% of models tested. The phrase "voice barely above a whisper" appears in 68.7%.
They call it slop. And they built tools to suppress it.
Existing text, no GPU
The Antislop paper's tools — the Antislop Sampler and FTPO training method — work at inference time. They intervene during text generation, detecting overused patterns as they form and steering the model toward alternatives. Effective, but they require GPU infrastructure and direct access to model weights or logprobs.
That doesn't help with text that already exists. Articles written with AI assistance. Documentation co-authored with Claude or GPT. The words are already on the page.
This site has twenty-odd articles, most written with AI collaboration. The recursive mirror experiment was explicitly about this — feeding journal entries to Claude, getting a voice analysis back, using that analysis to guide the writing of the article about the analysis. The question that exercise left open: how do you know when the AI's patterns have replaced yours?
The Antislop paper gives us data to answer that question. Not with vibes. With frequency ratios.
Twelve edits in one article
We built a styleguide from two sources: the paper's empirical banlist data and the voice patterns from our existing voice analysis. One tells you what to scan for. The other tells you what to aim for. Then we ran the rules against the articles on this site.
The word-level banlists were clean. No "flickered", no "murmured", no "gaze" or "shimmered". Technical writing doesn't trigger the same slop as creative fiction — the paper's most overrepresented patterns cluster around literary description and character action, which isn't the register these articles are written in.
The structural patterns were more interesting. In the convergent evolution article, we found:
- "Same instinct" repeated three times across the parallel comparison sections — a structural tic where the model's go-to phrasing for drawing parallels becomes a pattern itself
- "Neither is/approach" used three times to hedge comparisons — a diplomatic construction that reads as AI equivocating rather than a human making a judgment
- "Worth noting" and "What's interesting is" — throat-clearing before observations that would be stronger without the preamble
- "Landscape" — a single word from the paper's overuse list, buried in a sentence about document scope
- "This isn't X — it's Y" — the paper's most overrepresented sentence construction (6.3x human rates), found twice in the article. One was a deliberate philosophical claim that earned the structure. The other was a defensive disclaimer that read better without it.
Twelve edits total. The article was already strong — specific timestamps, honest self-critique, no hedging on the parts that matter. The slop was in the connective tissue, not the substance. Transitions and diplomatic framings where the model was smoothing edges that didn't need smoothing.
SlopSquid
Doing this manually works for one article. It doesn't scale to a site with twenty. And it doesn't generalize to the next article.
SlopSquid started as an abandoned browser extension prototype — vague ambitions about "detecting AI artifacts" without the data to back it up. The Antislop paper provides exactly that data. Word-level frequency ratios across 67 models. Trigram overuse percentages. Regex patterns for structural constructions like "It's not X, it's Y" (6.3x more prevalent in LLM output than human writing).
The revamped version is a Go CLI that ships with the paper's banlist data baked in. No model dependencies. No GPU. Just static analysis with quantitative scoring. You point it at a file or directory, it tells you what it found and how sloppy the text is by the numbers.
$ slopsquid score src/ * 20.6 moderate 57 hits 1186 words styleguide.html . 5.7 clean 11 hits 1228 words antislop-editorial.html . 4.6 clean 2 hits 214 words economic-honesty.html . 3.2 clean 3 hits 708 words it-smash/addendum.html . 0.5 clean 1 hits 3957 words convergent-evolution-personal-ai.html . 0.0 clean 0 hits 582 words recursive-mirror.html
The styleguide scores "moderate" because it literally lists slop words as examples. This article scores low because its mentions are in quotation context. Everything else reads clean. The convergent-evolution article — 3,957 words of AI-collaborative writing — lands at 0.5/100 with a single hit.
Each hit is weighted by the paper's frequency ratio. A word that's 85,000x overrepresented scores higher than one at 60x. The aggregate gives you a 0-100 slop density score: 0-20 is clean, 20-50 is moderate, 50+ means the model was doing most of the writing.
Hard bans break writing
The Antislop paper itself flags this problem. Hard bans on vocabulary cause worse output degradation than the slop they prevent. Token banning — the naive approach of just blocking words — crashed writing quality scores from 68 to 28 in their tests. The model collapses into repetition and incoherence when you take away too many of its preferred paths.
Their solution: soft bans. Reduce the probability of overused patterns without eliminating them entirely. If "flickered" is genuinely the right word in context, it should still be available. The system suppresses the default, not the deliberate choice.
Same principle for post-hoc editing. Over-policing kills voice. If you hunt every word on a banlist, you'll end up with text that's technically clean and reads like it was written by committee. The goal is awareness, not avoidance. Notice when the model is on autopilot. Decide deliberately whether the phrasing is yours or a default.
This connects to something the recursive mirror experiment revealed: the AI captures your refined voice, not your raw one. The aspirational version. What emerges when writing for an audience. Slop detection works the other way — it catches the moments where the AI's default voice replaced yours. Between the two, you get a clearer picture of where the collaboration stands.
Beyond fiction banlists
The paper focused on creative fiction. Their banlists cluster around literary description — character names, sensory phrases, atmospheric constructions. Technical and analytical writing has different slop patterns: hedging phrases, diplomatic both-sides framings, throat-clearing transitions. The framework generalizes. The specific banlists don't.
SlopSquid v1 ships with the paper's data. Future versions should build domain-specific profiles. What does slop look like in documentation? In blog posts? In commit messages? The methodology for finding out is the same: compare model output against human baselines, measure the frequency ratios, build the banlist.
For now, we have a number for how sloppy our text is. The articles on this site score low — against fiction banlists. That's the wrong test for technical writing. Nobody's "flickering" in a blog post about SQLite. The question is whether a detector tuned for analytical prose — hedging phrases, diplomatic framings, the word "ecosystem" — finds the same clean result.
So we built one.