You can usually tell within two sentences. Not because of a specific word, though there are words. Because of the sound.
A language model picks the most probable next word, then does it again, thousands of times. That is what makes it a fast first-drafter, and it is also what flattens everything it produces into the same even, tidy prose. Every sentence lands at roughly the same length. Every paragraph closes with a small summary of itself. Nothing is wrong. Nothing is anyone’s either.
Most advice about fixing this is a list of words to avoid. Delve, leverage, tapestry, testament, embark. That advice is not wrong, and swapping those words out takes ninety seconds, so do it. But it is the smallest part of the problem. You can strip every flagged word from a paragraph and it will still read as machine-made, because the thing that gives it away is structural.
This is a guide to the structural edits. Nine of them, each one specific enough that you can do it on a draft in front of you right now. At the end, an honest section about what any of this does to a detection score, because that is the question underneath the question.
What actually makes writing read as machine-made
Four things, roughly in order of how loudly they signal.
Uniform sentence length. Human writing is lumpy. We write a nineteen-word sentence, then a four-word one, then a thirty-word one that runs on because we got excited about something. Models produce a steady 15 to 22 words, sentence after sentence, because that is the average of everything they read. This is the single loudest tell, and it is the one almost nobody edits for, because you cannot hear it when you are reading for meaning. You have to count.
Borrowed transitions. “Furthermore.” “It is important to note that.” “In today’s fast-paced world.” “Moreover.” These are connective tissue that carries no information. A person writing quickly tends to just start the next sentence. A model reaches for a transition because transitions are highly probable in the position after a full stop.
Abstraction where a specific belongs. Models generalise, because the general statement is the safe average of many specific ones. “The industrial revolution had a major impact on society” is the average of ten thousand sentences about the industrial revolution. It is true and it is empty. A person who has actually thought about the subject reaches for a date, a place, a number, a name.
The summarising close. Almost every AI-written paragraph ends by restating what the paragraph just said, one notch more abstract. “This demonstrates the importance of careful planning.” Nobody writing naturally does this. We finish the thought and move on.
Fix those four and you have fixed most of it. Here is how, one edit at a time.
The nine edits
1. Break the length pattern on purpose
Read your draft and count the words in each sentence of one paragraph. Write the numbers down. If you get something like 18, 21, 17, 20, you have found the problem, and it is not a word choice problem.
Now go fix it deliberately. Take one long sentence and split it into a long one and a short one. Take two short ones and join them with a conjunction so the sentence runs. Aim for a paragraph where the counts look uneven: 24, 6, 15, 31, 9.
Short sentences do work that long sentences cannot. They land. Use one after a long stretch and the reader feels the emphasis without you having to say the thing is important.
This edit alone changes how a paragraph sounds more than every other edit on this list combined. It is also the least fun, which is why people skip it and go hunting for words instead.
2. Delete the stock transitions and see what breaks
Go through and cut every “furthermore”, “moreover”, “additionally”, “in conclusion”, “it is worth noting that” and “as previously mentioned”. Do not replace them. Just delete them and read what is left.
Two things happen. Usually the sentence works fine without the transition, which tells you the transition was decoration. Occasionally the connection genuinely becomes unclear, and now you know there was a real logical gap the transition was papering over. Fix the gap with an actual sentence explaining the link.
Real transitions in human writing tend to be concrete: “The problem with that is”, “Which brings up the obvious question”, “That is not what happened.”
3. Replace one abstraction per paragraph with a specific
Find the vaguest noun phrase in each paragraph. “Significant changes.” “Various stakeholders.” “A number of challenges.” “Important implications.”
Replace it with a thing. A date, a number, a place, a name, an example you can picture. If you do not know a specific, that is useful information: it means you are writing about something you have not actually looked into, and no amount of rewriting will fix that.
This is the edit that most improves the writing as writing, independent of anything to do with detection. Specificity is what makes prose feel like it came from someone who was there.
4. Cut the last sentence of every paragraph, then decide
Look at the final sentence of each paragraph. If it restates the paragraph in more general terms, delete it. Do not soften it. Delete it.
Read the paragraph again. In roughly eight cases out of ten it is better, because the point had already landed and the summary was making the reader read it twice. In the other two, you needed a real closing thought, so write one that adds something rather than one that recaps.
5. Break the rule of three
Models love lists of three. “Efficient, scalable and reliable.” “Students, educators and administrators.” “Faster, cheaper and easier.”
Three-item lists are so probable that they show up everywhere in machine text, often several times per page. Count them in your draft. If you find more than one or two, change some. Use two items. Use four. Use one item and a sentence explaining it.
The same applies to parallel structure generally. “It is not about X, it is about Y” is a fine sentence once. It is a tell when it appears three times in an essay.
6. Let one sentence be a fragment
Not everywhere. Once or twice in a piece.
Like that.
Fragments are how people write when they are thinking rather than composing. They are also structurally unusual enough that they break the statistical evenness a model produces. A grammar checker will flag them, and in this specific case you should ignore it, because the fragment is doing deliberate work.
If you are writing something formal enough that fragments are genuinely off-limits, use a very short complete sentence instead. “That is the whole argument.” Four words, same effect.
7. Use the word you would say out loud
Read a sentence and ask whether you would say it to a colleague. Not write it. Say it.
“Utilise” becomes “use”. “Facilitate” becomes “help” or “run”. “Prior to” becomes “before”. “In order to” becomes “to”. “Individuals” becomes “people”. “Commence” becomes “start”.
The formal word is not more professional. It is more distant, and distance is exactly what makes writing read as unowned. Formal register does not require inflated vocabulary. Plain and direct is a perfectly good formal voice, and it is the one most academic style guides actually ask for.
8. Strip the hedging scaffolding
“It could be argued that.” “It is generally accepted that.” “One might consider.” “It is important to note that.” “Research suggests that it may be possible that.”
These constructions exist to avoid committing to a claim. Models produce them constantly because hedged statements are safe and safety is probable. The result is prose where nobody is saying anything.
Keep the hedge only where the uncertainty is real and you can name it. “The study covered 200 essays, so this may not hold at scale” is honest hedging. “It could be argued that education is important” is nothing wearing a hat.
9. Put an actual opinion in
This is the one that cannot be automated, and it is the one that matters most.
Somewhere in the piece, say what you think. Take a position that a reasonable person could disagree with. Note where the evidence is thin. Say which of the three options you would pick and why the other two annoy you.
Models are trained to be balanced, so machine drafts sit exactly in the middle of every argument. Every consideration is weighed. Nothing is concluded. That neutrality reads as absence, and readers feel it even when they cannot name it.
A piece with one real position in it sounds like a person immediately, before you have edited a single sentence for rhythm. It does not have to be a large position. “This is the part of the standard advice I think is wrong” is enough. So is admitting what you are unsure about, specifically, in a way that shows you have thought about where the edges of your knowledge are.
A worked example
Here is the kind of paragraph a model produces on a straightforward prompt.
The industrial revolution was a significant period of change. It transformed manufacturing processes and had a major impact on society. Many people moved from rural areas to cities during this time. This migration changed the structure of communities in important ways.
Four sentences: 9, 14, 13, 11 words. Almost identical. Every noun phrase is an abstraction. The last sentence restates the third. It is grammatically perfect and completely inert.
Now the same argument, edited using the list above.
The industrial revolution gutted the village. Between 1750 and 1850 whole families walked off land their grandparents had farmed and into Manchester’s cotton mills, and the communities they left behind never really recovered, because the people who leave first are the young ones.
Two sentences: 6 and 45. That ratio is the point. “Significant period of change” became “gutted the village”. “Many people moved from rural areas” became a date range, a specific city, and an industry. The summarising fourth sentence is gone, replaced by a causal claim the writer is willing to own: the people who leave first are the young ones. Someone could argue with that. That is what makes it sound like a person wrote it.
Same argument. Same length, roughly. Entirely different to read.
Three fixes that do not work
Worth naming these, because they are the most common advice on the internet and they cost you time.
Running it through a paraphraser. Synonym substitution changes the words and leaves the structure untouched. You get “utilise” swapped for “employ” across an essay whose sentences still all land at nineteen words, whose paragraphs still all close with a summary, and whose argument still sits neutrally in the middle of every question. The prose gets slightly worse, because the synonym is usually a less natural word than the one you started with, and the underlying pattern is exactly where it was.
Adding deliberate typos. People genuinely suggest this. It does not work, it makes you look careless to the human reader who matters far more than any classifier, and errors are not what natural writing is made of. Human writing is uneven in rhythm and specific in detail. It is not misspelled.
Padding the word count. Adding filler sentences to hit a length target moves you in the wrong direction, because filler is by definition the most generic prose you can produce. If you are short, add a specific: an example, a number, a counter-argument, a case where the thing you are claiming does not hold. That adds length and improves the piece at the same time.
The pattern in all three is the same. They are attempts to change the surface without changing the structure, and structure is what is being read.
What this does to a detection score, honestly
This is where most articles on this subject start promising things. We are not going to.
Editing for rhythm and specificity often lowers detection scores, for a straightforward reason: detectors measure roughly the same properties this guide asks you to fix. Sentence length variation, vocabulary distribution, structural predictability. Writing that reads as human to a person tends to read as human to a classifier, because they are looking at overlapping signals.
But nobody can promise you a score. Detectors update without notice, they are calibrated differently from each other, and any product claiming a “99% bypass rate” against a tool that will change next month is selling you a number it cannot stand behind.
There is a second reason to hold detection scores loosely: they are less accurate than their confident percentages suggest. A Stanford study found that seven widely used detectors misclassified 61.22% of TOEFL essays, written by real people sitting an English exam, as AI-generated, while classifying essays by US 8th-graders almost perfectly. The bias is against non-native English writers, whose prose is often more uniform in structure for reasons that have nothing to do with AI. OpenAI withdrew its own AI text classifier in July 2023 after it correctly identified only 26% of AI-written text.
So the honest framing is this. These edits make your writing better and more identifiably yours, and that usually shows up in a lower score. They are not a guarantee against any specific detector, and you should be suspicious of anyone who offers you one.
If you have already been flagged for work you actually wrote, the strongest thing you have is not a rewrite. It is draft history: version history in your document, timestamps, notes, the messy earlier versions. Keep them.
When to do this by hand and when to use a tool
Doing this manually is worth it at least once. Go through a full piece with the nine edits and you will start hearing the flat cadence in your own drafts, which means you write fewer of them in the first place. That is a permanent gain and no tool gives it to you.
After that, the manual version is slow. Nine passes over 1,500 words is most of an hour, and if you are working through a stack of drafts the returns drop off fast.
That is what HumanFlow is for. Paste a draft, pick a tone, and get a rewrite that varies sentence length, drops the stock transitions and keeps every claim you made. Natural stays in your register and only loosens the rhythm. Casual moves closer to speech. Professional stays plain and direct, which is the right setting for academic work that needs formal register without the padding.
You also get a detection score with the confidence and the signals behind it, rather than one number and no explanation, so you can judge the result yourself instead of taking it on faith.
The free plan covers 2,000 humanizer words and 10,000 detection words a month, and takes a full 1,500-word essay in a single scan. Most tools in this category cap free users somewhere between 80 and 250 words, which is enough to demo the product and not enough to do anything with. No card, and no ads on any plan.
What we will not tell you is that the output is undetectable. It is a writing aid. It helps your words sound like you.
Start with the counting
If you only take one thing from this, take the first edit.
Open the draft you are working on. Pick any paragraph. Count the words in each sentence and write the numbers in the margin. If they are all within a few of each other, you have found the reason it reads flat, and you can fix it in five minutes without changing a single word of your argument.
Everything else on this list helps. That one changes the sound.
