GLM-5.2 is live — Z.AI's flagship with a 1M-token lossless contextfrom $0.900 per 1M tokens
Can You Tell If a Writer Used ChatGPT? Patterns, Not Proof
2026/08/02

Can You Tell If a Writer Used ChatGPT? Patterns, Not Proof

Can you tell if a writer used ChatGPT? Learn the recurring style signals, why no phrase proves AI authorship, and a fairer way to review suspicious text.

You probably cannot prove that a writer used ChatGPT by circling one sentence. You can, however, notice when an article repeatedly chooses the safest possible opening, the neatest possible transition, and the vaguest possible conclusion.

That distinction matters. A cluster of predictable writing habits may justify a closer read. It does not establish who—or what—wrote the page. Even purpose-built classifiers can produce false positives, and ordinary readers are less consistent than classifiers. The better question is not “Which forbidden phrase gives ChatGPT away?” It is “Does this article contain enough specific thought, evidence, and editorial judgment to earn trust?”

This guide answers that question without pretending style is forensic evidence.

TL;DR

  • No single word, transition, metaphor, or punctuation mark proves ChatGPT use. Human writers share every commonly cited “AI phrase.”
  • Suspicion usually comes from a repeated combination: interchangeable opening, over-signposting, symmetrical lists, vague authority, inflated significance, and a conclusion that restates rather than decides.
  • A polished article can be human, AI-assisted, heavily edited, or fully generated. The finished text rarely reveals a clean binary history.
  • AI detectors are useful as screening signals, not verdicts. OpenAI withdrew its own classifier in 2023 because of low accuracy, and published research has found false-positive risks for non-native English writers.[2][3]
  • For editorial review, test claims, sources, specificity, provenance, and revision history before arguing over “AI-sounding” vocabulary.

Why readers think they can spot ChatGPT writing

Once a verbal habit has been labeled “AI,” it becomes difficult to stop seeing it. A transition such as “However, it is important to note” feels diagnostic because it is common, formal, and easy to remember. The same happens with em dashes, three-part lists, rhetorical questions, and endings about “the future.”

The trouble is base rates. Those devices were common in corporate copy, student essays, journalism, and search-optimized articles long before ChatGPT. Language models learned them because people wrote them. Finding one in a new article is therefore compatible with several explanations:

  • the writer used ChatGPT and pasted the result;
  • the writer used an AI draft, then revised it;
  • the writer used AI only for grammar or structure;
  • the writer follows a conventional editorial template;
  • the writer simply likes that phrase.

Matt Lillywhite's essay about commenters acting as “AI detectives” captures the social version of the problem: once someone learns a few supposed tells, almost any polished sentence can be made to look suspicious.[1] The observation is useful. The certainty is not.

Seven patterns that make writing feel AI-generated

These are editorial symptoms, not proof of authorship. One occurrence means little. Repetition across the whole page is what makes the prose feel manufactured.

1. The opening could introduce almost any topic

An interchangeable introduction announces that a subject is changing the world, asks whether the reader has ever wondered about it, then promises a comprehensive exploration. Replace the topic with cloud computing, espresso, or retirement planning and the paragraph still works.

A stronger opening spends its first lines on a claim the article can actually defend:

Interchangeable openingSpecific opening
“In today's rapidly evolving digital landscape, AI is transforming how we write.”“A reader cannot prove ChatGPT use from one phrase; the useful evidence is a repeated lack of specificity.”
Announces a broad trendMakes a contestable claim
Delays the answerStarts with the answer

2. Every paragraph arrives with a signpost

Transitions are helpful until they begin narrating the existence of the article: “First and foremost,” “It is also worth noting,” “On the other hand,” “Ultimately,” “In conclusion.” The reader can feel the outline under every paragraph.

Human editing often removes half of these bridges. When the next sentence logically follows, it does not need a crossing guard.

3. The rhythm never changes

Generated prose often settles into similarly sized paragraphs and similarly balanced sentences. Each section explains a point, adds a caveat, and closes with a tidy implication. Nothing is technically wrong. Nothing surprises the ear.

Good prose is not random, but it has pressure. A short sentence can stop an argument. A longer one can carry evidence and qualification together. A fragment, used deliberately, can change pace. Uniform rhythm makes even accurate information feel preassembled.

4. Lists are suspiciously complete

Three benefits. Five challenges. Seven best practices. The number is chosen before the reporting is done, so every item receives equal visual weight even when only two matter.

This is not an argument against lists. It is an argument for hierarchy. If one finding changes the decision and four are minor, the article should say so instead of presenting five democratic bullets.

5. Authority is invoked but never located

Phrases such as “experts agree,” “studies show,” and “research suggests” create the shape of evidence without giving the reader anything to inspect. This problem is more serious than style because it can hide a fabricated or distorted claim.

Replace vague authority with a named source, date, method, and limit. “A 2023 study tested seven detectors on TOEFL essays” is auditable. “Research proves AI detectors are biased” is a slogan.

6. Every observation becomes historic

AI drafts are prone to significance inflation: a feature is “groundbreaking,” a workflow is a “game changer,” and every launch “marks a pivotal moment.” When every paragraph reaches for consequence, none of the consequences feel earned.

Concrete effects are more persuasive. Say that a change removes one API call, cuts a batch from 20 minutes to 12, or lets an editor verify a source. Scale the adjective to the evidence.

7. The conclusion summarizes but never chooses

Weak conclusions repeat the headings, say the future is exciting, and advise readers to embrace innovation responsibly. Strong conclusions resolve the tension introduced at the start.

For this article, the resolution is simple: readers can recognize weak, generic writing, but should not convert that editorial judgment into a confident claim about authorship.

A visual framework separating writing signals from evidence of authorship

What is not reliable evidence of ChatGPT use

Online checklists often treat ordinary language features as fingerprints. None of the following is reliable on its own:

  • em dashes or semicolons;
  • words such as “delve,” “nuanced,” “landscape,” or “tapestry”;
  • a three-item list;
  • grammatically clean prose;
  • headings and bullet points;
  • an analogy or decorative metaphor;
  • a low or high score from one detector;
  • publication frequency without access to the writer's workflow.

The last item deserves care. Publishing several researched pieces every day may raise a reasonable operational question, especially for a solo author. It still does not tell you whether the work was generated, dictated, ghostwritten, collaboratively edited, or assembled from prior research. Ask for provenance rather than inventing it.

Why AI detectors should not be treated as verdicts

AI-text detection is classification under uncertainty. A tool estimates whether statistical patterns resemble examples in its training data. It does not recover a document's revision history.

OpenAI retired its public AI-writing classifier after reporting a low accuracy rate. In its own challenge set, it caught only 26% of AI-written text as “likely AI-written” and mislabeled 9% of human text.[2] Those figures describe one old classifier, not every current product, but they illustrate why a score needs context.

A peer-reviewed 2023 study also found that several detectors disproportionately mislabeled writing by non-native English speakers.[3] Later detector designs and evaluations may perform differently. The responsible conclusion is not “all detectors are useless.” It is that consequences should not rest on one opaque score.

Use a detector to prioritize review, compare passages, or evaluate your own content pipeline. The reAPI AI text detector, for example, returns a score that can be combined with editorial checks. Do not describe that score as proof.

A fairer five-step review for suspicious writing

1. Separate quality from authorship

Mark generic claims, repetition, weak sourcing, and factual errors without guessing how they were produced. A poor article remains poor even if a human wrote every word. A useful, accurately sourced article does not become useless merely because software assisted with a draft.

2. Verify the consequential claims

Open the sources. Check whether the cited page supports the number, whether the date is current, and whether a comparison uses the same conditions. Fabricated references and confident mismatches are stronger editorial evidence of an unreliable process than any transition word.

3. Look for information gain

Ask what the page contributes beyond the first five search results. Original testing, a worked example, a decision framework, a primary-source comparison, or a clear limitation all count. A beautifully formatted summary of summaries does not.

Google's current guidance takes the same practical position: generative AI can help with research and structure, but scaled pages without added value may violate its spam policy; accuracy, quality, relevance, and context matter more than the mere presence of automation.[4]

4. Check provenance when stakes are high

For a newsroom, school, regulated workflow, or competition, ask for outlines, source notes, version history, or a disclosure of allowed tools. Process evidence is more useful than stylistic intuition because it addresses what actually happened.

5. Use a human decision

Treat tool output and stylistic patterns as inputs. Let a qualified reviewer weigh the stakes, policy, evidence, and the author's response. If a false accusation could affect a grade, job, or reputation, the review threshold should be correspondingly high.

A compact editor's checklist

Before publishing, ask:

  • Does the first paragraph make a claim specific to this topic?
  • Does each section add evidence or a decision, rather than restate the premise?
  • Can the reader identify the source behind every important number?
  • Are limitations stated next to the claims they limit?
  • Does sentence length and paragraph shape vary naturally?
  • Have decorative transitions and inflated adjectives been cut?
  • Does the conclusion decide something?
  • If AI assisted the work, has a human verified every factual claim?

If the answer is no, improve the article. Do that regardless of who drafted it.

FAQ

Can people accurately tell when writing is from ChatGPT?

People can notice generic or repetitive prose, but style alone cannot reliably establish authorship. Human and AI writing overlap, and many documents are produced through mixed workflows.

What words make writing look AI-generated?

Words such as “delve,” “landscape,” and “nuanced” are frequently cited online, but no word proves AI use. Repetition, vagueness, weak evidence, and uniform structure are more useful editorial signals than vocabulary blacklists.

Are AI detectors accurate?

Accuracy varies by detector, text length, language, model, editing, and evaluation set. Use detector scores as screening signals and combine them with source checks and process evidence, especially when a false positive has consequences.

Does Google penalize AI-written content?

Google's guidance focuses on whether content is accurate, useful, original, and made for people. It warns against scaled generation without added value rather than declaring all AI-assisted content ineligible.[4]

Should writers disclose ChatGPT use?

Follow the policy of the publisher, employer, school, or platform. When automation materially shaped reporting or the final text, a clear disclosure can give readers useful context. Grammar assistance may be treated differently from generated claims or passages.

The useful tell is editorial, not forensic

You may instantly know that a page is generic. You may know that it lacks sources, flattens every idea into the same cadence, and reaches a conclusion without making a decision. Those are valid reasons to stop reading or send the draft back.

They are not proof that ChatGPT wrote it.

The fairest standard is also the more demanding one: inspect what the article claims, what it adds, and how its author can support the process. That catches low-value content whether it came from a model, a content farm, or a tired human working from a template.

References

  1. Matt Lillywhite. I'll Instantly Know A Writer Used ChatGPT When I See This. The Daily Draft, July 2026. medium.com
  2. OpenAI. New AI classifier for indicating AI-written text. Updated July 2023. openai.com
  3. Weixin Liang et al. GPT detectors are biased against non-native English writers. Patterns, 2023. pubmed.ncbi.nlm.nih.gov
  4. Google Search Central. Guidance on using generative AI content on your website. Updated December 2025. developers.google.com

Further reading