← Video Library · AI Tools
AI Content Flagging by Google | Interview with Jon Gillham
Published March 14, 2023 · 225 views on YouTube
Key takeaways
- AI detection tools like Originality.AI score content on a 0-100 probability scale rather than a simple yes/no verdict.
- Google's guidance centers on E-A-T (expertise, authority, trust) — the stated target is low-quality content, not AI use itself.
- A common publisher workflow is human-written openings paired with AI-assisted supporting sections, edited and fact-checked by an expert.
Will Google penalize AI-generated content?
Jon Gillham, founder of Originality.AI, says Google’s stated position is that content quality is the issue, not AI use itself — but he expects Google will eventually be forced to take a stronger stance against AI-generated content as search quality declines. He frames AI as a shortcut similar to exact-match domains or paid links: shortcuts in SEO tend to get closed once they become widespread.
How do AI content detectors like Originality.AI work?
Most AI detection tools, including Originality.AI’s baseline approach, are built around a linear probabilistic model similar to how large language models generate text: the tool estimates the probability that each word in a sequence was the “next most likely” word an AI would have chosen, then aggregates that into a score. Originality.AI adds to this by training its own detection model on roughly 100,000 GPT-3 generated inputs, rather than relying purely on the probabilistic method.
How accurate is AI content detection?
According to Gillham, Originality.AI tests at 94% accuracy on raw GPT-3 output, with an update in progress expected to bring that to 97%. The tool assigns a confidence score from 0 to 100 rather than a binary verdict — a score of 75, for example, means the model is 75% confident the content was AI-generated, not that it is certain. He notes results can vary depending on the AI writing tool used and how much a writer customizes prompts or adds their own voice, since tools like Jasper build extra features on top of the same underlying models (OpenAI’s API) that add variability to detection outcomes.
How should publishers use AI content without risking a penalty?
Gillham’s advice is to evaluate AI content from a human-value standpoint rather than chasing a specific detection score, since Originality.AI’s score is not the same signal Google uses. He describes an advanced workflow some publishers use:
- Write the opening portion of an article (he cites roughly the first 500 words of a 1,500-word piece) with human-generated content that carries personal anecdotes, opinion, and unique data.
- Use AI to help produce supporting sections — background information, FAQs, and general commentary — that build out topical coverage of the subject.
- Have the AI-assisted sections edited and fact-checked by a subject-matter expert before publishing.
He ties this back to Google’s E-A-T framework (expertise, authority, trust): Google’s stated goal is to surface content produced with genuine expertise — a veterinarian writing about animal care, for example, rather than AI generating advice with no expert review — and to keep penalizing low-quality content regardless of whether AI was involved in producing it.
Should site owners treat every site the same way?
No — Gillham describes making a per-site decision. On some of his own sites he is comfortable publishing AI-assisted content because he has a process for adding value; on other sites, particularly ones with an established history of human-only content, he chooses not to introduce AI content at all because of the added risk. He also raises an open question about GPT-4 and future models: whether they will eventually produce content in a way that’s effectively undetectable, and says nobody in the space can say for certain yet.
Full video transcript
There’s a great use case for AI — I use it on some of my own sites, but it can be a challenge when people are paying $100-$200 for an article from a writer and that article is being produced by AI and they’re not aware of that. Then the big problem comes up: if there’s a risk that Google is going to hurt AI traffic (which I think there’s some question on whether or not that’s the case, but regardless there is a risk), publishers want to be able to make that risk-based decision on whether they publish AI content or not.
Yeah, I think it’s a huge topic right now, and there’s even been a boom in ebooks, especially with ChatGPT coming out — people just putting out content that, in a lot of cases, is just garbage, not much real thought put behind it, just generating for the sake of it. I think that’s why Google’s putting the hammer down on bad content like they’ve always done — nobody wants a bunch of trash out there that’s not useful for people. So I’m wondering, from your point of view as a publisher who’s been around the space a long time, I’ve seen a lot of conversation about how Jasper output sometimes gets flagged and sometimes doesn’t, ChatGPT tends to get flagged more on the first round sometimes and sometimes doesn’t — people see mixed results with these AI content detectors. Can you break down how Originality specifically operates, how it determines what’s original and what’s not? I’ve had content from Jasper without any edits that scored 99 to near 100, and I’ve also had the same or similar inputs come out totally different. Can you touch on that?
Yeah, so most tools in the space are based on learning GPT-3 and 3.5 output, which Jasper plugs into along with other tools, and then they add some extra stuff on top, as do other tools — but that’s the core starting point for a lot of these tools, the OpenAI API used to create the content. The way some AI detection tools work is with a linear probabilistic model — the way the NLP works is: what’s the highest-probability next word, plug that in, what’s the highest-probability word after that, plug that in — and detection tools build a giant linear probabilistic model of what the probability of each next word was, matching what they’d expect OpenAI’s GPT-3.5 or GPT-3 to have produced, and that gives it a score.
Originality does that, but we also built our own AI, trained on a large set of inputs — around 100,000 inputs from GPT-3 — and trained our own AI on what that output looks like. One of the frustrating questions that comes up is the same one you’d ask of any AI doing stock picking or prediction: the honest answer is that’s what the AI was trained to do. We test accuracy on raw GPT-3 output at 94% accurate, about to bump up to 97% with a new update. The way our scoring works is we assign a probability — so if we hit 75, that means we have 75% confidence that piece of content was AI-generated. We don’t always get it right, but we assign a probability from zero to a hundred representing the AI’s confidence that the content was AI-generated or not.
So is the goal with generating these probabilities — obviously to get the answer right more often than not — but is it also because some of Jasper’s outputs get flagged as pretty much original, and even ChatGPT, whatever tool you’re using, sometimes gets a totally different result? Is that how all the tools operate — kind of rolling the dice, not the best analogy, but doing a lot to train it to pick up patterns, and maybe depending on the quality of the prompt and the context you bring to it, it’s not necessarily going to come out as something AI would have written because you’re bringing your own uniqueness to the table. Does that sound right?
Yeah, I think for sure. As more models come online — it’s currently plugged into the same AI in the form of OpenAI’s, but as more models come online — the rate of detection… at what point will AI be good enough that it becomes undetectable? I don’t know the answer. GPT-2 detectors continue to be reasonably accurate on GPT-3 output. AI tools right now all kind of start from a similar point — OpenAI’s API — and then add extra features around it, and those extra features increase the chance of tricking a detection tool. Long term, it depends on your use case as a publisher and as a writer — but if you’re producing content, worrying too much about AI detection scores doesn’t make a ton of sense, because we’re not Google, and our detection score doesn’t necessarily mean that’s the same way Google is going to view it. It’s more important to look at it from a “does this add value” standpoint. On some of my sites I’m using AI content and happy to be using it — we have a method for adding value, and I don’t care as much what my AI detection score is on those. But yeah, you’re thinking about it correctly that some AI tools with extra features will be more likely to bypass AI detection, and that leads into: what will GPT-4 produce? Will that get us to the point of asking it to write in any format and it produces something undetectable? I don’t know, we’ll see. There’s a lot of unknowns in this space, and I think it’s perfectly normal that nobody really knows the answer right now — anyone claiming they do has probably got something a little sketchy going on, just because of how fast everything changes here.
So thinking in terms of people already creating content, whether it’s ChatGPT, Jasper, or other tools — first and foremost you want to add value with your content in some way. Some websites add value in different ways versus people expecting to hear from their favorite blogger and thinking “who wrote this, this doesn’t even sound like something a human wrote.” Other places people are just looking for a blurb, click here, buy this, and it’s not as relevant. Are you concerned Google will penalize the type of sites you’re running, or is your strategy less reliant on Google? I’m curious because it seems like even with these tools you can test content, and then Google can come out and say “we decided that’s AI” regardless.
Yeah, I think anytime there’s a shortcut in SEO it usually gets closed pretty quickly. I think AI is a shortcut, the same way exact-match domains were years ago, the same way paid links and other methods have previously been used to shortcut producing good content — that shortcut usually gets closed. So I am concerned that Google will be forced to punish AI-generated content, and their approach will struggle to differentiate between partially used AI, fully used AI, and human-generated content. Anytime Google makes an update there are incorrect losers — people who get punished, likely not intentionally, and Google doesn’t roll it back. I think it’s a pretty safe assumption that everyone wants to write with AI and no one really wants to read AI. There are exceptions with ethically used AI that adds extra value — Google says that will be fine — but when they’re faced with declining search result value, will they be forced into a strong position against AI-generated content? I think the answer is yes.
The other mental exercise I do on this: if I were to buy a site I cared about, at a price point I significantly cared about, and it has not historically used AI-generated content, do I want to introduce that risk onto that site? I make a decision on some sites that this is human-generated-only, and I make a different decision on other sites where AI-generated content makes sense. I think it’s in line with Google’s guidelines, and I’m happy to push out AI-generated content on some sites and not others.
Yeah, it sounds like you need to pick which sites get this type of content, and at what level you’re letting AI write the finished product. As more people tap into the same source, a lot of content is going to sound very generalized and similar. In our AI Author Community, leading into a book challenge right now, we obviously don’t want a plagiarized book, and we don’t want a book that’s straight-up 100% AI content — but we’re using tools like Jasper to extend our writing abilities and enhance our own voice. Do you have a tip, with Originality or otherwise, for people using this along a more research-paper approach — something with a lot of human element — where does this tool plug in for authors working in collaboration with these tools, not just AI-heavy content sites?
Yeah, I think what we’ve seen — and I think this answers your question — is some advanced approaches where people say “this is the portion of the article that needs the human value component” — the human opinion, the personal anecdotes, the E-A-T layered in around a certain topic — and then there’s FAQ content or general commentary below that can be generated and human-reviewed. So if it’s a 1,500-word article, the first 500 words are human-generated, using the personal anecdotes and unique data to layer in expertise, authority, and trust into that section. Then all the content below, which builds out complete topical authority on the subject, uses AI to assist in creating it — and then it’s edited so the accuracy is there, verified by the expert author. That’s where we’ve seen people take the most advanced, nuanced approach to using AI detection and producing content efficiently that’s the best fit for the given search term.
You mentioned E-A-T — what’s that acronym? It’s expertise, authority, and trust — how Google communicates what they want to see in content. Google has said AI is not bad, content is bad. They don’t like crappy content, and they’re good at detecting crappy content. The opposite of crappy content is good-quality content produced with a lot of expertise, authority, and trust. They don’t want me writing about a veterinary product — they want a veterinarian to write about animals and provide commentary. They don’t want AI saying “here’s the optimal diet for a dog with Lyme disease” — they want a vet who is an expert on that topic to produce that piece of content.
That makes perfect sense, and I think that’s a great way to approach it, especially for authority figures, authors, or business owners and thought leaders — taking that approach with any content, whether AI is part of the equation or not, is going to help content do better on Google and be more useful to the people you’re putting it out for. All that said, tools like the one you’ve built are really cool — I’ve been testing it out over the past couple weeks and it’s been handy, especially the plagiarism checker. For those interested in checking it out more, where can they find you, and any last words for people?
Yeah, that sounds great — people can check out originality.ai and then find me on LinkedIn or Twitter, those are probably the best spots to connect. We’re adding in additional functionality around readability scores and other detection features, working toward being that last place to have full awareness of the content you’re putting out — where it sits on plagiarism, AI, and readability — and keep becoming the most complete tool, especially for web publishers looking to make sure the content they’re publishing has the best chances of success online.
Love it, this is awesome, John — thank you so much for your time breaking down Originality.ai and the world of AI content detectors, a fast-paced industry. Go check out John and this tool, make sure you’re running your stuff through sound and human review, and don’t plagiarize obviously — make sure you’re using tools to check your work before you click publish. John, thanks so much again, we’ll see you on the other side.