← Video Library · AI Tools

AI Showdown with ChatGPT Claude Grok and Perplexity

Published April 25, 2025 · 718 views on YouTube

Key takeaways

Which AI tool won a head-to-head deep research showdown: ChatGPT, Claude, Grok, or Perplexity?

There wasn’t one clear winner. Grok finished first but gave the shortest, most condensed output. Claude finished second and produced the most comprehensive, best-structured report. Perplexity finished third but didn’t follow the requested recipe structure. ChatGPT took the longest but delivered a highly detailed, citation-heavy report with images once it finished.

What was the test setup?

During a live session at the Scale with AI Summit 2025, the same exact input was given to all four tools at once: the Buyer’s Brief / Market Sauce 9000 recipe (a structured research prompt covering company profile, demographics, psychosocial deep dive, market statistics, customer segments, and a strategic creative brief) plus a single website link for a real estate business. Each tool was set to its deep research or extended-thinking mode: ChatGPT’s Deep Research, Claude with extended thinking, Grok’s Deep Research and Think modes, and Perplexity’s Pro Research mode.

How did each tool’s speed and depth compare?

Tool Finish order Speed notes Depth / structure notes
Grok 1st Finished in 2 minutes 37 seconds Followed the recipe’s section labels but skipped roughly two-thirds of the requested content; described as “less in-depth” and more condensed once its Think mode was tried instead of Deep Research
Claude 2nd Hit the max message length once and needed a “continue” / re-extend-thinking prompt to finish Most comprehensive output of the four; closely matched nearly every section of the buyer brief blueprint
Perplexity 3rd Finished before ChatGPT Good for common pain points but did not follow the recipe’s structured-output format the way Claude and ChatGPT did
ChatGPT 4th (still running near the end of the session) Slowest to finish Very detailed once complete — citations, images of the business’s area, and a full appendix of sources

What did the test reveal about picking an AI tool for research work?

No tool was best across every dimension, so the presenter’s practice is to keep several subscriptions active at once and switch between them depending on the task. Monthly cost was cited as roughly $20–30 per tool for most of them, with the presenter’s own Claude subscription at about $20/month; Grok was being used without a paid plan at the time of the test.

Full video transcript

We’re going to take the buyer brief, the Market Sauce 9000 recipe. We’re going to give it a link to a website and we’re going to tell it to go. Then we’re going to observe the outputs of each, like an AI cocktail if you will, using a few different tools — Perplexity, Claude, Grok, and ChatGPT. We’re going to use the same exact inputs for each one of these. This is something anybody here can do inside of our community. Get access to the free digital copy of the buyer’s brief book. You can also buy it on Amazon.

Hello and welcome in everybody to day two of the Scale with AI Summit 2025. We’re going to use your website. We’re going to create a mega brief — I don’t even know what to call this, a superstar buyer brief. So I’m going to start with ChatGPT, then we’ll move on from there, but I’m going to do the exact same prompt to structure it every time: do deep research, follow these instructions for the Market Sauce custom recipe, generate a complete report at the end with direct citations where they are cited. That seems like it’d be helpful.

So this is going to be the start of this prompt — what we call “setting the table.” Here at Market Sauce, we’ve got our manuscript for the buyer brief (subscribe to our community). Scroll down to the recipe, part two, instructions and context. This whole thing is copied, and by the way, you can also chunk this into much smaller segments — I’m just giving it a mega prompt and seeing what these tools can do. I’m taking parts one and two. This is going to be interesting if it actually includes all these sales arguments. All these tools want to be the best; we’ll see who’s the best and can actually handle this right now, because whatever handles this to the highest degree is usually the one I work with the most at any given time, and that changes on a daily basis.

I don’t want to get down into the social micro content yet, but I do want to finish it off with the brand voice, identity, and core content. I’m pasting that in. I’m going to do one more thing — go to the bottom for the insights and action items. I’m going to edit that slightly (not in the book, because everyone else would lose access to that), so I just added this to the bottom: a strategic creative brief.

So we started from the top with the realty website. I selected deep research. The first thing ChatGPT asks me is what my goal is, and I’m just going to say “everything, figure it out, AI.” This is really where the operational piece comes into play, because if I’m doing this for a client and I don’t give it that context, the output could be totally off from what they want. If they say “this is what I want to do” and you’re actually only looking to buy houses, the output’s going to be geared toward a different direction. So what’s your number one goal with any of this, and make sure that context is in there. ChatGPT does that, and clarifies questions if needed.

This might take 20 minutes, guys, so I’m going to keep the ball rolling and see if we can do it much quicker. Deep search, choose style — I feel like Perplexity is going to be the struggle one here. Pro deep research. Okay, so now we’re cooking — Claude’s thinking, Claude’s cooking, Grok’s cooking, ChatGPT is still trying to figure out what it wants for dinner. Perplexity is cooking too — it’s starting now, it just took a little more to get going. It feels like watching a go-kart race: who’s going to make it first, and who finishes first versus who finishes best. I think first will be Grok, best will be ChatGPT.

These things are thinking now. When you ask what the difference is between the buyer’s brief and anything else, it’s the method for how we go about this research. At the end of the day we’re not locked into one platform — we’ve created a studio as an easy gateway trained on these, and I’d still recommend everyone have several of these tools at any given time. From a monthly subscription perspective, access to all of this right now is like less than $20 to $30 a month per tool. My Claude is $20 a month; I tested the $200/month ChatGPT Pro tier but I’m not using that one on this account right now — this is still my team’s account, which I think is $60 a month since there’s another member on it. Grok — I’m not paying for anything on Grok right now.

Grok finished first, so — actually Claude finished first. Perplexity is still cooking. And look at that — we got “Ander Lucia realty has established itself as a premier real estate agency,” so it’s going right now. Claude hit the max message length and paused the response. What I like to do with Claude right now is, after one response, go to retry and have it rethink and extend its thinking again. I don’t have a whole research-backed statistical report showing the difference in outputs, but I’ve noticed it thinks a little longer and gets a bit more aligned in its responses, so I default to that. Now it’s going back and doing the research side.

Grok finished in 2 minutes and 37 seconds — researching tips, services. Looking at this, I see all the product description, citations, target audience analysis — these are all from the recipe, so it’s following those structured outputs. Psychosocial deep dive — and notice it skipped over probably two-thirds of the actual prompt; it just didn’t incorporate that into the outputs, like the negative statistics, the false solutions, offer benefits. One of the clarity benefits of our reports is you can read the full report and at the end get micro segments you could approach. A lot of that can get skimmed over if you’re not focused on where you’re going to build a campaign around — you’re probably not going to build a campaign for ten different avatars at once. You want to start with one blueprint for one avatar; if you try to talk to everybody with one message, your copy won’t resonate the same. We get to the strategic creative brief and it still glosses over a lot of that, but then we get market statistics — interesting, it’s citing this. Thanks, Grok.

Claude’s still rocking and rolling. Perplexity’s done. ChatGPT’s still in the kitchen cooking, and I like watching it do this — this is wild, look at all of this. It’s going through the proxy method, breaking down its analysis, identifying different targets. What’s interesting is even when we were doing some of the initial buyer’s briefs manually, with prompting and making sure outputs were good, feedback we’d get was “what channels are they on?” and naturally you’d go ask ChatGPT and engage with it, which is where we went to the simulation approach with bot versions of these. But now these tools are pulling from so many points of reference that they’re answering questions we wouldn’t have even thought to ask.

I’ve never given this level of a prompt to Claude before. I feel like we knew Grok was going to finish first — cool, but Grok, you need to work on this a bit more, or maybe this isn’t the best use case for Grok. Maybe it’s much better at other types of tasks. I think it’s less in-depth. I’ve noticed this with Perplexity too — it’s good for common pain points, but Perplexity doesn’t follow recipe-style structured outputs the same way Claude or ChatGPT does.

Section one: company profile, executive summary, core services, demographics — that’s how it starts. Deep dive psychosocial — this is already much more in-depth. Claude is still thinking, not anywhere near done. It ran for 10 minutes last night breaking down a report from our five-hour-plus transcript and started to repeat itself after 8 minutes, losing some contextual information, which it didn’t do for a 3-hour interview I tested the same process with. So there are still some constraints — first-world AI problems, like only being able to write 5,000 words at a time — but we’re here troubleshooting them.

When I started doing the buyer brief work in the community and asked everyone to submit their 30-32 page word input, I used Claude for that first huge batch of testing because its outputs are just so comprehensive — this is very good and it’s getting everything. That’s the key here: you saw how truncated Grok was, and I don’t even know how output-heavy ChatGPT’s deep research mode is going to be. I’ve never done this type of input with ChatGPT on deep research mode before — sometimes deep research focuses way more heavily on finding research for everything, so I’m curious how much it moves over.

Maybe with Grok we should have tried the thinking part — let me try that. Grok only does deep research or think, so I’ll have it think now. If you were at our book launch, would you believe that was literally about a year and four days ago, when we launched the book with a live session and did ten of these on autopilot using our Market Sauce automation in Make? That worked, and that was one of the big milestones for the Make automation on there.

This is incredibly comprehensive for a quick Google search. Something Mario demonstrated that we don’t have in ours was the “3 a.m. thoughts” and the layers to those thoughts — that’s not part of the Market Sauce 9000, but that could be part of a future edition, or something you build into your own version with your own recipes and insights taken from how other tools operate and output. This is really good right here — primary cause, triggers, information overload, explanation — it’s adding those explanations. I’ve got nine minutes before Roger comes on for the first session of the day, and I hope this finishes by then. Success myths — this is brilliant. Offer benefits — this has hit every single section I’ve seen, maybe missing one or two, from the master buyer blueprint inside the buyer brief book.

That’s part one. Now, since I gave it the entire strategic creative brief, we have ten subsegments — detailed sub-subsegments I can assess and work seamlessly into part two. There’s a point where you probably want to be more hands-on producing some of these, but don’t forget all I did was give it a website. Literally that’s all I did, then ran it once and said “no, do this again with extended thinking.”

Meanwhile ChatGPT is still in the kitchen cooking — everyone will have eaten, filled up, and launched five products by the time ChatGPT is done, or they’ll have to reassess their entire launch protocol. We’ve hit the max limit for a message with a positive response, so you can continue, and now it’s breaking down each customer segment. Investment protocol — all the different subsegments right there, highlighting and making edits. Now it’s going into the sales argument, and with this section — from working through it with Ron Lynch and Marketing Mercenary — you want to write a sales argument that sells your product using a problem-agitate-solution structure mixed with a hero’s-journey prompt structure, articulating how your product solves the problem for each market segment. If you can write a short sales letter for each one from this section, it gives you additional context to understand how deeply your product serves that particular market segment. This is a ton of research to work through, so the next step is Assad and his team reviewing it: did you miss anything with your marketing, or have you even considered some of these audiences? Some of this can come up while you’re in Market Sauce or ChatGPT or otherwise, which is why we start with this process and then expand into other tools. If I want to make a campaign, I’d pick one of these market segments, then use something like Mario’s copy bots — “here’s all this information, now write all the copy” — or use your own tools, or use it inside Benson, filling out your blueprint there and comparing the two. This is just straight from Claude and a website plus our Market Sauce 9000 recipe.

ChatGPT is going to be working through the night on this. In terms of legs up in the race — you mind sending me this over? It’s going to be $49.99, I’m just playing with you, happy to share this with you. One of the benefits of joining these calls live is seeing what’s possible with these tools. What we just did there was ask Grok to redo it with think instead of deep research, and that seemed to do a lot better — Grok is a lot more condensed with its wording. I’m not going to be able to read all of this right now, but you can see it’s doing more of those sections and not truncating as much — solution, strategic brief, USP, still shorter. Claude definitely had the most content output and detail, and ChatGPT finished with 60 seconds left.

So we see the output — it’s even got pictures of the area, look at that, pictures, and see how it’s citing all the sources here inside target audience, super detailed comprehensive analysis, a lot more content and links, threats, marketing review strategy. This is on the deep research front — if I did this exact same thing without deep research, the same way Grok worked, I bet it doesn’t follow through the same way; it follows into a bit more of a structured output. But we did get the citations appendix. So what I’d say to do is start and prime the chat with deep research, then put that exact same context in, and it’s already going to have that research to inform it. I just gave this to four different AIs, pitting them against each other doing the research, and you can see the outputs and what they’re capable of does vary.

Watch on YouTube

Related tutorials