We Asked 4 AI Assistants Which CRM to Buy — Every Day for 18 Days

Created with Claude, reviewed by ·

  • ai-visibility
  • geo
  • data-study
  • competitive-intelligence

Between August 1 and August 18, 2026, we sent the same 10 CRM buying questions to four AI assistants, once a day, every day — 720 answers in total. Between them they recommended 39 distinct CRM products, about 5.3 per answer. On the six brands we tracked, they agreed with each other 74% of the time. Across all 39 products, that agreement falls to 49%.

That two-level result is the finding. The assistants converge on the two or three names everyone already knows, and split roughly down the middle on everything below them — which is precisely the range where most software companies live. The most-recommended brand in our set, HubSpot, appeared in 98.8% of answers. The least, Monday CRM, in 12.0%.

But the sharpest result is not about the assistants at all. It is about which model you are talking to. The four endpoints we queried are not the same size — one of them is a small, cheap model — and the gap between them is not spread evenly. On the household names they are close. On newer and smaller products, the small model returns almost nothing: three CRMs that the other three assistants each recommended in roughly one answer in five were named zero times in 144 answers.

The short version for anyone selling software: there is no such thing as “what AI says about your product.” There is what a specific model, at a specific size, said on a specific day — and if you are not a household name, that distinction is the whole ballgame.

Full disclosure before the data: Agonai is our product, it does this kind of tracking, and we used it as the instrument for this study. We picked a category we do not compete in — CRM for small businesses — precisely so the findings would not be about us. Method, raw counts and limitations are all below; take the numbers and check them.

What exactly did we measure?

We defined the metrics before looking at any data, to avoid fishing for a headline:

  1. Share of voice — in what percentage of answers does each brand appear?
  2. Volatility — how often does the recommended set change from one day to the next?
  3. Divergence — when two assistants get the same question on the same day, how much do their answers overlap?
  4. Sentiment — how is each brand described, not just whether it is named?
  5. Citations — which sources does each assistant point to when it recommends?

Method

We are declaring all of this so the study can be checked and repeated.

Window2026-08-01 → 2026-08-18 (18 consecutive days, no gaps)
CategoryCRM for small businesses
Brands trackedHubSpot, Salesforce, Pipedrive, Zoho CRM, Monday CRM, Attio
AssistantsFour model endpoints, same id sent every day (see below)
Queries10, fixed wording, never edited mid-window
Frequency1 run per query per assistant per day, same time of day
Total answers720 (18 × 10 × 4), every day complete, no partial days
AccessModel APIs, default parameters, no browsing tools except where the provider does it natively

The exact models. Every answer in this study was requested from one of these four endpoint ids, and the id we sent never changed across the 18 days:

AssistantModel IDAnswers
ChatGPT (OpenAI)gpt-4o-mini-2024-07-18180
Claudeclaude-opus-4-8180
Geminigemini-3.5-flash180
Perplexitysonar180

One caveat on that table, because “pinned” would be too strong a word: only gpt-4o-mini-2024-07-18 is a dated snapshot. The other three are aliases, and a provider can repoint an alias mid-window without announcing it. We did not record the underlying snapshot each alias resolved to on each day, so what we can honestly claim is that the request was identical every day — not that the model behind it provably was. If we run this again we will log the resolved snapshot per call.

These are also not the same class of model, and that matters more than anything else here. claude-opus-4-8 is a frontier model; gpt-4o-mini is a small, cheap one, and OpenAI documents its knowledge cutoff as October 2023 — the 2024-07-18 in the id is the snapshot date, not the cutoff. So where we write “ChatGPT” below, read “gpt-4o-mini”: not the model behind the consumer ChatGPT app, and not a claim about OpenAI’s policy. Finding 1 shows exactly where that distinction bites.

The 10 questions, verbatim:

  1. What’s the best CRM for a small business?
  2. What’s the best CRM for startups in 2026?
  3. Recommend a CRM for a 10-person sales team.
  4. What’s the best free CRM?
  5. What are the best HubSpot alternatives?
  6. Salesforce vs HubSpot — which should I choose?
  7. What’s the best CRM for solopreneurs and freelancers?
  8. Which CRM has the best email integration?
  9. What’s the best affordable CRM under $50 per user?
  10. Which CRM should an early-stage SaaS company use?

Two things we had to decide before counting

We excluded questions 5 and 6 from the brand comparisons. Both name a brand in the question itself, so counting them would have manufactured a result: across all 720 answers HubSpot appears in 99.0%, but two of the ten questions put the word “HubSpot” in the model’s mouth. Every share-of- voice number below is computed on the 8 neutral questions — 576 answers. The full 720 are used for the volatility and divergence work, where the question wording is held constant anyway.

We normalised product variants to their vendor’s CRM, across the whole corpus — not just the six brands we tracked. This matters, so here is exactly what it means. Raw extraction pulled 282 distinct name strings out of the answers. Most of that number is noise of two kinds. First, the assistants do not name products consistently: “HubSpot CRM”, “HubSpot Sales Hub”, “HubSpot Starter Suite” and plain “HubSpot” are one product; so are “Freshsales”, “Freshworks” and “Freshsales by Freshworks”; so are “Monday.com” and “Monday.com CRM”. Second, a large share of those strings are not CRMs at all — Gmail appeared in 175 answers, Google Workspace in 105, Outlook in 103, and the list also contains Slack, Notion, Airtable, Stripe, Zapier and even Forbes and PCMag, because assistants name integrations and cite outlets in passing.

Counting either of those as a product a buyer was recommended would be wrong, and counting them as disagreement between assistants would be badly wrong — it would score “HubSpot” versus “HubSpot CRM” as two systems recommending different things. So every figure in this article runs on the normalised, CRM-only set: 39 distinct products. We did not merge in vendors’ non-CRM products — Zoho Books, Zoho Mail, HubSpot Marketing Hub, Salesforce Marketing Cloud and Pardot count as nothing. Judgment calls we made and would defend: Bigin counts as Zoho (it is Zoho’s small-business CRM), “Sales Cloud” counts as Salesforce, Freshworks counts as Freshsales, and bare “Zoho” counts as Zoho CRM in a CRM question. If you disagree with any of those, the effect is a point or two, not a reordering.

Why 18 days and not 30

We originally planned a 30-day window starting in mid-July. It broke, and the honest version is more useful than a rounded number.

Our own account hit its monthly quota on July 23. Collection stopped silently and did not resume until August 1, when the counter reset — eight days with no data, and nothing told us at the time. Two days in the surviving July stretch were also partial (7 and 5 of 10 queries).

We could have stitched July onto August and declared the gap in a footnote. We did not, because a series with a hole in it invites exactly the objection that would sink the study. What makes this data worth citing is that it is continuous, not that it is long. So the window is the 18 days that are clean.

There is a lesson in the failure worth more than the eight days: a monitoring program that dies quietly is worse than no monitoring, because you keep trusting a number that stopped updating. If you run something like this, alert on the absence of data, not just on changes in it.

Percentage of the 576 neutral answers in which each brand appears at least once.

BrandChatGPTClaudeGeminiPerplexityAll four
HubSpot100.0%100.0%96.5%98.6%98.8%
Pipedrive86.8%87.5%77.1%77.8%82.3%
Zoho CRM100.0%84.7%59.0%54.9%74.7%
Salesforce68.8%64.6%47.9%26.4%51.9%
Attio0.0%25.0%25.0%22.9%18.2%
Monday CRM11.8%2.8%22.2%11.1%12.0%

The leaderboard is the least interesting column. HubSpot is named in essentially every answer to every question; that is what category dominance looks like from inside a language model, and no amount of clever positioning is going to dislodge it this quarter.

The number to sit with is the spread across the row. Zoho CRM is in 100% of ChatGPT’s answers and 54.9% of Perplexity’s — a 45-point gap on identical questions in the same 18 days. Salesforce runs from 68.8% down to 26.4%. And Attio, a real product with real customers, is recommended by Claude, Gemini and Perplexity at roughly a quarter of the time each, and never once by gpt-4o-mini in 144 answers.

The zeroes are about model size, not about ChatGPT

That Attio result is the most quotable number in this study and we are going to argue against it ourselves, because taken at face value it is misleading.

Attio is not the only zero. Look at what else gpt-4o-mini did not name, next to the three larger endpoints:

Productgpt-4o-miniClaudeGeminiPerplexity
Attio0.0%25.0%25.0%22.9%
Folk0.0%18.8%12.5%18.1%
Less Annoying CRM0.0%18.1%13.2%14.6%
Close0.7%16.7%15.3%20.1%
Streak2.1%16.7%16.7%13.9%
HoneyBook0.7%11.8%13.2%1.4%

Now look at the same model on the established names: Zoho 100% and Freshsales 82.6%, the highest of the four endpoints on both, and HubSpot 100%, level with Claude.

And here is the part that rules out the easy explanation. It is not naming fewer products. Averaged per answer, the four endpoints recommend almost the same number of CRMs:

EndpointCRMs per answer
Claude5.6
ChatGPT (gpt-4o-mini)5.4
Perplexity5.2
Gemini4.9

So the small model is not being terse. It fills the same number of slots — and fills them with the household names, over and over, while the long tail never appears at all.

The honest reading is therefore not “ChatGPT refuses to recommend Attio.” It is that a small model whose knowledge cutoff predates much of the current long tail recommends from a narrower pool, and every product outside that pool vanishes at once. We cannot separate model size from cutoff date with this data, and we are not going to pretend otherwise.

Which leaves a more useful finding than the one we started with: the size of the model your buyer is using decides whether you exist. If you are a household name, it barely matters. If you are Attio, Folk or Less Annoying CRM, the difference between a frontier model and a cheap one is the difference between one-in-five and never. That is not a thing you can optimise your way out of with content, and it is not visible at all if you check your AI visibility on one assistant.

Finding 2: do the assistants agree with each other?

It depends entirely on how far down the list you look.

Overlap is measured pairwise: for the same question on the same day, what share of the products named by either assistant were named by both (Jaccard). Averaged over all 864 same-question, same-day pairs — 18 days × 8 questions × 6 assistant pairings.

One denominator note, because it is the first thing we would check in someone else’s study: there are no empty comparisons hiding in that average. Every one of the 576 question-day-assistant combinations named at least one of the six tracked brands, so no pair was ever an empty set scored as 0% or 100%, and none had to be dropped. 864 is the real denominator, not a rounded one.

PairOverlap on the 6 tracked brands
Claude × ChatGPT87.3%
Claude × Gemini76.5%
Claude × Perplexity75.8%
ChatGPT × Perplexity69.4%
ChatGPT × Gemini68.1%
Gemini × Perplexity66.8%
Average74.0%

Claude and ChatGPT are close to interchangeable on the big names. Gemini and Perplexity are the two that most often go somewhere else — and note that this is not because either of them says more. Gemini recommends the fewest CRMs per answer of the four (4.9, against Claude’s 5.6); it just picks a different set. Perplexity is the most reluctant of all to name Salesforce, at 26.4%.

Now run the same calculation over all 39 normalised CRM products, not just the six we tracked, and the average overlap drops from 74.0% to 49.0%:

PairOverlap on all 39 CRM products
Claude × ChatGPT61.6%
Claude × Gemini59.8%
Claude × Perplexity45.6%
Gemini × ChatGPT45.2%
Gemini × Perplexity41.5%
ChatGPT × Perplexity40.3%
Average49.0%

That is the shape of the thing. The four assistants share a head and not a tail. They will all tell a small business to look at HubSpot and Pipedrive. What they say after that — the three or four other CRMs in a typical answer, the ones a buyer has not already heard of — agrees about half the time. Toss a coin.

If you have seen a larger divergence number quoted for this kind of study, check whether the brands were normalised. Ours drops to 28.7% if you count raw name strings — but that figure is mostly measuring the fact that assistants write “HubSpot CRM” and “HubSpot”, and that one of them mentioned Gmail. It is not divergence, and we are not going to report it as if it were.

This breaks the mental model most teams carry over from SEO. A search ranking is roughly one shared reality: you are 4th for a keyword and everyone sees you 4th. Here, four systems answer the same buyer question and agree on about half of the products they name. Optimising for “AI” as a single channel is a category error — you are looking at four channels that happen to speak the same language. We unpack the practical difference in what AI visibility tracking actually is.

Finding 3: how much did the answers move day to day?

Take every combination of question, assistant and tracked brand — 8 × 4 × 6 = 192 combinations — and for each one we have 18 daily observations of “named” or “not named”, 3,456 observations in all. 52 of the 192 combinations, or 27.1%, changed at least once during the window. The other 140 were identical every single day.

So the instability is real but concentrated: roughly a quarter of the brand-question pairs are genuinely in play, and the rest are settled.

The practical version of that number is what we call one-shot error: if you checked on one random day, how likely is your reading to disagree with the 18-day verdict? Across everything, 5.4%. But it is very unevenly distributed:

AssistantOne-shot error
Gemini9.3%
Perplexity7.4%
ChatGPT3.1%
Claude2.0%
QuestionOne-shot error
What’s the best CRM for startups in 2026?9.3%
What’s the best affordable CRM under $50 per user?7.9%
Which CRM has the best email integration?6.9%
What’s the best CRM for solopreneurs and freelancers?6.7%
What’s the best CRM for a small business?4.2%
Recommend a CRM for a 10-person sales team.3.9%
Which CRM should an early-stage SaaS company use?3.7%
What’s the best free CRM?0.9%

Volatility clusters where the answer depends on a fact that changes — “in 2026”, “under $50”, “best email integration”. It disappears on “what’s the best free CRM?”, which is effectively a fixed list. The questions closest to a purchase decision are the least stable ones, which is an unhelpful arrangement if you were hoping to check once a quarter.

A single check on Gemini is nearly five times more likely to mislead you than a single check on Claude. If you screenshot one assistant on one day and put it in a board deck, you are reporting a sample of one from the noisiest source you had.

We tracked how each mention was framed, not just whether it happened. Five of the six brands are described in almost uniformly positive terms when they come up. One is not.

BrandNamed inAvg. position in the answerExplicitly recommendedPositive framing
HubSpot98.8%1.599.3%99.5%
Pipedrive82.3%4.099.4%99.6%
Zoho CRM74.7%3.399.8%99.1%
Salesforce51.9%4.472.9%68.6%
Attio18.2%4.0100.0%100.0%
Monday CRM12.0%7.097.1%92.8%

Salesforce is named in more than half of all answers and endorsed in fewer than three-quarters of those. It is the only brand in the set that the assistants regularly bring up in order to steer the buyer away from it. Verbatim, from the collected answers:

“Avoid Salesforce at this stage — it’s powerful but overly complex and expensive for early teams.” — Claude

“Best only if you need deep customization and have the budget plus admin support; it is typically overkill for a 10-person team.” — Perplexity

“Enterprise CRMs like Salesforce are overkill, too expensive, and have a steep learning curve.” — Gemini

Two things follow. First, a share-of-voice number on its own can be actively misleading — half of Salesforce’s visibility in this category is a warning label. Second, the trend is going the wrong way for them: split the window into three six-day blocks and Salesforce falls from 56.3% to 53.1% to 46.4%, the only brand in the set with a consistent decline. Monday CRM is the only one consistently rising (10.4% → 11.5% → 14.1%).

Average position matters too, and it is the metric nobody tracks. HubSpot is named 1.5 products into the answer; Monday CRM is named 7.0 products in. Both are “mentions”. Only one of them is going to survive a buyer skimming the first paragraph.

Finding 5: which sources do the assistants cite?

Only one of the four showed its work.

Perplexity was the only assistant that returned source citations. ChatGPT, Claude and Gemini returned none through their APIs — they answer from training and do not browse by default. That is a finding in itself: for three of the four, there is no list of pages to go and influence, because the answer is not being assembled from pages at read time.

It also means there is no cross-assistant comparison to make here: three of the four cited nothing at all. What Perplexity points to when it recommends is a real question, but it is a study of one assistant rather than a corner of this one, and we are not going to squeeze it in here.

Finding 6: did website changes move the recommendations? — too short to say

We also monitored the six brands’ sites and pricing pages over the same window. Nothing in 18 days supports a causal claim in either direction, and we are not going to build a narrative out of two coincidences. Establishing whether a pricing-page change moves an assistant’s recommendation needs a window measured in months, with changes you can date precisely. We have started that clock; it is not this article.

What this means if you sell software

Four things follow from the data, in order of how much they should change your week:

  1. Stop treating “AI” as one channel — and stop treating one vendor as one answer. The four endpoints overlap 74% on the six household names and 49% across all 39 CRMs. Worse, the gap between a frontier model and a small one from the same vendor can be the difference between being named in one answer in five and never being named at all. A win in one is not a win in the others. Measure them separately, and find out which model your buyers are actually on.
  2. Stop doing one-off audits. 27.1% of brand-question pairs changed during 18 days, and a single-day check on Gemini disagrees with the 18-day verdict 9.3% of the time. A screenshot of ChatGPT recommending you is a lottery ticket, not a metric.
  3. Track how you are described, not just whether you are named. Salesforce appears in 51.9% of answers and is endorsed in 72.9% of those. If you are only counting mentions, you cannot tell the difference between being recommended and being used as the cautionary example.
  4. Alert on silence. Our own collection died for eight days without a word, and nothing told us at the time — we found out by going to look. Whatever you build, make missing data loud.

Limitations

We would rather list these than have them listed for us:

  • One run per day is a sample, not a census. These models are non-deterministic; a different hour of the day would produce somewhat different answers.
  • 18 days is short. It is enough to demonstrate instability; it is not enough to establish trends, seasonality, or cause and effect. The three-block movements in Finding 4 are directional, not significant.
  • The four endpoints are not matched on model size, and this is the study’s biggest weakness. claude-opus-4-8 is a frontier model and gpt-4o-mini-2024-07-18 is a small one with an older training cutoff. Any comparison between those two columns therefore confounds vendor with model tier and cutoff date, and we cannot separate them with this data. We chose these endpoints because they are what the product queries by default, not to make a point. Read the per-assistant columns as “this endpoint, as configured” — never as “OpenAI’s position.”
  • Six tracked brands are pre-registered; the 39-product set is not. The six were fixed before collection. The 39 CRMs came out of the answers themselves and were normalised afterwards, by us, by hand. We have described the merges and would defend them, but a second analyst would not produce a byte-identical list.
  • One category. CRM for small business is a mature, heavily-reviewed category. A newer or more technical one would very likely behave differently.
  • The share-of-voice table covers six brands, not the whole market. That list was fixed before collection started. The other 33 CRMs appear in the divergence figures and in Finding 1’s long-tail table, but not in the share-of-voice table.
  • Brand normalisation involves judgment. We have described exactly what we merged and what we did not. Reasonable people would draw one or two lines differently.
  • Sentiment and “explicitly recommended” are model-extracted, not hand-coded. We spot-checked them and quote the raw text above so you can see the framing yourself, but they are not hand-verified at scale.
  • API access, not consumer apps. People using ChatGPT’s web interface with browsing enabled may see different answers than the API returns.
  • Three of the four endpoint ids are aliases, not dated snapshots. Only gpt-4o-mini-2024-07-18 names a fixed build. claude-opus-4-8, gemini-3.5-flash and sonar can be repointed by their providers mid-window without notice, and we did not log the snapshot each one resolved to. The request was constant for 18 days; the model behind it is asserted, not proven.
  • We build in this category. Agonai was the instrument. We chose a category we do not compete in and we are publishing the method so you do not have to take our word for it.

Reproduce this

Everything needed to repeat the study is above: the 10 questions, the 6 brands, the window, the providers, the frequency, and the two counting decisions that shape the numbers. You can run it by hand against four chat interfaces — tedious but free — or with any tool that schedules the queries and stores the answers, which is what Agonai does.

If you want the same thing for your own category and brand, that is the product: start a trial and point it at your buyers’ questions instead of ours. If you are coming at this from traditional competitive intelligence, we also compare directly against Crayon, Klue and Kompyte — none of which track this.

Related reading: how to see what ChatGPT says about your brand is the manual version of what we automated here.


Data collected 2026-08-01 → 2026-08-18, 720 answers, portfolio “Study: CRM SMB”. Last verified: 2026-08-19.

See what your competitors just changed

Published pricing from $19/mo. Start a 14-day free trial — no credit card.

Start free trial