AI & Machine Learning

AI Search Has No Single Ranking: Why Different People See Different Brands

Leo Wang August 31, 2026
AI Search Has No Single Ranking: Why Different People See Different Brands

AI Search Has No Single Ranking: Why Different People See Different Brands

You and a colleague send the same product question to the same AI assistant. Your answer recommends Brand A first. Your colleague sees Brand C at the top. You ask again the next day, and the shortlist changes.

The system may be working as designed. Two screenshots alone also cannot prove that an AI engine personalized a brand ranking for each person. Brand results in AI search depend on several groups of conditions: the wording and current conversation, any account context the product is allowed to use, the platform and mode, and the retrieval and generation path for that run. A change in any one of them can alter the brands, sources, and order in the answer.

A more accurate model is this: brand visibility in AI search is a conditional probability distribution. No single fixed list represents every user, account, mode, and point in time.

---

The Short Answer

“AI search has no single ranking” means there is no fixed brand leaderboard that represents every user, account, mode, and time. An individual response can still contain an order, a shortlist, and a top recommendation. That local order applies to the conditions of that response.

Four boundaries matter from the start:

  1. Different answers alone cannot establish personalization. Probabilistic generation, retrieval-source changes, and web updates can create run-to-run variation even when the prompt, account, location, and mode stay constant.
  2. Login status, subscription tier, and enabled memory are separate variables. Account access, history, settings, plan, and product mode require separate records.
  3. API output and consumer-product output describe different environments. An API can expose search calls or citation metadata. Those fields document the API environment; they do not provide a request-level audit log for a ChatGPT, Gemini, or Google AI Mode consumer session [6].
  4. One screenshot cannot establish where a brand “ranks.” A 2026 information-retrieval preprint frames the central method clearly: repeat the measurement, describe visibility as a distribution, and avoid point-result conclusions [9].

In one sentence: Traditional rank answers where a page sits on a result list; AI visibility asks under which conditions a brand enters the answer, how often it appears, and what role it receives.

---

Why Can the Same Question Produce Two Brand Lists?

Suppose two people enter this exact prompt:

Which CRM is the best fit for a 50-person cross-border ecommerce team expanding into the U.S. market?

The visible sentence is identical. The full request state may differ.

ConditionUser AUser BWhat it may affect
Current conversationEarlier messages mention Shopify, Chinese-language support, and a monthly budgetEarlier messages mention a Salesforce migration and enterprise permissionsCandidate products, comparison criteria, and evidence needs
Account contextThe product can use a saved “small team, cost-sensitive” detailNo available memory, or memory is offRecommendation rationale and exclusion criteria
LocationSearch uses an approximate U.S. locationSearch uses an approximate U.K. locationProduct availability, price, regulation, and local sources
Product modeFast-response modeReasoning or research modeSearch volume, source types, and citation range
Individual runThe system retrieves source set AThe system retrieves source set BBrands and cited pages that enter the response

This illustrative scenario is based on public product mechanisms; no real-user internal logs were used. Its purpose is to show that identical text in the input box does not establish an identical request state.

OpenAI says that, when the relevant features are enabled, responses can use context such as saved memories, past chats, files, and custom instructions. Users can turn memory off or delete it [1]. ChatGPT Search may also infer an approximate location from an IP address. Precise device location is off by default and requires permission. When memory is enabled, relevant memories may inform search-query rewriting [3].

Google describes a similar capability through personal context in AI Mode. Where the feature is available and the user chooses to use it, AI Mode can draw on past searches and can connect to Gmail with the user's permission. The interface indicates when personal context is in use, and the connection can be removed [4].

These official materials establish that specific products can use personal context under specific settings. They do not publish a universal personalization weight, and they do not show that every brand change comes from a user profile.

---

“No Single Ranking” Still Allows Order Within One Answer

An AI answer can have an order. A model may name one product as the best fit and list two alternatives, or place candidates in a table. Readers notice that order, so it deserves measurement.

The mistake is extending one response order into a global leaderboard.

ConceptWhat it measuresSupported conclusionUnsupported conclusion
Order in one answerThe sequence of brands in one responseA brand's position in that responseThe brand has the same position for every user
Top-pick shareThe percentage of repeated tests where a brand is recommended firstHow often the brand becomes the top pick under defined conditionsThe platform maintains a public, fixed master list
Mention rateThe percentage of repeated tests that include a brandConditional visibilityEvery mention is a recommendation or citation
Visibility distributionMentions, order, roles, and sources across conditions and runsStability, range, and differencesA guarantee for one person's next answer

A useful measurement expression is:

Probability of a brand appearing = P(brand mentioned | prompt, conversation, account settings, location, platform, mode, time, run)

This expression is a measurement framework. No AI platform has published it as an operating formula. It forces a team to state which product, which conditions, which time window, and how many runs before reporting where a brand appears.

---

Four Layers Shape Brand Results

Same-prompt differences can be organized into four layers. The first two define the task and user conditions available to the system. The next two determine which capabilities and evidence paths the product uses.

LayerMain variablesDiagnostic questionEvidence boundary
Request layerOriginal prompt, wording, language, current conversation, uploaded filesDid both tests receive the same full conversational context?The conversation is part of the input; public descriptions do not expose the complete internal prompt
User/account layerMemory, history, settings, connected services, available locationWas the product allowed to use these details?Applies only where the product supports the feature, it is available, and settings allow it [1][4]
Product layerPlatform, model, consumer app or API, mode, subscription-enabled capabilitiesWere both people using the same product configuration?Google confirms that different AI features can use different models and techniques [5]
Run/retrieval layerProbabilistic generation, whether search triggers, source routing, web changes, timeDoes the result vary when every known condition stays fixed?Third-party repeated tests observe routing and source-set variation at specific points in time [8]

Request Layer: Identical Final Text Can Represent Different Tasks

If a user has just discussed budget, industry, country, or excluded options, the next message—“Which one do you recommend?”—depends on that history. A clean conversation, a continuing conversation, an attached file, and another language create different experimental conditions.

Copying only the final sentence from another person's input box therefore does not reproduce the full request. A personalization comparison first needs consistent current-conversation context. Otherwise, the test primarily measures a conversation difference.

User/Account Layer: Memory and Personal Context Are Optional Capabilities

OpenAI's Memory FAQ says personalization can draw from saved memories, past chats, custom instructions, and user-provided files. Users can inspect, delete, or disable relevant features [1]. These sources are user-level context. They differ from general knowledge learned in model parameters during training.

Temporary Chat shows why login status alone provides an incomplete label. OpenAI's current documentation distinguishes a default nonpersonalized Temporary Chat from a personalized Temporary Chat. The former does not use memory, custom instructions, or plugins. The latter can use existing memories but does not create new memories from that temporary conversation [2]. “Temporary,” “logged in,” and “has history” each describe only part of the configuration.

Personal context in Google AI Mode also has prerequisites involving feature availability, location, user choice, and connection status [4]. A rigorous test records those controls. It does not put every signed-in user into one undifferentiated group.

Product Layer: Platform, Mode, and Plan Can Change the Source Pool

Google explicitly states that AI Mode and AI Overviews may use different models and techniques. Different answers and links across those experiences are therefore expected [5]. Cross-platform results require separate reporting as well.

A 2026 preprint under review tested GPT-5.2, Gemini 3 Flash, and Perplexity sonar-pro. It used 250 brand-free category questions, repeated each question five times on each model, and collected 3,750 responses across five industries and 50 brands. In that sample, cross-model agreement on the top-recommended brand was 41.6% [11]. This exploratory study does not establish a benchmark for every industry. It supports a narrower conclusion: first place on one model does not transfer reliably to another model.

Mode can create a large difference too. Semrush and Kevin Indig ran 100 prompts through GPT-5.2 under minimal- and high-reasoning conditions, producing 200 responses. Only 25.6% of cited domains overlapped between the two conditions for the same prompts [10]. This vendor study ran each prompt once per mode. It illustrates that mode can change source structure; its percentage should not be treated as a stable long-term rate.

Login status, subscription tier, and response mode need separate fields. Login may make account context available. A subscription may unlock capabilities. The selected mode determines which tools the request uses. A single “paid-user answer” label cannot explain the differences among them.

Run/Retrieval Layer: Results Can Vary While Known Conditions Stay Fixed

Chris Green tested 1,000 prompts up to 10 times and recorded 9,946 completed ChatGPT search runs. 11.6% of prompts switched primary retrieval source across repeated runs. When the source switched, URL-set overlap fell from 0.273 to 0.149, while domain-set overlap fell from 0.265 to 0.155 [8].

These findings came from reverse engineering of product network traffic at that time. They are not OpenAI routing specifications. They show that retrieval-path changes can alter the candidate sources even without a difference in personal information. Web updates, index changes, model refreshes, and time-sensitive facts add further variation.

Attributing every different response to “the AI remembered me” ignores the product and run layers. Labeling all differences as random can hide repeatable effects from account context, location, or mode.

---

How Can You Separate Personalization From Run-to-Run Variation?

Personalization is a systematic result change associated with a user condition. Run-to-run variation occurs while the known conditions remain constant. Both effects can occur together.

ObservationReasonable initial explanationNext validation step
Repeated runs under one configuration rotate brand listsRun or retrieval variationKeep every condition fixed, add repeats, and record sources
One type of brand repeatedly rises only after synthetic memory is enabledPossible memory-condition effectRun an on/off comparison with the same account, conversation, and mode
Local brands and prices change systematically with locationPossible location effectFix language, account, mode, and time window; change location only
Switching from fast mode to reasoning mode changes many cited domainsProduct-mode effectUse the same prompt set and the same number of repeats in both modes
ChatGPT and Gemini choose different top brandsPlatform-level differenceReport platforms separately; do not label it personal context
Every group changes after several weeksTime, model, index, or web updatePreserve version, date, and source snapshots; establish a new baseline

A single A/B pair can still confuse random variation with a condition effect. Run the same number of repeats under each condition and compare the resulting distributions. A difference that appears once or twice provides weak attribution evidence. A difference that persists across repeated, controlled runs justifies further testing.

---

What Is Confirmed, and What Remains Unconfirmed?

Separating official capability documentation from independent observation prevents “the product can do this” from becoming “every response does this.”

Supported statementStatement the available evidence cannot support
ChatGPT can use memories, past chats, files, and custom instructions when the relevant settings are enabled [1]Every brand recommendation reads the full account history, or each user has a fixed personal brand ranking
ChatGPT Search can infer approximate location from IP; precise location needs permission; relevant memory may inform query rewriting [3]The platform sends the full IP or account identity to search providers, or every location difference is personalization
Google AI Mode can use past searches under defined conditions and lets users choose whether to connect Gmail [4]Personal context is active for every location, account, and question, or it always receives a fixed weight
AI Mode and AI Overviews can use different models and techniques [5]The two experiences should return the same brand list
Repeated studies observe changes in retrieval routing, cited domains, and top-recommended brands [8][10][11]Percentages from those samples are permanent product specifications or predictions for every industry
The Gemini API can return search calls, queries, and citation metadata [6]The API mirrors consumer-app requests or reconstructs a real user's private context

None of the controlled studies cited here isolates signed-in versus signed-out status as the only variable and demonstrates that login alone changes the brands. Authentication belongs in the test matrix, but memory, history, location, subscription, and mode need separate records.

---

Why Does a Traditional Rank Tracker Distort This Picture?

A traditional rank tracker repeatedly queries a relatively standardized search interface and records a URL's position in a result list. That method remains useful for web search. A direct copy of it for AI answers faces four structural problems.

  1. The complete candidate list is unavailable. An answer shows only the brands selected for the narrative. Absence from the response does not prove absence from internal retrieval.
  2. One run is one sample. Repeating the same configuration can change sources and brands, so one position gives variation a false appearance of precision [8][9].
  3. Standardized testing deliberately removes personal conditions. That creates a comparable baseline, but it does not represent an account with memory, history, or location permission.
  4. Real-account testing can become hard to explain. When account history, plan, and mode all differ, the observed difference has no isolated cause.

Native platform reporting does not fill the entire gap. Google's documentation for the Search Console generative AI performance report lists page, country, date, and device dimensions. The current documentation does not list user account, memory state, or a complete personalization path [7]. Aggregated reporting can show exposure trends for a site. It cannot explain why one person saw Brand A.

A standardized AI tracker therefore measures visibility in a defined test environment. That baseline remains valuable when the report states platform, location, language, mode, account conditions, time window, and repeat count. Calling it a universal ranking creates the distortion.

---

How Do You Design a Test That Can Explain the Difference?

The method has two core rules: change one variable at a time, and repeat every condition.

Establish the Baseline First

At minimum, record:

  • The full prompt, language, and punctuation
  • Whether the conversation is new or continuing, including prior test inputs
  • Platform, consumer product or API, model, and mode
  • Login status, subscription tier, memory, and history settings
  • Approximate location or an authorized test location
  • Test date, time window, and any visible product version
  • A repeat count defined before the test starts
  • Full response, brand order, citation URLs, and whether search occurred

Build a One-Variable Test Matrix

Variable to testControlComparisonPrimary question answered
Current conversationNew, empty conversationAdd a fixed synthetic business backgroundDoes conversation context alter candidate brands?
MemorySame test account, memory offSame account, only a predefined synthetic memory enabledDoes available memory create a persistent difference?
LocationFixed location AFixed location BDoes region change local brands, prices, or sources?
LoginComparable signed-out environmentSigned-in account with other controllable settings held steadyDoes authentication and its attached capability set require further isolation?
Subscription tierA shared mode on tier AThe same shared mode on tier BDoes the plan alter available capabilities? If the mode differs, plan cannot receive sole attribution
Response modeFast or default modeReasoning or research modeDo source and brand sets change with mode?
PlatformChatGPTGemini or another consumer productHow much do final results differ by product?
TimeBaseline in a narrow windowA predefined later windowDo model, index, and web changes shift the distribution?

“Platform A versus Platform B” is a product comparison. It cannot serve as a clean single-cause experiment because the model, retrieval system, and interface all change together. It can test whether results transfer across products, but it cannot identify the internal component responsible for the difference.

Complete the Audit in Six Steps

  1. Define the business question. Choose a question real users ask and one that can influence brand choice. Avoid beginning with a branded prompt.
  2. Freeze the prompt and recording template. Set fields, repeat count, and end time before looking at results.
  3. Build a clean baseline. Use new conversations and documented settings to measure run variation first.
  4. Change one condition at a time. Give memory, location, mode, tier, and platform separate groups.
  5. Save the full result. Record mentions, order, recommendation language, and citation sources. Keeping only the “winner” removes the explanation.
  6. Compare distributions. Report numerator, denominator, time window, and condition differences. Avoid drawing conclusions from one screenshot.

---

Which Metrics Can Replace a Universal Position?

The useful metrics are conditional and expose uncertainty.

MetricCalculation unitQuestion it answersCommon misreading
Mention rateRuns that mention the brand ÷ total runsHow often does the brand appear under this condition?Treating every appearance as an endorsement
Top-pick shareRuns that recommend the brand first ÷ total runsHow often does the brand become the top choice?Calling a within-condition share the platform's master ranking
Recommendation-role distributionShare of top pick, alternative, poor fit, and comparison-only rolesIn what role does the brand appear?Counting the name while ignoring sentiment and function
Citation rateShare of runs containing a relevant owned or third-party sourceWhich answers use verifiable sources?Equating a citation with a visible brand mention
Brand-set overlapIntersection and union of brand sets across two groupsHow much does the shortlist change with the condition?Using a single pair as a long-term stability estimate
Conditional differenceMention rate under condition B minus condition AWhat direction is associated with the controlled condition?Claiming causation without repeated runs
Time variationDistribution change across time windowsIs the result drifting?Treating a product update as a content outcome

A report can replace “Brand A ranks first” with a statement such as:

In August 2026, under the specified consumer product, U.S. location, English-language new conversation, and default mode, Brand A appeared in X of N runs and was the top pick in Y. With the documented enterprise background added, its top-pick share became Z of N.

N, X, Y, and Z must come from a real test. A template should preserve empty fields until data exists; filling them with invented values would create false precision.

---

Where Should the Privacy Boundary Sit?

Personalization research can drift into collecting people's private histories. A valid study does not require that data.

  • Prefer synthetic backgrounds. Prewrite fictional conditions such as “U.S. small business, budget-sensitive, uses Shopify” so the setup can be reproduced.
  • Obtain explicit consent for real accounts. Participants should know which settings and outputs are recorded, and they should be able to withdraw.
  • Do not collect private conversation text. Record only the controls, synthetic conditions, and aggregated outcomes required by the experiment.
  • Keep personal memories out of the data request. Saved memory may contain work, health, family, or identity details that have no place in routine marketing monitoring.
  • Aggregate results by group. Report condition effects without building identifiable recommendation profiles.
  • Remove credentials and permissions. Delete temporary accounts, location access, service connections, and exported files after the test.

These limits also improve experimental quality. Years of uncontrolled history in a real account add confounding variables. A synthetic history is cleaner and makes the changed condition easier to identify.

---

What Does Innflows Address Here?

Same-prompt variation creates an operational problem for brand teams: results require repeated observation by condition, platform, and time. One screenshot cannot complete that work.

Innflows can support external query simulation and grouped monitoring. A team can maintain a stable set of business questions, organize tests by platform, language, location, and other publicly controllable conditions, and record brand mentions, recommendation roles, and citation sources over time. Within this article's framework, that creates a “question–test condition–repeated run–brand distribution–citation source” baseline and helps identify which observable layer contains the difference.

The boundary is explicit: Innflows cannot read a user's private memories, complete chat history, or an AI platform's internal personalization profile. It also cannot reconstruct the full request state of a real person. External monitoring describes the distribution within selected test environments. It cannot guarantee which brand any individual will see next. Controlled tests involving memory settings or synthetic account history should use separately authorized test accounts.

---

Five Common Misconceptions

Misconception 1: AI Has No Ordering at All

One response can still contain sequence, a top pick, and alternatives. What is missing is a leaderboard that stays fixed across users, modes, and time. Measure top-pick share and role distribution; order still carries meaning within a response.

Misconception 2: Two Different Answers Prove Personalization

A different conversation, product mode, retrieval-source switch, or probabilistic generation can also change the answer. A user-condition effect becomes plausible only after other conditions are controlled, runs are repeated, and the difference persists with that variable [8][9].

Misconception 3: Signed-In Users Form One Group and Signed-Out Users Form Another

Signing in may bring history, memory, settings, subscription capabilities, and a different default mode together. Without separate records, a group difference has no clear explanation. The sources cited here do not establish login alone as the cause of a brand-result change.

Misconception 4: Enough API Runs Will Predict the Consumer App

An API is a separate environment suited to automation and proxy research. Google Search grounding in the Gemini API can return search calls and citation metadata [6]. Those fields are not per-request logs for Google AI Mode or the Gemini consumer app. Report API results as their own surface.

Misconception 5: One Average AI Visibility Score Is Enough

Cross-model agreement on the top brand was only 41.6% in one preprint sample [11]. A separate software index collected on August 6, 2026 found different first choices from ChatGPT and Gemini in about one-third of software categories. It included 9,978 usable consumer-app answers, but had no error bars or peer review [12]. A blended platform average can hide the most important disagreement.

---

When Do These Conclusions Have Limits?

Official documentation describes capabilities and controls. It does not report usage rates. OpenAI and Google confirm memory, location, and personal-context capabilities. They do not publish trigger rates, weights, or complete decision logic for every type of brand question [1][3][4].

Controlled evidence has not isolated login as the sole cause of a brand change. Signed-in and signed-out states belong in a test matrix; causal attribution must wait for a controlled result.

Every independent study has a sample boundary. The source-routing study reverse-engineered ChatGPT network traffic at a particular time. The reasoning-mode study came from a tool vendor and ran 100 prompts once in each mode. The cross-model study remains a preprint under review and covers five industries and 50 brands [8][10][11].

“First” is a coarse label. Some answers identify one top choice; others recommend several products by use case. Extracting the first brand mechanically loses meaning, so recommendation role should be recorded too.

External testing cannot expose every internal condition. Consumer products do not reveal complete system prompts, routing, candidate sources, or personalization profiles to monitors. Every measurement needs an observation window.

The conclusions are time-bound. This article uses official materials and studies available through August 21, 2026. Product names, feature availability, models, and default modes can change. Each third-party percentage belongs to its own sample and collection period.

---

Frequently Asked Questions

Will Two People Always Get Different Answers to the Same Prompt?

No. They may receive the same brands, or their answers may differ because of conversation context, account context, location, product mode, or retrieval on that run. The conclusion is that no universal leaderboard can guarantee identical results for everyone. It does not predict a difference in every pair of responses.

Does Signing In to ChatGPT or Google Guarantee Personalized Brand Recommendations?

No. Personalization depends on whether the specific product supports the capability, whether the feature is available, whether the user enables or connects the context, and whether the current task uses it [1][4]. Login status is one condition to record.

How Can I Test Whether a Memory Changed the Brand Result?

Use the same authorized test account and mode. First repeat the prompt with memory off. Then add only a predefined synthetic memory and repeat the same number of runs. Compare mention rate, top-pick share, recommendation role, and citation set. One run with memory off and one with it on cannot rule out normal run variation.

Is the First Brand in an Answer Ranked Number One by the AI?

It is the first brand or top pick in that response. If repeated tests under the same conditions consistently place it first, report its top-pick share for those conditions. That evidence cannot establish first place across the entire platform and every user.

Can an API Replace Consumer-Product Testing?

An API supports scaled experiments and analysis of the search metadata it returns. It cannot replace consumer-product testing because interface, system instructions, tools, account context, and modes may differ. Keep the two result sets separate [6].

Can GEO Make Every User See the Same Brand?

No. GEO can improve content accessibility, evidence quality, topic coverage, and visibility within selected test environments. It cannot control private user context, platform models, retrieval routing, or probabilistic generation. A defensible goal is to improve the probability of stable appearance across important conditions and report the observed variation honestly.

---

Core Takeaways

Four layers shape brand results in AI search. The request layer defines the task. The user/account layer determines which permitted personal context is available. The product layer defines platform, model, mode, and tools. The run/retrieval layer determines which sources enter this response and how the answer is composed.

The question “Where does our brand rank in AI?” therefore lacks essential conditions. More answerable questions are:

  • Which platform, mode, location, and time window?
  • Was the test a new conversation or one with documented context?
  • What were the memory, history, and connected-service settings?
  • Across how many runs, what were the mention rate, top-pick share, and role distribution?
  • Which difference persisted with one controlled variable, and how much variation remained under identical conditions?

Moving from one screenshot to a condition matrix will not make AI search fully predictable. It will show the team what was actually measured, and it will prevent one response from a standardized account from being described as the ranking every user sees.

AI search still selects brands, but each selection belongs to a set of conditions. The durable goal is stronger probability of entering the shortlist as important conditions change.

---

References

[1] - Memory FAQ — OpenAI Help Center

[2] - Temporary Chat FAQ — OpenAI Help Center

[3] - ChatGPT Search — OpenAI Help Center

[4] - AI in Search: Going beyond information to intelligence — Google

[5] - AI Features and Your Website — Google Search Central

[6] - Grounding with Google Search — Google AI for Developers

[7] - Generative AI performance report (Search) — Google Search Console Help

[8] - ChatGPT citations change when hidden search pipelines switch — Search Engine Land, 1,000 prompts and 9,946 completed runs

[9] - Don't Measure Once: Measuring Visibility in AI Search (GEO) — arXiv:2604.07585, preprint submitted April 2026

[10] - Only 25% of cited sources overlap between ChatGPT's different reasoning modes — Semrush and Kevin Indig, 100 prompts and 200 responses

[11] - Who Owns the AI Recommendation? A Multi-Industry Empirical Map of Brand Category Ownership Across Large Language Models — arXiv:2606.23057, preprint submitted June 2026 and under review

[12] - ChatGPT and Gemini disagree on the best software in a third of categories — The Next Web on the GetIntel AI Software Index, data collected August 6, 2026

Related Articles