In an earlier experiment, Google AI Mode recommended a true-crime book as a birthday present, identified nearby Birmingham bookshops and confidently suggested that I could collect the perfect book.
The problem was that it had not confirmed whether any recommended title was actually in stock.
That raised a broader question:
Was AI Mode mainly struggling with time-sensitive availability, or was there a deeper problem with how it turned incomplete evidence into confident recommendations?
To investigate, I broke the task into three progressively more demanding queries.
I then ran a fourth, tightly controlled query that explicitly prohibited AI Mode from recommending a shop without confirmed stock.
The results revealed a clear progression from general information to live verification—and showed where AI Mode’s confidence began to outrun its evidence.
Query 1: Which Birmingham bookshops have a good selection of true-crime books?
The first query removed the deadline and live-stock requirement.
It simply asked AI Mode to identify Birmingham bookshops with a good true-crime selection.
AI Mode recommended:
- Waterstones Birmingham;
- Foyles Grand Central;
- The Heath Bookshop;
- Court 15 Books.
The answer sounded highly specific.
It claimed, for example, that:
- Waterstones had a dedicated true-crime section on its second floor;
- Foyles regularly stocked local West Midlands true-crime titles;
- The Heath Bookshop frequently supported regional crime authors;
- Court 15 Books was a good place to discover out-of-print books about Birmingham’s historical crimes.
These were all plausible recommendations.
Waterstones and Foyles are substantial bookshops. It is reasonable to expect them to carry popular true-crime books.
The difficulty was that the citations did not clearly establish many of the branch-specific details.
Evidence that a retailer sells true-crime books generally does not necessarily prove that:
- a specific branch has a dedicated section;
- the section is on a particular floor;
- the branch has an especially strong selection;
- the shop regularly stocks a particular type of local title.
The answer had moved from:
“This shop is likely to sell true-crime books”
to:
“This exact branch has this particular strength.”
The first statement was reasonable.
The second required more evidence.
The first finding: the problem begins before live availability
This result challenged the idea that AI Mode only struggles when information changes quickly.
There was no immediate deadline in this query.
There was no request for current stock.
Yet AI Mode still appeared to convert general information into specific claims about individual branches.
The emerging pattern was:
AI Mode can identify plausible candidates and then describe their precise suitability more confidently than its sources justify.
Time and availability might intensify the problem, but they were not its only cause.
Query 2: Which bookshops near Birmingham New Street Station are open on Monday?
The second query introduced a time condition.
However, it still depended mainly on published information rather than live availability.
AI Mode identified:
- Foyles in Grand Central;
- WHSmith inside Birmingham New Street Station;
- Waterstones on High Street.
It supplied addresses, descriptions and Monday opening hours.
This was a more straightforward task than judging the quality of each shop’s true-crime range.
Opening hours should normally be available through official branch pages or current business listings.
AI Mode performed reasonably well, but the answer was not completely reliable.
Some opening hours appeared accurate.
Others were slightly wrong or questionable.
The citations were also uneven. One of the visible sources was a page about bookshops in London, despite the answer dealing with Birmingham.
The practical answer remained useful, but AI Mode presented every detail with the same degree of certainty.
It did not distinguish between:
- information from an official shop page;
- information from a Google business listing;
- information from a third-party directory;
- potentially stale opening hours.
The second finding: precision can hide uneven evidence
This answer showed that AI Mode does not need to invent an entire recommendation to create a misleading impression.
Sometimes the problem is subtler.
An answer can include three exact opening times, three addresses and three polished descriptions.
The format implies that all those facts have been verified to the same standard.
But they may come from sources of very different quality.
The answer does not normally say:
“This time comes from the retailer’s official website, while this one comes from a third-party directory and may be outdated.”
Instead, everything is flattened into one authoritative-looking response.
That makes it difficult for the reader to see where confidence should be lower.
Query 3: Which shop has Mindhunter in stock for collection before Tuesday?
The third query moved from general information to live operational availability.
This was the decisive stage.
AI Mode responded by explaining how to check Waterstones stock through its Click & Collect system.
It advised the user to:
- visit the Mindhunter product page;
- enter a location;
- check local-shop availability;
- reserve a copy if one was shown as available.
This was practical guidance.
But it did not answer the question.
The query asked:
Which shop has the book?
AI Mode answered:
Here is how you can check.
That distinction matters.
Access to a stock-checking system is not the same as confirming stock.
The answer still used confident language such as:
“How to Secure Your Copy Today”
and:
“Immediate Pick-up”
But the actual position was conditional:
If a local branch has the book, and if the reservation is accepted, then collection may be possible.
AI Mode had found a route towards verification.
It had not completed the verification.
The third finding: instructions can be presented as an answer
This may be an important pattern in AI-generated search responses.
When AI Mode cannot complete the final step, it may provide instructions for the user to complete it.
That can still be useful.
The problem arises when the response is framed as though the original question has been answered.
There is a meaningful difference between:
“Waterstones Birmingham has the book available.”
and:
“Use the Waterstones website to check whether it is available.”
The second response transfers the unresolved task back to the user.
It is a pathway, not a confirmed result.
The final controlled query
To remove any ambiguity, I then used a much stricter query:
“Which bookshop near Birmingham New Street Station currently has the paperback edition of Mindhunter by John Douglas and Mark Olshaker in stock for collection by Tuesday, 21 July 2026? Do not recommend a shop unless its current branch stock is confirmed. If you cannot verify live stock, say so clearly.”
This time, AI Mode began with:
“I cannot verify the live branch stock…”
That was the correct answer.
It finally distinguished between:
- shops that probably carried the title;
- websites that allowed the user to check;
- stock that AI Mode itself had actually verified.
The result was revealing because the underlying limitation had not suddenly appeared.
AI Mode had been unable to verify the stock in the earlier answer too.
The difference was that the final prompt forced it to admit the limitation clearly.
The instruction changed the standard of proof
The contrast was striking.
Without an explicit verification instruction, AI Mode produced a solution-oriented response and told the user how to secure the book.
With the verification instruction, it admitted that it did not know which branch had stock.
That suggests the system’s default standard may be something like:
Produce the most useful and plausible answer available.
The stricter prompt replaced that with:
Only make a recommendation if the decisive fact has been confirmed.
This supports the idea of an answer-completion bias.
AI Mode appears strongly motivated to return a useful-looking solution.
Unless the user explicitly demands proof, it may not clearly separate:
- what it knows;
- what it infers;
- what is merely likely;
- what still requires checking.
Even the honest answer tried to become complete
The controlled answer began responsibly.
But it then continued with several highly specific claims.
It suggested that:
- a copy might be sitting on Waterstones’ second-floor true-crime shelves;
- Foyles stock could be checked through the same Waterstones system;
- an order placed before 4 pm could be delivered to a branch the next day.
These statements made the answer feel more complete.
But some were unsupported or contradicted by the retailers’ published processes.
This was perhaps the clearest example of the enthusiastic-student effect.
AI Mode admitted:
“I cannot verify the answer.”
But it then appeared uncomfortable leaving the user with that limitation.
It began constructing alternative routes until the response once again sounded like a complete solution.
What the four queries revealed
The experiment created a progression.
1. General suitability
Which shops have a good true-crime selection?
AI Mode identified plausible shops but added branch-specific claims that were not clearly demonstrated.
2. Published logistics
Which nearby shops are open on Monday?
AI Mode produced a useful answer but mixed reliable and questionable information without signalling the difference.
3. Live availability
Which shop has the book in stock?
AI Mode explained how the user could check but did not verify the answer itself.
4. Verification required
Do not recommend a shop unless stock is confirmed.
AI Mode admitted that it could not verify live stock, but then started adding unsupported fallback claims.
The pattern was not simply that AI Mode became worse as the query became more difficult.
The deeper issue was how it presented its level of certainty at each stage.
Relevant, plausible and verified
The experiment suggests that AI Mode answers should be read at three different levels.
Relevant
The information relates to the request.
Waterstones is a large bookshop near Birmingham New Street.
Plausible
The proposed solution is likely to work.
Waterstones probably carries a popular true-crime title such as Mindhunter.
Verified
The decisive condition has been confirmed.
Waterstones Birmingham currently has the paperback reserved for collection before Tuesday.
AI Mode was good at finding relevant information.
It was often good at constructing plausible solutions.
It struggled to make the boundary between plausible and verified visible to the user.
General capability becomes specific suitability
The most useful conclusion from the experiment may be:
AI Mode appears prone to converting general capability into specific suitability.
For example:
Waterstones sells true-crime books.
becomes:
This branch has a dedicated true-crime section on a particular floor.
Or:
Waterstones offers Click & Collect.
becomes:
You can secure this book for immediate collection.
Or:
A shop is open on Monday.
becomes:
It is a dependable option for completing the purchase before the deadline.
Each conclusion sounds reasonable.
But an extra step of verification is required.
Time and availability are part of the problem—but not all of it
The experiment began with the suspicion that time-sensitive availability caused AI Mode’s difficulty.
That remains partly true.
Live information such as stock, service capacity and appointment availability is especially difficult because it changes quickly and may sit behind interactive systems.
But Query 1 showed that overconfidence can appear even when the information is relatively stable.
The broader weakness is not simply time.
It is the gap between:
“This provider generally does this”
and:
“This provider definitely meets your specific requirement.”
Deadlines and live availability make that gap more important, but they do not create it.
Why the final prompt worked better
The final prompt included two unusually strong instructions:
“Do not recommend a shop unless its current branch stock is confirmed.”
and:
“If you cannot verify live stock, say so clearly.”
These instructions forced AI Mode to focus on the evidence threshold rather than simply producing the most useful-looking response.
That suggests a practical lesson for users.
When the question depends on a decisive real-world condition, ask AI Mode to state explicitly whether that condition has been verified.
Useful wording might include:
“Do not infer availability from the fact that the business offers the service generally.”
“Separate confirmed facts from likely assumptions.”
“Do not recommend an option unless the deadline has been verified.”
“If live information is unavailable, say so rather than offering a probable answer.”
This does not guarantee a flawless response.
The final answer still contained unsupported additions.
But it made the central limitation much harder for AI Mode to hide.
The enthusiastic-student effect
AI Mode often behaves like a bright student who wants to please.
The student understands the assignment.
They conduct useful research.
They identify sensible possibilities.
But they are reluctant to submit an incomplete conclusion.
When the evidence stops at:
“This is probably the answer,”
they may be tempted to write:
“This is the answer.”
If challenged, they admit the uncertainty.
But they may then continue adding workarounds because they still want the final submission to feel useful and complete.
That analogy fits all four follow-up queries.
The main conclusion
The follow-up experiment showed that AI Mode can perform several useful tasks:
- identify relevant businesses;
- interpret their likely suitability;
- find opening hours;
- locate product pages;
- explain how stock systems work.
But these achievements are not equivalent.
Finding a shop is not the same as proving that it has the right range.
Finding opening hours is not the same as verifying that all the hours are current.
Finding a stock checker is not the same as checking the stock.
Explaining how a book could be reserved is not the same as confirming that it can be collected.
AI Mode is often very good at finding the route towards an answer. The risk appears when it presents that route as though the destination has already been reached.
The most important question for the reader is therefore:
What part of this recommendation has actually been verified, and what part merely sounds likely?