AI Models Neutral 5

3 Chatbots, 2 Cities, 1 Lesson: Better Prompts Fix AI Trip Planning

A controlled test of ChatGPT, Gemini and Claude shows prompt quality, not model capability, is the binding constraint in AI travel planning. Paid tiers ran deeper and faster with better stamina, while outputs skewed bland and vulnerable to generative engine optimisation. The finding reframes 'AI can't do X' complaints as prompt-engineering and product-design problems.

· 5 min read · Verified by 2 sources ·

Beat this week

Last 7 days · AI Models

15 stories
7 avg impact
0% positive
73% negative
vs prior 7 days +2 +2 stories vs prior 7 days

Impact 7.0/10 (+0.3 vs prior). Counts are stories in our record, not a market forecast.

Open the change report

Coverage balance Negative coverage leads. Negative coverage exceeds positive coverage by 73 percentage points.

  • 27% neutral
  • 73% negative

This story sits in AI Models — the counts compare this beat's last 7 days with the previous 7 in our verified record, not a market forecast.

Figures are computed live from our source-verified story record (as of ) The volume change compares this window with the prior 7 days in the same record. — see our methodology for how impact and sentiment are derived.

AI briefing

Key takeaways

5 impact
Neutralsentiment
2sources
5min read
  1. A controlled test of ChatGPT, Gemini and Claude shows prompt quality, not model capability, is the binding constraint in AI travel planning.
  2. Paid tiers ran deeper and faster with better stamina, while outputs skewed bland and vulnerable to generative engine optimisation.
  3. The finding reframes 'AI can't do X' complaints as prompt-engineering and product-design problems.
Drawn from
  • theage.com.au
  • brisbanetimes.com.au

In this briefing

Mentioned

Key Intelligence

Key Facts

  1. 1A journalist tested three chatbots — ChatGPT, Gemini and Claude — in both free and paid tiers on two fictional five-day trips (São Paulo for a first-timer, New York for a repeat visitor).
  2. 2Testing used fresh accounts and 'temporary chats' so the models could not infer the writer had lived in both cities for years.
  3. 3Wharton professor Ethan Mollick: 'They're not oracles... Think of it like working with a smart person.'
  4. 4Paid versions 'tended to go deeper and work faster, and were less likely to run out of steam' than free tiers.
  5. 5Outputs were described as 'impressive' but 'bland' — e.g., generic recommendations like 'who doesn't like good food?'
  6. 6The article flags 'generative engine optimisation' (GEO) — companies gaming models to appear in AI-generated recommendations — as a growing vulnerability.

They're not oracles... Think of it like working with a smart person.

Ethan Mollick Professor, Wharton School, University of Pennsylvania

On why AI travel planning underperforms — users provide too little context

Capability
Depth of responses Shallower Deeper
Speed Slower Faster
Stamina on long tasks More likely to run out of steam Less likely to run out of steam

Analysis

For ML practitioners, this reads less like a travel story and more like an unscientific but useful real-world benchmark of three frontier assistants on a messy, multi-constraint planning task. The undercover methodology — fresh accounts, two distinct traveller personas, and free and paid tiers — controls for the model's memory of the user and isolates the variable that actually moved results: the prompt. The observation that paid versions go 'deeper, faster, and don't run out of steam' is a practical signal about reasoning-tier inference and context retention, while the 'blandness' the author flags is the alignment-temperature trade-off showing up in the wild.

The stubborn belief that AI 'sucks at planning holidays' is really a failure of prompting rather than of the underlying models — that is the central finding of a hands-on experiment published by The Age and Brisbane Times on September 23, 2026. The journalist built two fictional five-day itineraries — São Paulo for a first-time visitor and New York for a veteran traveller who had exhausted the main sights — then queried three chatbots (ChatGPT, Gemini and Claude) in both free and paid tiers. To remove the models' ability to lean on her own history, she used fresh accounts and 'temporary chats' so the systems could not infer that she had lived in both cities for years. The result was a striking reversal: given enough upfront context — dates, budget, interests, pace and prior experience — the bots produced recommendations the author called 'impressive,' even if not enough to retire guidebooks, friend tips and late-night research.

For ML practitioners, this reads less like a travel story and more like an unscientific but useful real-world benchmark of three frontier assistants on a messy, multi-constraint planning task.

Wharton professor Ethan Mollick supplied the mental model that the test validates: 'They're not oracles... Think of it like working with a smart person.' A smart collaborator needs context before it can be useful, and so does a large language model. This reframes the popular 'AI is bad at X' complaint as partly a user-input problem. Prompt engineering, long a niche practitioner skill, is becoming an everyday consumer competency, and the quality of a model's output on open-ended planning tasks now scales as much with the brief a user writes as with the model's raw capability. That is a meaningful shift for an industry that has spent most of its marketing energy on benchmark scores and model size.

For AI builders, the test doubles as an unscientific but informative real-world benchmark of three frontier assistants on a messy, multi-constraint task. Two findings stand out. First, the paid-versus-free gap was observable and directional: paid versions 'tended to go deeper and work faster, and were less likely to run out of steam.' That is a practical signal that reasoning-tier inference, longer effective context and stronger instruction-following materially improve subjective output quality on long-horizon tasks — and a warning that free tiers may quietly underdeliver for heavy use. Second, outputs converge on bland middle ground. The author's dig — 'who doesn't like good food?' — is a symptom of the alignment-temperature trade-off: models default to the lowest-risk, most universally palatable answer when the prompt lacks the constraints that would force specificity.

The article also surfaces a genuinely important and under-covered risk: generative engine optimisation (GEO). Just as search engine optimisation reshaped the web, GEO sees businesses 'scramble to game the models so they appear' in AI-generated answers. In travel, that means restaurants, hotels and tour operators optimising their digital footprint to win placement inside chatbot itineraries — a recommendation that can look editorial but is effectively manipulated placement, usually invisible to the consumer. It is a trust and disclosure problem with no settled regulatory answer, and its scale will grow in direct proportion to how many people delegate planning to assistants.

What to Watch

The competitive subtext matters too. Pitting ChatGPT, Gemini and Claude on equal footing, the fact that all three cleared a 'usable itinerary' bar once prompted properly suggests differentiation is narrowing and that user skill — not model choice — is becoming the biggest variable in output quality. For product teams, that argues for shifting investment from raw capability toward UX that elicits constraints: structured intake forms, clarifying follow-ups before generation, and persistent memory that learns traveller preferences over time. The models already have the competence; the bottleneck is getting the right information out of the user.

Looking ahead, three implications stand out. Prompt literacy will harden into a marketable skill and a product-design responsibility, much as search literacy did two decades ago. Travel and other content-dependent industries face a GEO arms race that mirrors the birth of the SEO industry, complete with optimisation vendors, ranking games and eventual attempts at disclosure standards. And the free-versus-paid quality gap the author observed may widen if vendors gate reasoning-tier inference more aggressively — or narrow as base models improve. The practical takeaway is deceptively simple: treat the model as a collaborator, and the quality of a holiday plan, or any plan, will scale with the quality of the brief you give it.

Source cluster

Primary reporting

2articles

Cite This Page

"3 Chatbots, 2 Cities, 1 Lesson: Better Prompts Fix AI Trip Planning." AI Intelligence Brief, September 26, 2026. https://getaibrief.com/story/ai-travel-planning-prompt-test-chatgpt-gemini-claude

How we covered this story

Every story in our AI coverage is assembled from multiple primary sources, cross-referenced for factual consistency, and scored along three independent dimensions: sentiment, operational impact, and source-cluster confidence. Single-source rumors and unverifiable claims do not pass our editorial gate. When a story shows "Verified by N sources" with N≥2, the development is independently corroborated; when N=1, we mark it explicitly so readers can weigh the signal accordingly.

Impact scoring uses a 1-10 scale weighted toward regulatory, financial, and operational consequence rather than coverage volume. A topic that runs in every outlet but moves no real decisions ranks lower than a niche regulatory filing that reshapes how operators in the AI space have to behave. Read our full methodology for the scoring rubric, our glossary for term definitions, and our trends index for the longitudinal view across the beat.

Sources are only linked to a story once they clear our classification pipeline at a minimum 35 percent relevance threshold. According to that methodology, reviewed July 2026, this follows multi-source corroboration standards recommended by journalism research bodies such as the Reuters Institute for the Study of Journalism.

See something wrong in this story — a wrong fact, a broken source link, a misattributed entity? Report a data issue.