What problem does it solve?
Museums, heritage sites, cities and events hold far more knowledge than a visitor ever sees. Wall labels are short, recorded audio guides cover a fixed route in a handful of languages, and human guides are limited by schedules, group sizes and the languages they speak. Visitors who are curious but not experts, who speak another language, or who cannot read small print or see the object well get the thinnest experience. National Gallery Singapore puts it plainly: people walking into a museum often feel lost or intimidated.
Updating guide content is slow as well. Every new exhibition, route or audience (children, specialists, first time visitors) means rewriting and rerecording the same stories. Tourism boards face the same problem at city scale, with visitor questions arriving in dozens of languages and outside office hours. A generative guide changes the unit of work: curators maintain one body of approved content, and the guide tells it in the visitor's language, at the visitor's level, about the thing in front of them.
- UN Tourism estimates that 1.52 billion international tourists were recorded worldwide in 2025, almost 60 million more than in 2024.UN Tourism World Tourism Barometer (2026)
How does it work?
- Know where the visitor is. The guide receives context with each question: the room or stop, a GPS position, a QR code or object number, or a photo of the artwork that is matched against the collection (Art Basel's Lens returns artist and gallery details in about two seconds).
- Retrieve approved content. It looks up the object or place in the collection database or points of interest list and retrieves the curated texts, research and practical information (opening hours, accessibility, routes) that belong to it.
- Tell the story in the visitor's terms. The model turns that content into a short spoken or written answer in the visitor's language and at their level, and can connect it to interests the visitor mentions, as National Gallery Singapore's G(ai)le does with pop culture references.
- Answer follow up questions. Visitors ask in their own words, by voice or text; the guide stays within the approved content and says so when it does not know.
- Suggest what next. It recommends the next stop, event or exhibit based on the visitor's position, time and interests.
- Feed insight back. Anonymized questions show curators and marketing teams what visitors actually want to know, and where the content has gaps.
- Audience
- Customer facing
- Autonomy
- Autonomous
- Adoption
- Early adopters
- Channels
- Mobile app, Web chat, WhatsApp, Phone and voice, Kiosk and branch, Digital human
What is it worth?
Benchmarks are computed from the public deployments below: one data point per organization per KPI, with who made each claim.
No public deployment has disclosed a measurable outcome yet.
Value drivers: Customer experience, Inclusion and access, Revenue growth, Employee productivity.
Indicative value
A city museum with 1 million visitors a year and a guided route of about 150 stops
USD 10,000 to USD 200,000
Multilingual guide content production cost avoided per year
How this is calculated
Formula: stops * languages * costPerStopLanguage * refreshShare. The low scenario uses every low input, the high scenario every high input.
| Input | Low | High | Basis |
|---|---|---|---|
| Stops or objects with guide content stops, stops | 100 | 200 | Editorial assumption for a medium sized museum or city route. Replace with your own. |
| Additional languages offered languages, languages | 4 | 8 | Editorial assumption. For comparison, National Gallery Singapore's guide supports four languages in total. Replace with your own visitor language mix. |
| Cost to write, translate and record one stop in one language costPerStopLanguage, USD per stop per language | 100 | 250 | Editorial assumption for professional translation and voice recording. Replace with your own agency rates. |
| Share of stops rewritten or added each year refreshShare, fraction of stops per year | 0.25 | 0.5 | Editorial assumption covering new exhibitions, rotations and audience versions. |
What it leaves out: Counts only the avoided cost of writing, translating and recording guide content in extra languages. It leaves out the cost of running the AI and curating the source content, any revenue from longer visits, return visits or bookings, and the accessibility benefit, which is the main reason many institutions build a guide.
Who already uses it?
4 public deployments, strongest evidence first. Grades: A regulator or audit, B the organization itself, C vendor case study, D anonymous or estimate.
Art Basel
Switzerland · Media and entertainment · 2025
Art Basel's Companion app combines a conversational AI companion, grounded in gallery, artwork, dining and lodging information for its fairs in five cities, with the Art Basel Lens: a visitor photographs an artwork and receives artist and gallery details in about two seconds, matched by image embeddings against a vector index. Art Basel reports more engagement, return visits and time in the app, without publishing figures, and is exploring wayfinding and restaurant bookings.
No outcome disclosed.
National Gallery Singapore
Singapore · Government and public sector · 2025
National Gallery Singapore built G(ai)le, an AI docent that searches the Gallery's archives and explains artworks to visitors in conversational language, in English, Mandarin, Malay and Tamil. It adapts its stories to interests a visitor mentions, offers an audio only mode and an "eyes up" mode, and gives the Gallery a view of what visitors ask about. Staff also use it to draft versions of tour content for different audiences, which writers then refine.
No outcome disclosed.
Bloomberg Connects
United States · Media and entertainment · 2024
Bloomberg Connects uses Gemini to help create immersive audio guides, with the stated aim of making museums more accessible to visually impaired visitors. Google Cloud's listing gives no detail on scale, languages or results.
No outcome disclosed.
Madrid Destino
Spain · Government and public sector · 2024
Madrid Destino, the city's municipal tourism office, runs VisitMadridGPT, a virtual assistant that answers visitors in more than 95 languages from the city's official, expert curated tourism site, available when physical offices are closed. The city analyses the questions to find the most requested topics and adjust its website content. According to the story, Madrid attracted 10.6 million visitors in 2023.
No outcome disclosed.
How do you implement it?
A model agnostic playbook: what to prepare, the order to build in, and what goes wrong.
Data you need
- A collection management or points of interest database with stable identifiers per object or place
- Approved, curated texts per object or place, with an owner and a review date
- Practical visitor information (opening hours, routes, accessibility, events) from one source
- A pronunciation list of names of artists, places and local terms for speech recognition and synthesis
- For image recognition, reference photos of each object and the rights to use them
Systems to integrate
- Collection management system or content management system
- The organization's visitor app, website or kiosk software
- Positioning (GPS, beacons, QR codes) or an image matching service
- Ticketing and events calendar for recommendations and practical questions
- Analytics for anonymized question and usage reporting
Complexity: Medium
The conversation is the easy part. The work is in the content: a clean collection or points of interest database with identifiers, approved texts per object, rights to use images and research, and a reliable way to know where the visitor is (QR codes, beacons, GPS or image matching). Spoken delivery in a noisy hall and accurate recognition of local names and accents need testing on site.
- 1
Start from the content, not the model
Pick one gallery, route or district and bring its content into shape: an identifier per object or place, the approved texts, practical information and a short brief on tone. The quality of the guide will never exceed the quality of this content.
- 2
Decide how the guide knows where the visitor is
QR codes or object numbers are the most reliable and cheapest. Image matching needs no codes on the wall, but it needs reference photos and tests in real lighting and angles; GPS works outdoors but not between rooms. Many deployments use two methods with a fallback.
- 3
Write the voice of the guide
Agree with curators how the guide speaks, what it may interpret and what it must leave open. National Gallery Singapore spent a long time tuning prompts so its docent informs without imposing a single view of the art.
- 4
Test with real visitors and real languages
Build a test set of questions per stop in every supported language, including names that are hard to pronounce and questions the content cannot answer, and run it on every content or model change. Then run a pilot on the floor and listen to what visitors ask.
- 5
Design for access from day one
Offer audio only and large text modes, captions for spoken replies, and a way to use the guide without looking at the screen. These features serve visually impaired visitors and everyone who wants to look at the object rather than a phone.
- 6
Close the loop with curators
Review anonymized questions every month: frequent questions without a good answer become new content, and questions that show confusion feed back into labels and routes.
Guardrails
- Answers only from the approved collection and visitor content, with a clear "I do not know" when the content is silent
- Facts such as dates, attributions and prices are taken from the source record, never generated
- Clear disclosure that the guide is an AI, and a way to reach staff for practical or safety questions
- No identification of people in camera images; image matching limited to objects in the collection
- Location and conversation data kept only as long as needed, anonymized for analytics
KPIs to instrument
- Share of visitors who use the guide, by language and entry point
- Questions per session and share answered from content versus declined
- Factual accuracy on a monthly sample reviewed by curators
- Visitor satisfaction with the guide compared with the recorded audio guide or human tour
- Speech recognition errors on names and local terms, per language
Human in the loop
Curators and educators own the content and the tone of the guide, approve new stories before they go live, and review a sample of conversations each month for errors of fact and tone. Front of house staff handle anything practical the guide cannot, such as lost items, accessibility assistance and safety.
Common failure modes
- Confident but wrong stories
- The guide invents a date, an attribution or an anecdote that sounds right. Prevent with retrieval from approved records, facts taken from structured fields and curator review of samples.
- Wrong object, wrong story
- Image matching or positioning picks the neighbouring object and the guide tells the wrong story. Set a confidence threshold and ask the visitor to confirm or scan the code when unsure.
- A screen between visitor and art
- Visitors stare at their phones instead of the object. Design audio first and short answers, as National Gallery Singapore did with its "eyes up" mode.
- Stale practical information
- Opening hours, closed rooms or event times in the guide differ from reality. Pull practical information from one live source instead of copying it into the content.
What are the risks and rules?
EU AI Act
Limited risk (transparency)
A visitor facing assistant must make clear that people are interacting with AI (Article 50), and synthetic speech should be identifiable as AI generated. It is not high risk. It would change if the camera feature were used to identify or categorise visitors by biometric data, which a guide does not need: remote biometric identification, biometric categorisation and emotion recognition are high risk under Annex III point 1, and biometric categorisation that infers sensitive traits is prohibited under Article 5.
Guidance
- Article 50, transparency obligations for providers and deployers of certain AI systems (European Union, Europe). People must be informed that they are interacting with an AI system, and providers of systems that generate synthetic audio must mark it in a machine readable format as artificially generated.
- Guidelines 02/2021 on virtual voice assistants (European Data Protection Board, Europe). How GDPR applies to voice assistants, including transparency, purpose limitation and retention of voice recordings.
Controls to put in place
- AI disclosure at the start of every session and on spoken replies
- Content ownership per object or place, with review dates and a change log
- Test set per stop and language, run before every content or model change
- Data protection impact assessment covering location data, voice and camera images
- Retention limits and anonymization for conversation logs used in analytics
Frequently asked questions
- Does an AI guide replace human guides and docents?
- In the deployments on this page, no. National Gallery Singapore describes its AI docent as a new teammate alongside the tours team, and uses it to draft tour versions that staff refine. The Gallery notes that in person tours are not always available in the language a visitor prefers, which is where the guide helps.
- How does the guide know which object or place the visitor means?
- Through context sent with each question: a QR code or object number, GPS outdoors, or image recognition. Art Basel matches a photo of an artwork against its index and returns the artist and gallery in about two seconds. Codes are the most reliable; image matching needs reference photos and a confidence threshold.
- How do you stop the guide from making up facts about the collection?
- Ground every answer in approved records, take hard facts such as dates and attributions from structured fields, make the guide say when the content is silent, and have curators review a sample of conversations every month.
- Is an AI tour guide high risk under the EU AI Act?
- Not as described here. It carries the Article 50 transparency duties: visitors must know it is an AI, and synthetic speech must be identifiable. Location, voice and camera data still fall under the GDPR, so a data protection impact assessment is advisable. Using the camera to identify or categorise visitors would change the assessment.
How to cite this page
Blits.ai AI Use Case Library, "AI visitor and tour guide for cities, museums and events", last verified 26 September 2026, https://www.blits.ai/ai-use-cases/ai-visitor-and-tour-guide. Licensed under CC BY 4.0. Method: how we verify use cases.
Changelog
- 27 September 2026: First published