Using AI as a Tool for Visual Descriptions

PAC partnered with the Cleveland Museum of Art to train staff, create a Visual Description Style Guide, and responsibly integrate AI with human review and image-level metadata, establishing scalable workflows to deliver short visual descriptions across the museum’s online collection.
Media



Project Description
Online museum collections have opened access to extraordinary cultural resources around the world. But for blind and low-vision visitors, an online collection without visual descriptions can still amount to thousands of blank spaces where artworks should be.
Prime Access Consulting (PAC) and Cleveland Museum of Art (CMA) wanted to change that.
When CMA launched its redesigned website in 2024, accessibility was a central priority. With more than 68,000 artworks in its collection, and more than 230,000 digital images when additional views of three-dimensional objects and conservation photography are included, the scale of providing meaningful visual access was enormous. CMA set an ambitious initial goal: provide a short visual description for every primary collection image online.
The challenge was not convincing the museum that descriptions were important. It was figuring out how to produce tens of thousands of them without sacrificing quality.
PAC partnered with CMA to build the foundation for doing exactly that. Before asking what artificial intelligence could produce, PAC and CMA spent a year asking a more fundamental question: What makes a good visual description? That distinction shaped everything that followed.
Building the Human Expertise First
Generative AI can produce text almost instantaneously. That does not mean it knows what information matters to a blind person, how to prioritize visual details, how institutional voice should shape a description, or how to navigate the complex questions of identity and bias that emerge when describing people.
PAC began by building that expertise within CMA through our capacity building module, Visual Description Practice. Through staff training, PAC introduced the foundations of visual description: what it is, how it differs from interpretive museum content, how information should be prioritized, and how seemingly simple choices in language can affect the experience of blind and low-vision audiences.
Staff practiced writing descriptions themselves, developing the knowledge needed to recognize both strong descriptions and problematic ones. CMA then submitted staff-authored descriptions to PAC for review. PAC provided detailed feedback, allowing the team to refine its approach through repeated writing, critique, and iteration.
The process also forced conversations that cannot be resolved by a generic set of accessibility guidelines. How should the museum describe race, ethnicity, gender, or bodies without projecting assumptions onto the people represented? How should an abstract work be described? Which details matter most when a description needs to remain concise? How should the museum’s own voice and values come through without slipping into interpretation? The answers became the foundation for the next stage of the work.
Turning Practice into a Visual Description Style Guide
PAC worked with CMA to codify what the team was learning into a comprehensive Visual Description Style Guide.
Part teaching tool and part institutional standard, the guide established a consistent approach for both staff and, eventually, AI-generated descriptions. It addressed general best practices, the written architecture of a description, prioritization of information, core visual characteristics, descriptions of people, different contexts for description, and CMA’s institutional voice.
For collection images, the team focused on short descriptions of approximately 75 words that communicate what is visually available in the image rather than providing art-historical interpretation. Descriptions generally begin with an overview before moving into prioritized details such as color, medium, shape, texture, size, orientation, and spatial relationships.
This was more than an editorial exercise. The Style Guide translated knowledge into infrastructure. It created a shared standard that could guide staff, support quality review, survive changes in personnel, and provide the basis for testing emerging technologies. Only after establishing that foundation did the project seriously turn toward AI.
Using AI Without Outsourcing Judgment
The scale problem was stark. After approximately a year, CMA had produced around 500 human-authored descriptions that met its standards. More than 67,000 primary collection images remained. At that pace, accessibility would always be chasing the collection.
CMA began investigating whether generative AI could close that gap, testing a succession of models as the technology rapidly evolved. Early experiments produced uneven results. Models struggled with artwork, rejected certain imagery, introduced inaccuracies, and raised particular concerns around portraits and the description of people. Later models performed significantly better, with Gemini ultimately producing the strongest results during the period documented by the project.
The Visual Description Style Guide and PAC-reviewed descriptions became critical to this process. CMA used hundreds of PAC-approved examples to refine AI output, while further fine-tuning the Style Guide’s principles into instructions the models could apply. Key visual information—including shape, size, color, orientation, and positioning—could be explicitly prioritized rather than leaving the model to decide what was important. In this model, AI did not determine what constituted a successful description. People did.
PAC’s role was therefore not simply to review machine-generated text. It was to help establish the human-centered framework against which the technology could be trained, tested, challenged, and improved. The project approached AI as an amplifier: if the underlying practices are thoughtful, informed, and accountable, technology can help scale them. If those practices contain gaps or biases, AI can scale those just as easily.
For that reason, inclusive design expertise and input from disabled people cannot enter the process after the technology has already been built. They have to be part of defining what success means in the first place.
Testing What the Technology Produces
Scaling production also required scaling evaluation. PAC continued to provide human review of descriptions while the broader team explored multiple methods for evaluating AI output. Rather than assuming that a description was successful because it sounded plausible, the process tested whether it actually conveyed the artwork according to the criteria established through the Style Guide.
CMA also planned for evaluation to continue once descriptions reached the public. Redesigned artwork detail pages were developed to surface visual descriptions visibly rather than limiting them to screen-reader-accessible alt text, while giving visitors a way to contact the museum when something in a description was inaccurate or needed improvement.
Building Description into the Digital Collection
Producing descriptions was only part of the challenge. CMA also needed a sustainable way to store and publish them.
A description belongs to a particular image, not simply to an artwork. A three-dimensional object photographed from the front may require a substantially different description when photographed from behind. Connecting a single description to the collection object record could therefore result in inaccurate alt text when different images were displayed.
CMA addressed this by storing descriptions as image-level metadata in its Digital Asset Management System rather than attaching them only to the object record in the Collection Management System. Each description could then remain associated with the exact image it describes and surface as alt text wherever that asset appears across the website.
The approach transforms visual description from isolated editorial copy into structured digital collection data, creating an infrastructure that can grow alongside the collection.
The work is also expanding beyond primary artwork images. PAC and CMA have been developing approaches for additional views, interactive 3D models, and conservation images, each of which introduces different questions about what needs to be described and how a visitor understands visual and spatial information.
Changing the Scale of What Is Possible
The problem facing CMA is not unique. Even museums with longstanding accessibility programs often have only a fraction of their collections described. New images and digital assets are created faster than staff can manually describe them, leaving blind and low-vision audiences perpetually waiting for access. PAC and CMA’s work offers another possibility.
The project demonstrates that scale does not have to mean abandoning quality, and responsible AI does not begin with choosing an AI model. It begins with people: training staff, listening to disabled expertise, defining what good description looks like, codifying those decisions, creating sustainable content workflows, then determining where technology can responsibly accelerate the work.
The goal is not AI-generated visual description for its own sake. The goal is to reach a point where the question blind visitors have been forced to ask for decades “Why isn’t that described?”no longer needs to be asked.