Using AI as a Tool for Visual Descriptions

A screenshot of The Cleveland Museum of Art's collection page. The main content shows an oil painting, the Portrait of Dora Wheeler, by William Merritt Chase. In the painting, a woman in a blue dress sits in a dark wooden chair, resting her head on one hand, beside a small table holding a vase of yellow flowers. The background is a soft yellow patterned wall or screen with loose floral decoration. On the right side, the page lists the artwork title, date "1882-83," artist "William Merritt Chase," and details including culture, medium, measurements, credit line, public domain status, and location. Buttons for "ArtLens App," At the top is the museum's logo and navigation menu with options such as "Visit," "What's On," "Art," "Donate," "Join/Log In," and "Search." "Share," "Download," and "Print" appear below the information.

PAC partnered with the Cleveland Museum of Art to train staff, create a Visual Description Style Guide, and responsibly integrate AI with human review and image-level metadata, establishing scalable workflows to deliver short visual descriptions across the museum’s online collection.

Media

A museum collection webpage for Portrait of Dora Wheeler by William Merritt Chase, dated 1882-83. A white modal window titled "Visual Description" dominates the center, with an "X" close button in the top-right. The visual description reads, "A square oil painting depicts a woman with light skin tone reclining on a carved wood chair. Dressed in a floor-length, fur-lined deep blue gown, her body angles toward our left. She gazes past our right shoulder, chin resting on her left hand, propped up on the leather arm of the chair. On a small table behind her, a blue vase is filled with yellow flowers set against a muted yellow tapestry." Underneath the description it reads, "Notice something inaccurate with this visual description? Fill out our feedback form" with a hyper link. Behind the semi-transparent overlay, parts of the artwork, page title, artist name, dimensions, location, and other collection details are visible.
A museum-style "Contact Us" web form. A black header bar at the top contains large white text reading "Contact Us." Below, the page says, "Have a question about or issue with an artwork? Use this form." In the "Reason for Request" section, a required dropdown field is selected as "Feedback on Visual Description." A paragraph below reads, "At our institution, visual descriptions are approximately 75-word textual representations of static media meant to support accessibility, especially for users who are blind or have low vision. Information not available from observing the artwork such as contextual information about the artist, time period of creation, interpretation, etc. is the purview of the "Description" field." Farther down, an "Artwork" section asks users to search by artwork title or accession number. A selected artwork field shows "Portrait of Dora Wheeler William Merritt Chase (American..." with a required asterisk, and a black "Remove Artwork" button appears to the right.
An "Edit Metadata" interface of a website. At the upper left is a preview image of a decorative two-sided drinking cup with black handles and a painted satyr-like face, including large eyes, ears, facial hair, and an ornate cylindrical top. Beneath the preview is a blue "Save Changes" button. The lower portion contains a "GENERAL" metadata section with empty fields labeled "ANNOTATE_TEXT" and "BOOK_VIEWER_RECORD_TYPE," followed by an "ALT TEXT" field filled with a detailed description of the cup's appearance, "Two-sided drinking cup with the face of a satyr, a man with horse ears, as the body, a cylindrical lip with black geometric patterns on orange extending from his head, flanked by two vertical, black handles. The satyr has medium-light skin tone, a fluffy black beard along his jaw, and a moustache and goatee the same color as his skin. His white, rectangular teeth press together and a white, dot-speckled band extends across his forehead."

Project Description

Online museum collections have opened access to extraordinary cultural resources around the world. But for blind and low-vision visitors, an online collection without visual descriptions can still amount to thousands of blank spaces where artworks should be.

Prime Access Consulting (PAC) and Cleveland Museum of Art (CMA) wanted to change that.

When CMA launched its redesigned website in 2024, accessibility was a central priority. With more than 68,000 artworks in its collection, and more than 230,000 digital images when additional views of three-dimensional objects and conservation photography are included, the scale of providing meaningful visual access was enormous. CMA set an ambitious initial goal: provide a short visual description for every primary collection image online.

The challenge was not convincing the museum that descriptions were important. It was figuring out how to produce tens of thousands of them without sacrificing quality.

PAC partnered with CMA to build the foundation for doing exactly that. Before asking what artificial intelligence could produce, PAC and CMA spent a year asking a more fundamental question: What makes a good visual description? That distinction shaped everything that followed.

Building the Human Expertise First

Generative AI can produce text almost instantaneously. That does not mean it knows what information matters to a blind person, how to prioritize visual details, how institutional voice should shape a description, or how to navigate the complex questions of identity and bias that emerge when describing people.

PAC began by building that expertise within CMA through our capacity building module, Visual Description Practice. Through staff training, PAC introduced the foundations of visual description: what it is, how it differs from interpretive museum content, how information should be prioritized, and how seemingly simple choices in language can affect the experience of blind and low-vision audiences.

Staff practiced writing descriptions themselves, developing the knowledge needed to recognize both strong descriptions and problematic ones. CMA then submitted staff-authored descriptions to PAC for review. PAC provided detailed feedback, allowing the team to refine its approach through repeated writing, critique, and iteration.

The process also forced conversations that cannot be resolved by a generic set of accessibility guidelines. How should the museum describe race, ethnicity, gender, or bodies without projecting assumptions onto the people represented? How should an abstract work be described? Which details matter most when a description needs to remain concise? How should the museum’s own voice and values come through without slipping into interpretation? The answers became the foundation for the next stage of the work.

Turning Practice into a Visual Description Style Guide

PAC worked with CMA to codify what the team was learning into a comprehensive Visual Description Style Guide.

Part teaching tool and part institutional standard, the guide established a consistent approach for both staff and, eventually, AI-generated descriptions. It addressed general best practices, the written architecture of a description, prioritization of information, core visual characteristics, descriptions of people, different contexts for description, and CMA’s institutional voice.

For collection images, the team focused on short descriptions of approximately 75 words that communicate what is visually available in the image rather than providing art-historical interpretation. Descriptions generally begin with an overview before moving into prioritized details such as color, medium, shape, texture, size, orientation, and spatial relationships.

This was more than an editorial exercise. The Style Guide translated knowledge into infrastructure. It created a shared standard that could guide staff, support quality review, survive changes in personnel, and provide the basis for testing emerging technologies. Only after establishing that foundation did the project seriously turn toward AI.

Using AI Without Outsourcing Judgment

The scale problem was stark. After approximately a year, CMA had produced around 500 human-authored descriptions that met its standards. More than 67,000 primary collection images remained. At that pace, accessibility would always be chasing the collection.

CMA began investigating whether generative AI could close that gap, testing a succession of models as the technology rapidly evolved. Early experiments produced uneven results. Models struggled with artwork, rejected certain imagery, introduced inaccuracies, and raised particular concerns around portraits and the description of people. Later models performed significantly better, with Gemini ultimately producing the strongest results during the period documented by the project.

The Visual Description Style Guide and PAC-reviewed descriptions became critical to this process. CMA used hundreds of PAC-approved examples to refine AI output, while further fine-tuning the Style Guide’s principles into instructions the models could apply. Key visual information—including shape, size, color, orientation, and positioning—could be explicitly prioritized rather than leaving the model to decide what was important. In this model, AI did not determine what constituted a successful description. People did.

PAC’s role was therefore not simply to review machine-generated text. It was to help establish the human-centered framework against which the technology could be trained, tested, challenged, and improved. The project approached AI as an amplifier: if the underlying practices are thoughtful, informed, and accountable, technology can help scale them. If those practices contain gaps or biases, AI can scale those just as easily.

For that reason, inclusive design expertise and input from disabled people cannot enter the process after the technology has already been built. They have to be part of defining what success means in the first place.

Testing What the Technology Produces

Scaling production also required scaling evaluation. PAC continued to provide human review of descriptions while the broader team explored multiple methods for evaluating AI output. Rather than assuming that a description was successful because it sounded plausible, the process tested whether it actually conveyed the artwork according to the criteria established through the Style Guide.

CMA also planned for evaluation to continue once descriptions reached the public. Redesigned artwork detail pages were developed to surface visual descriptions visibly rather than limiting them to screen-reader-accessible alt text, while giving visitors a way to contact the museum when something in a description was inaccurate or needed improvement.

Building Description into the Digital Collection

Producing descriptions was only part of the challenge. CMA also needed a sustainable way to store and publish them.

A description belongs to a particular image, not simply to an artwork. A three-dimensional object photographed from the front may require a substantially different description when photographed from behind. Connecting a single description to the collection object record could therefore result in inaccurate alt text when different images were displayed.

CMA addressed this by storing descriptions as image-level metadata in its Digital Asset Management System rather than attaching them only to the object record in the Collection Management System. Each description could then remain associated with the exact image it describes and surface as alt text wherever that asset appears across the website.

The approach transforms visual description from isolated editorial copy into structured digital collection data, creating an infrastructure that can grow alongside the collection.

The work is also expanding beyond primary artwork images. PAC and CMA have been developing approaches for additional views, interactive 3D models, and conservation images, each of which introduces different questions about what needs to be described and how a visitor understands visual and spatial information.

Changing the Scale of What Is Possible

The problem facing CMA is not unique. Even museums with longstanding accessibility programs often have only a fraction of their collections described. New images and digital assets are created faster than staff can manually describe them, leaving blind and low-vision audiences perpetually waiting for access. PAC and CMA’s work offers another possibility.

The project demonstrates that scale does not have to mean abandoning quality, and responsible AI does not begin with choosing an AI model. It begins with people: training staff, listening to disabled expertise, defining what good description looks like, codifying those decisions, creating sustainable content workflows, then determining where technology can responsibly accelerate the work.

The goal is not AI-generated visual description for its own sake. The goal is to reach a point where the question blind visitors have been forced to ask for decades “Why isn’t that described?”no longer needs to be asked.