Simplify the interface and guide the workflow to make evaluation feel controllable
The first step to making open-ended art evaluation feel controllable is to reduce the complexity of the tool itself. A 2026 action design research study at Ubisoft found that off-the-shelf generative AI tools have complex user interfaces that are misaligned with production workflows, causing issues like AI illiteracy and employee resistance [1]. By introducing a simplified interface with workflow guidance on top of Stable Diffusion, they saw increased user acceptance and productivity [1]. This means that when you design an evaluation tool, you should strip away unnecessary options and provide step-by-step guidance—this makes the process feel manageable and controllable, even when the underlying generation is open-ended.
Calibrate human judgments with references to make evaluation reliable
A major challenge in open-ended evaluation is that human raters often can't tell AI-generated art from human-made art without a reference point. A 2021 study on story generation found that Amazon Mechanical Turk workers, even with strict qualifications, failed to distinguish model-generated text from human-written text [5]. However, when workers were shown model output alongside human references, their judgments improved significantly [5]. This suggests that for art evaluation, you should always provide a side-by-side comparison or a calibration set—this gives raters a concrete anchor, making their evaluations more consistent and controllable.
Use text prompts and novelty search to steer diversity while keeping generation open-ended
To control the evaluation of open-ended generation, you need to guide the diversity of outputs without closing off the space. MarioGPT, a text-to-level generator for Super Mario Bros, showed that fine-tuning a language model with text prompts allows for controllable level generation [4]. When combined with novelty search, it produced an increasingly diverse range of levels with varying play-style dynamics [4]. This means that in an evaluation tool, you can let users specify high-level goals (e.g., 'more abstract' or 'more realistic') and then use novelty search to ensure the outputs remain varied—this gives a sense of control over the open-ended process.
About These Sources
This answer is built on 5 peer-reviewed studies — published from 2021 to 2026, 3 from 2024 or later, 1 in Q1 journals, collectively cited 389 times — selected as the most relevant from 5 studies that passed quality screening, drawn from 48 papers retrieved from a database of over 500 million.
Sources used in this answer
Design principles for text-to-image generative artificial intelligence creativity support tools for visual design
In an action design research study at Ubisoft, a simplified interface with workflow guidance on Stable Diffusion increased user acceptance and productivity, yielding seven design principles for text-to-image generative AI creativity support tools.
Learning to Evaluate the Artness of AI-Generated Images
ArtScore, a metric trained on interpolated images between photo and artwork, aligns more closely with human artistic evaluation than existing metrics like Gram loss and ArtFID, enabling instance-level and reference-free artness assessment.
Artificial Hivemind: The Open-Ended Homogeneity of Language Models (and Beyond)
Infinity-Chat, a dataset of 26K open-ended queries with 31,250 human annotations, reveals an 'Artificial Hivemind' effect where language models produce homogeneous outputs, and shows that reward models and judges are less calibrated to human ratings on idiosyncratic preferences.
MarioGPT: Open-Ended Text2Level Generation through Large Language Models
MarioGPT, a fine-tuned GPT2 model, generates diverse Super Mario Bros levels from text prompts, and when combined with novelty search, enables open-ended generation of increasingly diverse content with varying play-style dynamics.
The Perils of Using Mechanical Turk to Evaluate Open-Ended Text Generation
A survey of 45 open-ended text generation papers found most fail to report crucial AMT task details, and experiments showed that AMT workers, unlike English teachers, could not distinguish model-generated from human text unless shown references, which improved their calibration.
