A poem can receive a different judgment even when none of its words change. In a new experiment, readers gave lower liking scores to poems they believed were produced by artificial intelligence and higher scores to poems they believed came from people. The result exposes how assumptions about authorship can shape the experience of art before a reader knows the truth.
Readers judged poems before learning their origin
The study tested reactions to Czech poetry written by humans and generated by AI. Participants read 16 short poems, rated how much they liked each one and guessed whether its author was a person or a machine. That structure let the researchers compare actual authorship with perceived authorship for every response.
The peer-reviewed research article in Digital Scholarship in the Humanities reports a final sample of 126 participants and 2,016 poem evaluations. Readers identified the true source correctly 45.8 percent of the time, a result close enough to chance to show how difficult the distinction was in this setting.
The inability to identify authors reliably did not make the labels irrelevant. Poems believed to be human received an average liking score of 2.3, while those believed to be AI-generated averaged 1.0. Those were reactions to perceived origin, not a simple comparison between two fixed sets of poems.
The actual AI poems performed well
When researchers grouped ratings by the poems’ real origin, the machine-written selections did not fare worse. AI-generated poems received an average liking score of 2.0, compared with 1.4 for the human-written selections. The finding complicates the familiar idea that readers can recognize and reject machine-made writing through style alone.
The study does not show that AI poetry is generally superior to human poetry. Its authors chose a limited set of texts, and performance can change with the model, prompt, poet, genre and language. The stronger conclusion is narrower: in this experiment, the actual AI poems were competitive, while the belief that a poem was made by AI was associated with a penalty.
Authorship can act like a frame
People rarely encounter art as isolated language. A byline, reputation, price, venue or story about how a work was made supplies a frame for interpreting it. Human authorship can suggest intention, biography and effort; machine authorship can suggest statistical imitation or reduced creative labor. Those expectations may change how much meaning a reader searches for in the same lines.
The effect has practical consequences for publishers, contests and classrooms. A disclosure can encourage honesty while also influencing evaluation before readers consider the work itself. Hiding AI involvement creates a different problem because it denies audiences information they may reasonably consider relevant to creative ownership and labor.
The researchers made their study materials and supporting data available through OSF. That transparency allows other scholars to inspect the choices behind the experiment and test whether the pattern appears with other languages, forms and groups of readers.
The direction of influence remains uncertain
The result is an association between a participant’s guess and a rating, not clean proof that an AI label caused dislike. A reader might first dislike a poem and then infer that a machine wrote it. The reverse could also happen: suspecting AI authorship might lower the rating. Both processes may have operated at once.
A stronger causal test would randomly attach human or AI labels to identical poems and compare reactions between groups. Such a design would isolate the effect of the label, although deception and later disclosure would require careful handling. The present experiment is valuable because it captures readers’ spontaneous assumptions, but those assumptions cannot reveal their own causal order.
Future experiments could separate several possible judgments by asking about beauty, originality, emotional depth and technical skill independently. They could also reveal authorship before reading, after reading or not at all. Comparing those conditions would show whether an AI label changes initial attention, the interpretation of particular lines or only the final score a reader chooses.
A small experiment raises a larger question
The sample consisted mainly of nonexperts, and the poems were short stanzas removed from a wider literary context. Czech-language findings may not carry unchanged into other traditions. Long narrative poems, familiar authors or live readings could give audiences more signals and different reasons for judging authenticity.
Even with those limits, the study shows that the debate over AI art is not only about output quality. Readers also care about agency, effort and the relationship between a work and its maker. As generated text becomes harder to distinguish, perceived authorship may become an increasingly powerful part of the artwork itself.
The same poem can therefore occupy two emotional categories at once. When readers see a human hand behind it, they may reward intention and connection. When they imagine a model, they may discount those qualities even if the words have not moved at all.
This article was produced with the assistance of AI and reviewed by Morning Overview editors prior to publication.
More from Morning Overview