Vulsar AI has published a benchmark that tests the ability of 24 artificial intelligence models to write fiction and compares them against texts by human writers. The study, based on 475 prompts, puts GPT 6 Astra at the top with a predicted win rate of 87.8% in the overall ranking, versus 86.6% for the human baseline, though the gap changes significantly once professional and amateur writers are separated out.
Key facts about Vulsar’s creative writing benchmark, in 20 seconds
- Vulsar compared 24 AI models against human writers using 475 prompts.
- GPT 6 Astra gets an 87.8% predicted win rate in the overall result.
- Human writers reach 86.6% in that same ranking.
- Against professional writers, Astra drops to 76.2%, while the human baseline reaches 99.9%.
- Against amateurs, Astra reaches 93.9%, versus 79.7% for human writers.
The benchmark, published on September 17, 2026, attempts to measure a capability that is harder to compare than the usual math, coding or knowledge tests. Vulsar isn’t simply trying to determine whether a model can produce coherent text, but to estimate which stories human readers would be more likely to prefer.
The methodology uses a reward model trained on human preferences for creative writing. Vulsar does not use another language model as a final judge, nor a traditional rubric with several independent criteria. The system compares responses and calculates a predicted win rate for each participant.
GPT 6 Astra Tops the Overall Ranking
GPT 6 Astra gets an 87.8% predicted win rate, versus 86.6% for the baseline made up of human writers. The gap is 1.2 percentage points, within a result that Vulsar presents as a preference estimate rather than an absolute measure of literary quality.
| Rank | Model or baseline | Predicted win rate |
|---|---|---|
| 1 | GPT 6 Astra | 87.8% |
| 2 | Human writers | 86.6% |
| 3 | GPT 5.6 Sol | 77.6% |
| 4 | Muse Spark 1.3 | 70.6% |
| 5 | Claude Fable 5.1 | 70.3% |
| 6 | Claude Opus 5 | 63.4% |
| 7 | GPT 5.6 Terra | 62.6% |
| 8 | GPT 5.6 Luna | 61.8% |
| 9 | Kimi K3 | 61.0% |
| 10 | Grok 4.6 | 59.3% |
Further down the list come GLM 5.3 Flash at 54.6%, Qwen3.8 2.4T A95B at 53.9%, and Inkling at 52.3%. Near the bottom of the ranking are Gemma 4 31B at 23.4%, Qwen3.8 27B at 23.2%, DeepSeek V4.1 Flash at 19.8%, DeepSeek V3.2 at 16.5%, and Gemma 4 26B A4B at 10.9%.
The study also records differences in story length. Human writers produce an average of 2,592 tokens, while GPT 6 Astra stays at 1,537. Claude Opus 5 reaches 2,330 and Claude Fable 5.1 reaches 2,114.
Length, however, is not used as a criterion for the ranking. The benchmark is designed to measure the preference predicted by Vulsar’s reward model.
Professional Writers Show a Much Bigger Gap
The picture changes notably when the benchmark isolates professional writers. In this category, the human baseline gets a predicted win rate of 99.9%. GPT 6 Astra scores 76.2%, while Grok 4.6 reaches 78.5%.
| Rank | Model or baseline | Predicted win rate |
|---|---|---|
| 1 | Professional writers | 99.9% |
| 2 | Grok 4.6 | 78.5% |
| 3 | GPT 6 Astra | 76.2% |
| 4 | Claude Fable 5.1 | 69.8% |
| 5 | Muse Spark 1.3 | 64.9% |
| 6 | Qwen3.8 2.4T A95B | 63.6% |
| 7 | Claude Opus 5 | 63.3% |
| 8 | GPT 5.6 Sol | 63.1% |
The gap between the models and professional authors is even more visible when looking at text length. Professional writers average 4,528 tokens per story, versus 1,745 for GPT 6 Astra.
Vulsar explains that the prompts used with professionals include more detailed instructions and specific requirements. Because of that, this part of the benchmark shouldn’t simply be read as a repeat of the overall ranking with a different group of participants.
The comparison shows that a model’s performance can vary depending on the type of task and the experience level of the human baseline. An overall ranking can hide those differences.
Models Score Better Against Amateurs
The amateur writer category paints a different picture. GPT 6 Astra reaches a predicted win rate of 93.9%, GPT 5.6 Sol gets 85.2%, and the amateur writer baseline sits at 79.7%.
| Rank | Model or baseline | Predicted win rate |
|---|---|---|
| 1 | GPT 6 Astra | 93.9% |
| 2 | GPT 5.6 Sol | 85.2% |
| 3 | Amateur writers | 79.7% |
| 4 | Muse Spark 1.3 | 73.5% |
| 5 | Claude Fable 5.1 | 70.5% |
| 6 | GPT 5.6 Terra | 70.2% |
| 7 | GPT 5.6 Luna | 67.4% |
| 8 | Claude Opus 5 | 63.5% |
The difference between this category and the professional one is one of the study’s most notable aspects. GPT 6 Astra goes from 93.9% against amateurs to 76.2% against professionals.
Nor does the result turn the benchmark into a general answer on whether artificial intelligence writes better than people. The comparison depends on the set of prompts, the baseline used and the system employed to estimate preferences — a point worth keeping in mind alongside how Anthropic has positioned Claude Fable 5 and its successors for broader, general-purpose use rather than narrow benchmark chasing.
475 Prompts to Measure a Hard-to-Compare Capability
Vulsar used 475 unique prompts and asked participants to generate a short story for each task. The tests cover different types of fiction and use the same instructions for the models within each evaluation set.
The company disabled reasoning when a model offered that option. When it couldn’t, it used the lowest available reasoning level. Empty responses and refusals are not automatically counted as losses; instead, they’re excluded from the relevant comparisons.
This detail matters because Vulsar’s win rate is not equivalent to an accuracy percentage. An 87.8% score means that, within the benchmark’s set of comparisons, the evaluation system estimates a favorable preference for GPT 6 Astra’s responses at that rate.
Confidence intervals add further context. For GPT 6 Astra, Vulsar puts the 95% confidence interval between 86.7% and 88.9%. For the overall human baseline, the interval sits between 85.5% and 87.7%.
The Benchmark Doesn’t Independently Measure Originality
Vulsar also acknowledges a limitation related to originality. Its reward model has been trained on human preferences that may value the sense of novelty in a story, but the benchmark doesn’t independently check whether the ideas or phrasing in the texts are genuinely new compared with existing works or the models’ training data.
Because of that, a high win rate shouldn’t be read as a certification of originality.
The same caution applies to literary quality. The benchmark aims to estimate aggregated human preferences, not to establish a universal measure of quality. Individual tastes may differ from the reward model’s prediction.
For evaluating AI models, the study is especially useful because it introduces a dimension that doesn’t usually appear in technical benchmarks: comparing automatic generation against human preferences on an open-ended task.
The results also show that performance isn’t distributed the same way when the reference writers’ skill level changes. GPT 6 Astra lands roughly at the level of the general human baseline, beats the amateur baseline, and falls considerably further behind professional writers within the metrics Vulsar published.
That makes Creative Writing Bench V1 a specific test of creative writing, but not a general measure of language models’ capabilities.
Frequently Asked Questions
Which models take part in Vulsar’s benchmark?
Vulsar compares 24 artificial intelligence models against human writer baselines using 475 creative writing prompts. The ranking includes models from different families and sizes.
What result does GPT 6 Astra get?
GPT 6 Astra reaches an 87.8% predicted win rate in the overall ranking. Against professional writers it scores 76.2%, and against amateurs it reaches 93.9%.
How does Vulsar evaluate the stories?
Vulsar uses a reward model trained on human preferences for creative writing. It does not use another LLM as a final judge, nor a traditional scoring rubric.
Does the benchmark prove AI writes better than humans?
It does not establish a general conclusion of that kind. The study measures a predicted preference within 475 specific tasks, and the results shift significantly depending on whether the comparison is against professional or amateur writers.

