A new creative-writing benchmark suggests that advanced artificial intelligence models are now performing at a level comparable to, and in one case slightly above, amateur human writers.
What did the creative-writing benchmark test?
Vulsar AI tested 24 AI models against human writers using 475 prompts. GPT 6 Astra recorded an 87.8% predicted win rate, narrowly ahead of the combined amateur human-writer score of 86.6%. GPT 5.6 Sol followed with a 77.6% predicted win rate.
Did AI outperform professional writers?
No. The results do not show that AI has surpassed professional writers. Vulsar’s separate findings indicated that professional writers performed better than the amateur human baseline.
Why writing length matters
The benchmark also found a notable difference in output length. Human writers produced an average of 2,592 tokens per prompt, compared with 1,537 tokens for GPT 6 Astra.
That difference suggests that writing quality cannot be judged by win rates alone. Style, originality, detail, context and the reader’s expectations also influence how creative work is evaluated.
What the results mean
The findings suggest that the gap between advanced AI and casual human writing is becoming smaller. However, skilled professional writers remain a tougher challenge for current AI systems, and the benchmark should not be interpreted as proof that AI can replace human creativity.