|国家预印本平台
| 注册
首页|Decoupled Analysis-Judging: An Automated Creativity Evaluator Using LLMs in Complex Multi-step Creativity Tasks

Decoupled Analysis-Judging: An Automated Creativity Evaluator Using LLMs in Complex Multi-step Creativity Tasks

Wang,Xiangyu Wu,Jin Li,Xiaoyu Zheng,Chanjin Zhou,Yifeng

Decoupled Analysis-Judging: An Automated Creativity Evaluator Using LLMs in Complex Multi-step Creativity Tasks

Decoupled Analysis-Judging: An Automated Creativity Evaluator Using LLMs in Complex Multi-step Creativity Tasks

Wang,Xiangyu 1Wu,Jin 2Li,Xiaoyu 3Zheng,Chanjin 4Zhou,Yifeng5

作者信息

  • 1. Department of Educational Psychology, East China Normal University
  • 2. Shanghai Institute of Artificial Intelligence for Education, East China Normal University;School of Computer Science and Technology, East China Normal University
  • 3. School of Education and Intelligent Education Research Center, Yangzhou University
  • 4. Shanghai Institute of Artificial Intelligence for Education, East China Normal University
  • 5. School of Data Science and Engineering, East China Normal University
  • 折叠

摘要

Automated evaluation of creativity tasks remains challenging for LLM-as-a-Judge, as LLM is susceptible to biases such as verbosity bias and leniency bias. Such limitations are particularly evident in Contextually-Grounded and Procedurally-Structured Tasks (CGPST), a complex multi-step creativity task where inter-step dependencies, highly subjectivity, andwide scoring ranges lead to more unstable and biased judgments. Existing approaches either rely on task-specific training or directly apply LLM-as-a-Judge, both of which struggle to ensure reliable evaluation under such complexity. To bridge these gaps, we propose CreaEval, an automated creativity evaluator for CGPST that decouples typical LLM-as-a-Judge into analysis and judging. Correspondingly, CreaEval involves two critical phases: Memory-augmented Analysis, a SoT-LLM converts multi-step responses into structured evaluation evidence, incorporating cross-step memory; and Evidence-based Judging, a Judge-LLM uses the extracted evidence for judging without accessing raw responses. Comprehensive experiments show that CreaEval achieves an average performance improvement of 22.74% over the second-best baselines across CGPST and two classic simple creativity tasks, demonstrating its generalizability. The code is available at https://github.com/Jaong/CreaEval.

Abstract

Automated evaluation of creativity tasks remains challenging for LLM-as-a-Judge, as LLM is susceptible to biases such as verbosity bias and leniency bias. Such limitations are particularly evident in Contextually-Grounded and Procedurally-Structured Tasks (CGPST), a complex multi-step creativity task where inter-step dependencies, highly subjectivity, and wide scoring ranges lead to more unstable and biased judgments. Existing approaches either rely on task-specific training or directly apply LLM-as-a-Judge, both of which struggle to ensure reliable evaluation under such complexity. To bridge these gaps, we propose CreaEval, an automated creativity evaluator for CGPST that decouples typical LLM-as-a-Judge into analysis and judging. Correspondingly, CreaEval involves two critical phases: Memory-augmented Analysis, a SoT-LLM converts multi-step responses into structured evaluation evidence, incorporating cross-step memory; and Evidence-based Judging, a Judge-LLM uses the extracted evidence for judging without accessing raw responses. Comprehensive experiments show that CreaEval achieves an average performance improvement of 22.74% over the second-best baselines across CGPST and two classic simple creativity tasks, demonstrating its generalizability. The code is available at https://github.com/Jaong/CreaEval.

关键词

LLM-as-a-Judge/Creativity task/Automated evaluation

Key words

LLM-as-a-Judge/Creativity task/Automated evaluation

引用本文复制引用

Wang,Xiangyu,Wu,Jin,Li,Xiaoyu,Zheng,Chanjin,Zhou,Yifeng.Decoupled Analysis-Judging: An Automated Creativity Evaluator Using LLMs in Complex Multi-step Creativity Tasks[EB/OL].(2026-09-03)[2026-09-04].https://chinaxiv.org/abs/202609.00026.

学科分类

TP3
首发时间 2026-09-03
下载量:0
|
点击量:9
段落导航相关论文