BookAsSumQA: An Evaluation Framework for Aspect-Based Book Summarization via Question Answering

Abstract
Aspect-based summarization aims to generate summaries that highlight specific aspects of a text, enabling more personalized and targeted summaries. However, its application to books remains unexplored due to the difficulty of constructing reference summaries for long text. To address this challenge, we propose BookAsSumQA, a QA-based evaluation framework for aspect-based book summarization. BookAsSumQA automatically generates aspect-specific QA pairs from a narrative knowledge graph to evaluate summary quality based on its question-answering performance. Our experiments using BookAsSumQA revealed that while LLM-based approaches showed higher accuracy on shorter texts, RAG-based methods become more effective as document length increases, making them more efficient and practical for aspect-based book summarization.
Type
Publication
The 14th International Joint Conference on Natural Language Processing and The 4th Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics
Note
Click the Cite button above to demo the feature to enable visitors to import publication metadata into their reference management software.
Note
Create your slides in Markdown - click the Slides button to check out the example.
Add the publication’s full text or supplementary notes here. You can use rich formatting such as including code, math, and images.

Authors
I am currently a researcher at the Japan AI Safety Institute and the University of Electro-Communications under Satoshi Hara. I received my Master’s degree from the same university under the supervision of Kei Harada.
My research interests lie in evaluation and alignment for trustworthy AI and AI Safety. At Japan AISI, I evaluate Japanese LLMs, with a particular focus on robustness to misinformation and disinformation, as well as hallucination generation. At UEC, I work on robust hallucination detection using model-internal representations, with the goal of improving methods for detecting model misalignment.
My research interests lie in evaluation and alignment for trustworthy AI and AI Safety. At Japan AISI, I evaluate Japanese LLMs, with a particular focus on robustness to misinformation and disinformation, as well as hallucination generation. At UEC, I work on robust hallucination detection using model-internal representations, with the goal of improving methods for detecting model misalignment.
Authors
Authors
Authors
Authors