arXiv:1305.1319 · 中英对照阅读
New Alignment Methods for Discriminative Book Summarization Work in Progress
中文速览
把整本书和一篇短小的人写摘要对应起来很难,因为摘要中的一句话可能浓缩书里整段甚至整章内容,传统的词语或短语对齐方法因此不适用。研究提出了两种专为这种长文本场景设计的隐马尔可夫模型(HMM):一种把书切成可变长度的段落,用段落词汇来生成摘要句;另一种以书中的词为状态,并通过学习词之间的跳转规律来寻找对应关系。实验表明,这两种无监督对齐方法能提升抽取式图书摘要效果,尽管距离理想水平还有差距。更重要的是,它们无需人工标注就能揭示一本书中哪些段落、情节或信息最可能被摘要选中,为理解和改进长篇文学作品的自动摘要提供了基础。
摘要
We consider the unsupervised alignment of the full text of a book with a human-written summary. This presents challenges not seen in other text alignment problems, including a disparity in length and, consequent to this, a violation of the expectation that individual words and phrases should align, since large passages and chapters can be distilled into a single summary phrase. We present two new methods, based on hidden Markov models, specifically targeted to this problem, and demonstrate gains on an extractive book summarization task. While there is still much room for improvement, unsupervised alignment holds intrinsic value in offering insight into what features of a book are deemed worthy of summarization.
术语表
- unsupervised alignment
- 无监督对齐
- extractive summarization
- 抽取式摘要
- hidden Markov model (HMM)
- 隐马尔可夫模型(HMM)
- passage model
- 段落模型
- word-level alignment
- 词级对齐
- phrasal alignment
- 短语级对齐
- monolingual alignment
- 单语对齐
- comparable corpora
- 可比语料库
- document/abstract alignment
- 文档—摘要对齐
- discriminative summarization
- 判别式摘要
- unigram language model
- 一元语言模型
- Viterbi alignment
- 维特比对齐
- source document
- 源文档
- target summary
- 目标摘要
- hidden states
- 隐状态
- observation sequence
- 观测序列
- emission distribution
- 发射分布
- transition distribution
- 转移分布
- start-state distribution
- 初始状态分布
- likelihood maximization
- 似然最大化
- passage span
- 段落跨度
- sentence alignment
- 句子对齐
- Project Gutenberg
- Project Gutenberg
- Wikipedia
- Wikipedia
- Ziff-Davis corpus
- Ziff-Davis语料库
- machine translation
- 机器翻译