Texture Synthesis Using Convolutional Neural Networks
Leon A. Gatys Centre for Integrative Neuroscience, University of Tübingen, Germany Bernstein Center for Computational Neuroscience, Tübingen, Germany Graduate School of Neural Information Processing, University of Tübingen, Germany、Alexander S. Ecker Centre for Integrative Neuroscience, University of Tübingen, Germany Bernstein Center for Computational Neuroscience, Tübingen, Germany Max Planck Institute for Biological Cybernetics, Tübingen, Germany Baylor College of Medicine, Houston, TX, USA、Matthias Bethge Centre for Integrative Neuroscience, University of Tübingen, Germany Bernstein Center for Computational Neuroscience, Tübingen, Germany Max Planck Institute for Biological Cybernetics, Tübingen, Germany
Abstract
Here we introduce a new model of natural textures based on the feature spaces of convolutional neural networks optimised for object recognition. Samples from the model are of high perceptual quality demonstrating the generative power of neural networks trained in a purely discriminative fashion. Within the model, textures are represented by the correlations between feature maps in several layers of the network. We show that across layers the texture representations increasingly capture the statistical properties of natural images while making object information more and more explicit. The model provides a new tool to generate stimuli for neuroscience and might offer insights into the deep representations learned by convolutional neural networks.
中文速览
用卷积神经网络(Convolutional Neural Network, CNN)的特征空间来建模自然纹理,长期以来缺乏一个既能捕捉复杂统计结构、又有理论解释力的参数化方案。研究者借助专为目标识别训练的VGG-19网络,将纹理定义为网络各层特征图之间的相关性(即Gram矩阵),再从白噪声出发通过梯度下降生成与原始纹理Gram矩阵匹配的新图像。实验表明,随着所用网络层数的加深,合成纹理的视觉质量逐步提升,使用到第四个池化层时效果已几乎与原图难以区分,且性能优于此前最好的手工统计模型Portilla-Simoncelli;此外,即便将模型参数压缩约80倍,感知质量也几乎不受影响。这项工作揭示了判别式训练的深度网络同样具备强大的生成能力,为神经科学实验刺激的制作和理解深层视觉表征提供了新工具。
原文 arXiv:1505.07376;中英对照 + 大白话阅读 https://aha.fim.ai/paper/1505.07376v3