The Presentation of the First Large-Scale Text-to-Speech Transformers
Speech processing laboratories successfully adapted self-attention Transformer blocks to generate raw acoustic mel-spectrograms directly from raw text, outperforming older recurrent synthesis methods. Part of the 30 AI Roots Facts: 2018 Edition archive. HistoricallyVerified