Balacoon's neural text-to-speech model can generate human-quality speech in real-time using only a conventional CPU, without the need for expensive GPU acceleration.
Also worth reading: How effective is Descript for long-form text-to-voice conversion? · How can I use Jarvis TTS to generate text-to-speech audio in Paul Bettany's voice? · What are the best alternatives to Storyline and Camtasia for text-to-speech functionality?