Describe the voice, pacing, and content in natural language—then generate polished audio in seconds.