Choose what the listener should hear
Describe your idea freely, or start with an optional editable template for narration, dialogue, podcasts, sound scenes, sound effects or ambience. A spoken explanation needs a different prompt from an empty forest at dawn. Describe the essential event first, then add the voice, mood and acoustic setting. A focused request gives you a clearer basis for judging the result.
For speech, write the words you want spoken and use consistent speaker names. For sound-only clips, say that no speech is wanted and describe the sequence of events. Keep directions readable rather than adding many conflicting adjectives. The combined input, including selected character directions, has a 3,000-character limit per request.
Listen, compare and revise one thing
Generate a first version, then listen for pronunciation, pacing, speaker changes and unexpected sounds. Keep an earlier version available while trying a new direction. If a pause feels rushed, refine the performance request; if a scene feels crowded, reduce its events. Changing one important detail at a time makes your version comparison more useful.
Before generating, write down the detail that would make this clip successful. It might be an understandable question, a restrained emotional turn or a clean sound-only ending. That small listening goal helps you decide whether to keep the version, rewrite the prompt or move on to the next segment.
Build a complete piece from short clips
Divide longer ideas into segments you can review independently. The model supports podcast content up to four minutes and other audio up to two minutes. These are upper limits, so the result can be shorter. Arrange chosen versions in order, adjust pauses between segments and export up to 30 minutes including pauses as WAV or MP3. A clip ZIP contains separate processed WAV clips, the script and an arrangement manifest for further editing.