Write for a listener who cannot see the scene
Start with the action that matters: someone arrives, a familiar place sounds different, or a character makes a choice. Give the listener enough context to follow without describing every visible detail. Use short spoken sentences and read them aloud before generating. If a line depends on a gesture, rewrite it so the meaning comes through in speech, narration or an audible event.
Separate narration from character dialogue with consistent labels. Give each scene a small dramatic purpose rather than filling one prompt with the whole plot. Foleyix accepts up to 3,000 characters of combined input per generation. Smaller scenes leave room for character directions and make it easier to revise the moment that needs attention.
Give voices and sounds different jobs
Use synthetic characters to make exchanges distinguishable, then audition their saved samples. Up to three ready characters can be selected in each segment. Describe sound details that affect the story: a door that stops a conversation, footsteps that approach or rain that makes a room feel enclosed. A sound scene requests speech and setting together; separate effect or ambience segments provide individual clips you can place in sequence.
Find the story’s rhythm in the edit
Listen for intelligibility first, then performance and atmosphere. Compare versions of the same segment and select the take that serves the scene. Arrange clips in order and adjust pauses at changes of place, speaker or mood. The model supports podcast content up to four minutes and other audio up to two minutes; each episode export is up to 30 minutes including pauses. Export WAV or MP3, or choose a ZIP of processed separate WAV clips with the script and arrangement manifest. Review the finished sequence before sharing.