How to write a sound description

Describe the voices, speech, sound effects and music in the order you want to hear them. Clear roles and a short scene are a useful starting point.

1. Character

Give each speaker a stable name. Describe age range, tone, language, articulation and pace; two contrasting traits are enough to begin.

2. Dialogue

Put the speaker name before each line, quote the words to be spoken, and place delivery notes outside the quotation marks. Keep the same names throughout.

3. Sound effects

Name the sound source, distance, texture and timing: one close shutter click after the sentence is clearer than “cinematic effects”.

4. Music

Describe instruments, mood, rhythm and where the music enters or fades. Say that it stays under the dialogue if speech should lead.

5. Sequence and references

Write events in listening order; specify overlaps or pauses. Type @ in the editor to select a ready reference, then assign @voice1 to a named character. Up to three references can be selected.

Worked example · a fictional phone advertisement

[Character: Announcer (male, around thirty, clear and calm voice, modern feel, English, precise articulation)]
[Music: Airy electronic tones and futuristic synthesizer arpeggios enter slowly, followed by rhythmic electronic drums; keep them beneath the voice.]
[Dialogue: Announcer (measured, confident): "The future is in your hands, right now."]
[Dialogue: Announcer (professional and persuasive): "Yaoye X, the flagship phone with a new imaging engine. Capture the detail in every moment. A vivid display and all-day battery bring performance and design together."]
[Sound: A crisp unlock chime, a smooth screen swipe, then one close shutter click.]
[Dialogue: Announcer (calm and firm): "Yaoye. See a brighter future."]
[Music: A spacious synthesizer signature fades to silence.]

Keep the prompt within 3,000 characters. The official model documentation lists Chinese and English; other language samples are demonstrations, and results can vary. Generated audio may change the intended timing or delivery.

Based on Alibaba Cloud Model Studio documentation: Audio generation ↗ · Model specifications ↗

Open workspace →