Как проверить Duration Slider в Suno до пакетной генерации

(How to test Suno's Duration Slider before batch generation.) A canary test for length, tempo, structure and endings.
Setting a duration is easy. Getting a song that actually fits that duration is harder. A length cap may get you a 2:30 track that still sounds rushed, skips the bridge, or ends mid-phrase as if someone pulled the plug. You won't see any of that if you only check the number on the timer, and it gets expensive once you've queued forty generations with the same settings. This guide shows how to run a small, controlled comparison first. You'll make a handful of generations, measure them, and pick a combination of duration, lyric length, BPM and structure that holds up. Only after that do you scale to a batch.
1. The real problem: a length limit is not a song plan
When you set a target duration, you give the generator a constraint. It still has to decide how to spend that time. If your lyrics and requested structure need more time than you allowed, something has to give. These are the most common symptoms people report and that you should test for:
- Tempo compression. The vocal phrasing gets faster or denser than the genre suggests. Lines get crammed, syllables blur and breaths disappear.
- Section loss. A verse, the bridge, or the final chorus never shows up. Sometimes two sections merge into one.
- Broken ending. The track stops abruptly, fades during a sung line, or ends on a chord that clearly isn't a resolution.
- Filler. The opposite problem. Short lyrics in a long duration get padded with repeated choruses, long instrumental loops or improvised extra lines.
None of these failures shows up in the file length. A 2:28 track can be well paced and complete, or it can be a compressed fragment that happens to stop at 2:28. That's why you need a canary test: a small run with all variables fixed, checked against explicit criteria, before you commit to batch generation.
2. Treat the slider as a hypothesis, not a guarantee
We don't rely on any claim about how the duration control works internally, and you shouldn't either. Treat it as a black box and write down what you want to verify:
- H1, length: the output length lands within an acceptable tolerance of the target, for example ±10%.
- H2, tempo: the perceived tempo stays in the range you asked for in the style prompt and doesn't drift when the duration changes.
- H3, structure: every section you wrote appears in the intended order.
- H4, ending: the track ends with a deliberate resolution such as a final chord, a short tail or a written outro, not a cut.
Each hypothesis gets a pass/fail rule in Section 6. If a configuration fails any one of them in your canary run, it doesn't go into the batch.
3. Budget the song before you generate it
Most failures come from asking for more musical time than the duration allows. You can estimate that time with basic arithmetic before spending a credit.
3.1 The core formula
In 4/4 time, one bar lasts:
seconds_per_bar = 4 × 60 / BPM
At 90 BPM a bar lasts about 2.67 s. At 120 BPM it lasts 2.0 s. Next, estimate how many bars one lyric line takes. In much mid-tempo pop, one line fills roughly two bars, but this is an assumption that depends heavily on genre and phrasing. Rap packs more words into a bar, and ballads stretch lines out. Treat bars_per_line as a parameter you calibrate, not a constant.
vocal_seconds = lyric_lines × bars_per_line × seconds_per_bar
other_seconds = (intro_bars + instrumental_bars + outro_bars) × seconds_per_bar
planned_length = vocal_seconds + other_seconds
If planned_length is well above your target duration, expect compression or section loss. If it's well below, expect filler. Aim for the plan to land within about 10% of the target.
3.2 A small budget calculator
Save this as song_budget.py. It uses only the Python standard library.
#!/usr/bin/env python3
"""Estimate how much musical time a lyric + structure needs at a given BPM."""
import argparse
def budget(bpm, lines, bars_per_line, beats_per_bar, intro, instrumental, outro):
sec_per_bar = beats_per_bar * 60.0 / bpm
vocal = lines * bars_per_line * sec_per_bar
other = (intro + instrumental + outro) * sec_per_bar
return sec_per_bar, vocal, other, vocal + other
def fmt(seconds):
m, s = divmod(round(seconds), 60)
return f"{m}:{s:02d}"
def main():
p = argparse.ArgumentParser()
p.add_argument("--bpm", type=float, required=True)
p.add_argument("--lines", type=int, required=True, help="sung lyric lines, counting repeats")
p.add_argument("--bars-per-line", type=float, default=2.0)
p.add_argument("--beats-per-bar", type=int, default=4)
p.add_argument("--intro", type=int, default=4, help="bars")
p.add_argument("--instrumental", type=int, default=8, help="bars")
p.add_argument("--outro", type=int, default=8, help="bars")
p.add_argument("--target", type=float, help="target duration in seconds")
a = p.parse_args()
spb, vocal, other, total = budget(a.bpm, a.lines, a.bars_per_line,
a.beats_per_bar, a.intro, a.instrumental, a.outro)
print(f"seconds per bar : {spb:.2f}")
print(f"vocal time : {fmt(vocal)}")
print(f"non-vocal time : {fmt(other)}")
print(f"planned length : {fmt(total)}")
if a.target:
delta = (total - a.target) / a.target * 100
verdict = "OK" if abs(delta) <= 10 else ("TOO LONG: expect compression/cuts"
if delta > 0 else "TOO SHORT: expect filler")
print(f"vs target {fmt(a.target)} : {delta:+.0f}% -> {verdict}")
if __name__ == "__main__":
main()
Run it:
python3 song_budget.py --bpm 90 --lines 28 --target 150
With these inputs the script computes 2.67 s per bar, 28 × 2 × 2.67 ≈ 149 s of vocals, plus 20 bars × 2.67 ≈ 53 s of intro, instrumental and outro. That's about 3:22 planned against a 2:30 target, roughly +35%, so the verdict is "TOO LONG". This is plain arithmetic on assumed parameters. It shows that the lyrics don't fit, not how Suno will respond.
3.3 Count lines correctly
Count every line that will be sung. If a chorus appears three times, count it three times. The most common budgeting mistake is counting the chorus once because it's written once.
4. Worked case: a 2:30 track that kept losing its bridge
This is an illustrative scenario that shows how to reason through the protocol. It isn't a recorded experiment.
Say you need a set of 2:30 mid-tempo indie-pop tracks for short-form video. The structure is Intro → Verse 1 → Chorus → Verse 2 → Chorus → Bridge → Final Chorus → Outro. The draft lyrics have 8-line verses, 6-line choruses and a 4-line bridge, at about 90 BPM.
Sung line count: 8 + 6 + 8 + 6 + 4 + 6 = 38 lines. Run it through the budget:
python3 song_budget.py --bpm 90 --lines 38 --target 150
That gives about 203 s of vocals alone, far beyond 150 s. You now have a testable prediction: at a 2:30 cap, something will be dropped or compressed, and the bridge and final chorus are the obvious candidates because they come last. You have three ways to resolve it. You'll test them in the canary run instead of guessing:
- Option A: trim lyrics. Use 4-line verses, a 4-line chorus and a 2-line bridge: 4+4+4+4+2+4 = 22 lines → about 117 s of vocals + about 53 s of non-vocal ≈ 2:50. That's still long, so also reduce the instrumental and outro to 4 bars each (12 bars ≈ 32 s) → about 2:29.
- Option B: raise the tempo. At 112 BPM a bar is about 2.14 s. 38 lines × 2 × 2.14 ≈ 163 s of vocals is still too long. Raising tempo alone won't fix a lyric that's nearly twice too long, and it changes the genre feel.
- Option C: drop a section. Remove Verse 2 and the bridge so the shape becomes Intro → Verse → Chorus → Verse → Final Chorus → Outro, with shorter lyrics. This is a structural decision, not just a numeric one.
The arithmetic narrows the field to two or three plausible configurations. The canary test decides between them.
5. Design the canary run
5.1 Freeze everything you are not testing
Before you generate anything, write down and then stop changing:
- the model or version selector shown in your Suno interface;
- the exact style prompt text, including genre, mood, vocal type and a tempo hint such as "90 bpm";
- the exact lyric text for each lyric variant, saved as separate files;
- the structure markers you use, if any (see 5.3);
- any other toggles your UI offers, such as instrumental mode or vocal gender. Record whatever is visible.
Change one factor at a time. If you change the duration and the lyrics at once and the result improves, you won't know which change caused it.
5.2 A minimal grid
Generation is stochastic. The same inputs can produce noticeably different songs. If your interface doesn't expose a fixed random seed, you can't remove that variance, only sample it. That's why every cell in the grid gets repeated generations.
| Factor | Levels | Why |
|---|---|---|
| Target duration | your target, plus one shorter and one longer (e.g. 2:00 / 2:30 / 3:00) | shows whether failures depend on the cap |
| Lyric variant | full draft vs. budget-fitted trim | tests the arithmetic from Section 3 |
| Repeats | 3 per cell | separates a systematic failure from bad luck |
That's 3 × 2 × 3 = 18 generations, so the canary run is far smaller than a batch of fifty or a hundred. Check your own plan's credit cost per generation before you start. We don't quote prices here because they change. If 18 is too many, drop the longer duration level and keep three repeats. Cut levels before you cut repeats.
5.3 Structure markers
Suno users commonly put bracketed section labels in the lyrics, such as [Intro], [Verse 1], [Chorus], [Bridge], [Outro] and sometimes [End]. These metatags are hints, not a contract. Whether and how well they're followed is exactly what H3 and H4 test. Use the same markers in every variant so they stay a constant rather than a variable.
[Intro]
[Verse 1]
Line one of the first verse
Line two of the first verse
Line three of the first verse
Line four of the first verse
[Chorus]
Chorus line one
Chorus line two
Chorus line three
Chorus line four
[Verse 2]
...
[Chorus]
...
[Bridge]
Bridge line one
Bridge line two
[Final Chorus]
...
[Outro]
[End]
Write a distinctive word or phrase into each section, such as a place name in the bridge. Then you can tell by ear whether that section was sung, without relying on the model's labels.
5.4 Name files so the log can be joined
canary/
lyrics_full.txt
lyrics_trim.txt
style_prompt.txt
audio/
d150_trim_r1.mp3
d150_trim_r2.mp3
d150_trim_r3.mp3
d150_full_r1.mp3
...
log.csv
The pattern is d<target seconds>_<lyric variant>_r<repeat>. When you download from Suno, rename each file right away. Two clips with similar titles are easy to mix up later.
6. Pass/fail criteria, written before listening
Write the criteria down before you hear anything. Otherwise you'll quietly lower the bar for a track you happen to like.
| Hypothesis | Measure | Example pass rule (adjust to your project) |
|---|---|---|
| H1 length | file duration (ffprobe) | within ±10% of target |
| H2 tempo | estimated BPM + listening | within ±8% of the requested BPM and vocals not audibly rushed |
| H3 structure | section checklist | every marked section present, in order, with its marker word audible |
| H4 ending | tail level + listening | no cut mid-phrase; final chord or tail resolves |
A configuration passes only if at least 3 of 3 repeats pass all four checks, or 2 of 3 if your project tolerates manual rejection during the batch. Decide which rule applies before you start.
7. Measure: commands and a script
7.1 Setup
python3 -m venv .venv
. .venv/bin/activate
pip install librosa soundfile numpy
# ffprobe comes with ffmpeg; install it with your OS package manager
7.2 Exact duration
for f in canary/audio/*.mp3; do
printf "%s,%s\n" "$(basename "$f")" \
"$(ffprobe -v error -show_entries format=duration -of csv=p=0 "$f")"
done
7.3 Tempo and ending heuristics
Save this as measure.py. It reports duration, an estimated tempo, and two tail-level numbers that flag likely hard cuts. Automatic tempo estimation often lands at half or double the true tempo, and it can be misled by syncopation. Treat these numbers as hints and confirm them by ear or with tap tempo.
#!/usr/bin/env python3
"""Rough per-file metrics for a Suno canary run. Heuristics, not ground truth."""
import csv
import sys
import librosa
import numpy as np
def db(x, ref):
return 20 * np.log10(max(x, 1e-12) / max(ref, 1e-12))
def analyze(path):
y, sr = librosa.load(path, sr=None, mono=True)
duration = len(y) / sr
tempo, _ = librosa.beat.beat_track(y=y, sr=sr)
tempo = float(np.atleast_1d(tempo)[0])
rms = librosa.feature.rms(y=y)[0]
times = librosa.times_like(rms, sr=sr)
peak = float(rms.max())
tail = rms[times >= duration - 3.0]
tail_mean_db = db(float(tail.mean()), peak) # last 3 s vs loudest frame
last_frame_db = db(float(rms[-1]), peak) # very last frame vs loudest frame
# Heuristic: a natural ending usually decays; a cut often stays loud to the last frame.
suspect_cut = last_frame_db > -20.0
return {
"file": path.split("/")[-1],
"duration_s": round(duration, 2),
"tempo_est": round(tempo, 1),
"tail3s_db": round(tail_mean_db, 1),
"last_frame_db": round(last_frame_db, 1),
"suspect_cut": suspect_cut,
}
def main(paths):
w = csv.DictWriter(sys.stdout, fieldnames=["file", "duration_s", "tempo_est",
"tail3s_db", "last_frame_db", "suspect_cut"])
w.writeheader()
for p in paths:
w.writerow(analyze(p))
if __name__ == "__main__":
main(sys.argv[1:])
python3 measure.py canary/audio/*.mp3 > canary/metrics.csv
The −20 dB threshold is a starting point we picked for this protocol, not a calibrated value. Calibrate it on your own material. Find one track you're sure ends cleanly and one you're sure is cut, see where they land, and set the threshold between them.
7.4 The listening pass
The script can't tell whether the bridge was sung or whether the vocals sound rushed. Listen to each file once, start to finish, and fill in the log. Don't skip around: section loss is easy to miss if you jump straight to the end.
file,target_s,lyric_variant,duration_s,tempo_est,tempo_by_ear,rushed,intro,verse1,chorus1,verse2,chorus2,bridge,final_chorus,outro,ending,suspect_cut,pass,notes
d150_trim_r1.mp3,150,trim,,,,,,,,,,,,,,,,
d150_trim_r2.mp3,150,trim,,,,,,,,,,,,,,,,
d150_full_r1.mp3,150,full,,,,,,,,,,,,,,,,
Use y/n for section columns, and one of resolved, fade_ok, fade_mid_line, hard_cut or loop_stop for ending. These fixed values make the log easy to summarize.
7.5 Summarize by configuration
#!/usr/bin/env python3
"""Pass rate per (target, lyric_variant) from the filled-in log."""
import csv
from collections import defaultdict
cells = defaultdict(lambda: [0, 0])
with open("canary/log.csv", newline="") as fh:
for row in csv.DictReader(fh):
key = (row["target_s"], row["lyric_variant"])
cells[key][1] += 1
cells[key][0] += row["pass"].strip().lower() == "y"
for (target, variant), (ok, n) in sorted(cells.items()):
print(f"target={target:>4}s lyrics={variant:<5} pass {ok}/{n}")
8. Read the results and decide
Most canary runs end in one of these patterns. Each one points to a different fix:
- Full lyrics fail at every duration; trimmed lyrics pass at the target. The budget arithmetic was right. Use the trimmed lyrics and calibrate
bars_per_linefrom what you heard. - Trimmed lyrics pass at a longer duration but lose the final chorus at the target. The plan is still slightly too long. Remove a chorus repeat or shorten the outro, then re-run only that cell.
- Length passes but tempo drifts upward as the duration shrinks. That's compression. Reduce lyrics rather than accepting the faster feel, unless it suits the genre.
- Structure passes but endings are cut in most repeats. Try an explicit outro with instrumental bars, add an end marker, or leave more headroom between planned length and target. Test each change on its own.
- Results vary widely between repeats of the same cell. The configuration sits on a boundary. Move it clearly inside the safe zone (fewer lines or more time) instead of hoping for good luck in the batch.
Calibrate the model of your song. Using passing tracks, compute the real bars_per_line: (vocal-time-by-ear ÷ sung lines) ÷ seconds_per_bar. Put that value back into song_budget.py for the next project. That's how the arithmetic becomes yours instead of a generic guess.
9. From canary to batch
- Lock the winning configuration: duration, lyric template (line counts per section), BPM hint, structure markers and style prompt.
- Write a lyric template with line counts per section. New lyrics for the batch must match it, for example "verse = 4 lines, chorus = 4 lines, bridge = 2 lines". Changing line counts invalidates the canary result.
- Keep the pass criteria from Section 6 and apply them to batch output too. Spot-check at least the first few files and a random sample after that.
- Set a stop rule: if the batch pass rate falls clearly below the canary rate (for example, more than one failure in the first five), stop and investigate before spending the rest.
- Re-run the canary when anything upstream changes: model version, style prompt, genre, tempo range, or a visible change in the duration control.
That last rule matters most. A canary result applies to the configuration you tested, not the next model version. If output changes after an update, check whether the cause is the model or your own configuration before you rewrite the prompts. We wrote a separate guide on telling a model regression from a configuration change.
10. Failure cases and what they usually mean
| Symptom | Likely cause | First thing to try |
|---|---|---|
| Bridge or final chorus missing | planned length well above target | trim lyrics to fit budget |
| Vocals rushed, syllables crammed | too many words per bar for the tempo | fewer words per line, or fewer lines |
| Track stops loud, mid-phrase | no time left for an ending | explicit outro + headroom |
| Fade-out over a sung line | last section overruns | shorter final section, instrumental tail |
| Chorus repeated more than written | lyrics too short for duration | add a verse or shorten target |
| Long instrumental loops | lyrics too short, or instrumental bars over-requested | reduce instrumental markers |
| Estimated BPM half or double the request | beat tracker octave error | check by ear before concluding drift |
| Sections in wrong order | markers loosely followed | consistent marker names; distinctive words per section |
11. Mistakes that invalidate a canary run
- One generation per cell. You can't tell a systematic failure from randomness.
- Editing the style prompt mid-run. Every result after the edit belongs to a different experiment.
- Choosing criteria after listening. The bar slides toward whatever you happen to like.
- Trusting the tempo estimate blindly. Beat trackers make octave errors, so always confirm by ear.
- Counting the chorus once. The budget comes out short and you blame the slider.
- Judging by the first 30 seconds. The failures this test looks for happen at the end.
12. Limitations
- This protocol treats Suno as a black box. It can show whether a configuration works for you, but not why the model behaves as it does.
- Three repeats per cell is a practical minimum, not statistical proof. Rare failures will still turn up in a large batch, so keep the spot checks.
bars_per_line = 2is a starting assumption. Genres with dense or sparse phrasing need their own calibration.- The tail-level heuristic catches obvious hard cuts. It doesn't catch a quiet but musically unresolved ending. Only listening does.
- Interface elements, model versions and credit costs change. Record what you saw on the day you tested, and repeat the canary run after any change.
- Nothing here covers licensing or commercial-use terms. Check those separately in your plan's current terms.
13. One-page checklist
- Count sung lines, including repeats.
- Run
song_budget.pyagainst your target. Aim for within ±10%. - Prepare a full and a trimmed lyric variant with identical markers and a distinctive word in each section.
- Freeze the model, style prompt and toggles, and write them down.
- Write the pass criteria before generating.
- Generate the grid with 3 repeats per cell and rename the files immediately.
- Run
ffprobeandmeasure.py, then do a full listening pass intolog.csv. - Summarize the pass rate per cell and pick a configuration that passes cleanly, not one on the edge.
- Calibrate
bars_per_linefrom passing tracks. - Batch with a fixed lyric template, spot checks and a stop rule. Re-run the canary run after any upstream change.
Where to go next
For more reproducible workflows that test agents and generators before you scale them, see the Agent Lab Journal guides. Definitions of canary test, BPM, metatags and batch generation are in the glossary.