Как проверить Duration Slider в Suno до пакетной генерации

Как проверить Duration Slider в Suno до пакетной генерации
Temporary fallback cover; replace in editorial pass.

(How to test Suno's Duration Slider before batch generation.) A canary test for length, tempo, structure and endings.

Published: Reading time: 25 minutes Level: beginner, no prior experience needed

Setting a duration is easy. Getting a song that actually fits that duration is harder. A length cap may get you a 2:30 track that still sounds rushed, skips the bridge, or ends mid-phrase as if someone pulled the plug. You won't see any of that if you only check the number on the timer, and it gets expensive once you've queued forty generations with the same settings. This guide shows how to run a small, controlled comparison first. You'll make a handful of generations, measure them, and pick a combination of duration, lyric length, BPM and structure that holds up. Only after that do you scale to a batch.

1. The real problem: a length limit is not a song plan

When you set a target duration, you give the generator a constraint. It still has to decide how to spend that time. If your lyrics and requested structure need more time than you allowed, something has to give. These are the most common symptoms people report and that you should test for:

None of these failures shows up in the file length. A 2:28 track can be well paced and complete, or it can be a compressed fragment that happens to stop at 2:28. That's why you need a canary test: a small run with all variables fixed, checked against explicit criteria, before you commit to batch generation.

2. Treat the slider as a hypothesis, not a guarantee

We don't rely on any claim about how the duration control works internally, and you shouldn't either. Treat it as a black box and write down what you want to verify:

  1. H1, length: the output length lands within an acceptable tolerance of the target, for example ±10%.
  2. H2, tempo: the perceived tempo stays in the range you asked for in the style prompt and doesn't drift when the duration changes.
  3. H3, structure: every section you wrote appears in the intended order.
  4. H4, ending: the track ends with a deliberate resolution such as a final chord, a short tail or a written outro, not a cut.

Each hypothesis gets a pass/fail rule in Section 6. If a configuration fails any one of them in your canary run, it doesn't go into the batch.

3. Budget the song before you generate it

Most failures come from asking for more musical time than the duration allows. You can estimate that time with basic arithmetic before spending a credit.

3.1 The core formula

In 4/4 time, one bar lasts:

seconds_per_bar = 4 × 60 / BPM

At 90 BPM a bar lasts about 2.67 s. At 120 BPM it lasts 2.0 s. Next, estimate how many bars one lyric line takes. In much mid-tempo pop, one line fills roughly two bars, but this is an assumption that depends heavily on genre and phrasing. Rap packs more words into a bar, and ballads stretch lines out. Treat bars_per_line as a parameter you calibrate, not a constant.

vocal_seconds  = lyric_lines × bars_per_line × seconds_per_bar
other_seconds  = (intro_bars + instrumental_bars + outro_bars) × seconds_per_bar
planned_length = vocal_seconds + other_seconds

If planned_length is well above your target duration, expect compression or section loss. If it's well below, expect filler. Aim for the plan to land within about 10% of the target.

3.2 A small budget calculator

Save this as song_budget.py. It uses only the Python standard library.

#!/usr/bin/env python3
"""Estimate how much musical time a lyric + structure needs at a given BPM."""
import argparse


def budget(bpm, lines, bars_per_line, beats_per_bar, intro, instrumental, outro):
    sec_per_bar = beats_per_bar * 60.0 / bpm
    vocal = lines * bars_per_line * sec_per_bar
    other = (intro + instrumental + outro) * sec_per_bar
    return sec_per_bar, vocal, other, vocal + other


def fmt(seconds):
    m, s = divmod(round(seconds), 60)
    return f"{m}:{s:02d}"


def main():
    p = argparse.ArgumentParser()
    p.add_argument("--bpm", type=float, required=True)
    p.add_argument("--lines", type=int, required=True, help="sung lyric lines, counting repeats")
    p.add_argument("--bars-per-line", type=float, default=2.0)
    p.add_argument("--beats-per-bar", type=int, default=4)
    p.add_argument("--intro", type=int, default=4, help="bars")
    p.add_argument("--instrumental", type=int, default=8, help="bars")
    p.add_argument("--outro", type=int, default=8, help="bars")
    p.add_argument("--target", type=float, help="target duration in seconds")
    a = p.parse_args()

    spb, vocal, other, total = budget(a.bpm, a.lines, a.bars_per_line,
                                      a.beats_per_bar, a.intro, a.instrumental, a.outro)
    print(f"seconds per bar : {spb:.2f}")
    print(f"vocal time      : {fmt(vocal)}")
    print(f"non-vocal time  : {fmt(other)}")
    print(f"planned length  : {fmt(total)}")
    if a.target:
        delta = (total - a.target) / a.target * 100
        verdict = "OK" if abs(delta) <= 10 else ("TOO LONG: expect compression/cuts"
                                                 if delta > 0 else "TOO SHORT: expect filler")
        print(f"vs target {fmt(a.target)} : {delta:+.0f}%  -> {verdict}")


if __name__ == "__main__":
    main()

Run it:

python3 song_budget.py --bpm 90 --lines 28 --target 150

With these inputs the script computes 2.67 s per bar, 28 × 2 × 2.67 ≈ 149 s of vocals, plus 20 bars × 2.67 ≈ 53 s of intro, instrumental and outro. That's about 3:22 planned against a 2:30 target, roughly +35%, so the verdict is "TOO LONG". This is plain arithmetic on assumed parameters. It shows that the lyrics don't fit, not how Suno will respond.

3.3 Count lines correctly

Count every line that will be sung. If a chorus appears three times, count it three times. The most common budgeting mistake is counting the chorus once because it's written once.

4. Worked case: a 2:30 track that kept losing its bridge

This is an illustrative scenario that shows how to reason through the protocol. It isn't a recorded experiment.

Say you need a set of 2:30 mid-tempo indie-pop tracks for short-form video. The structure is Intro → Verse 1 → Chorus → Verse 2 → Chorus → Bridge → Final Chorus → Outro. The draft lyrics have 8-line verses, 6-line choruses and a 4-line bridge, at about 90 BPM.

Sung line count: 8 + 6 + 8 + 6 + 4 + 6 = 38 lines. Run it through the budget:

python3 song_budget.py --bpm 90 --lines 38 --target 150

That gives about 203 s of vocals alone, far beyond 150 s. You now have a testable prediction: at a 2:30 cap, something will be dropped or compressed, and the bridge and final chorus are the obvious candidates because they come last. You have three ways to resolve it. You'll test them in the canary run instead of guessing:

The arithmetic narrows the field to two or three plausible configurations. The canary test decides between them.

5. Design the canary run

5.1 Freeze everything you are not testing

Before you generate anything, write down and then stop changing:

Change one factor at a time. If you change the duration and the lyrics at once and the result improves, you won't know which change caused it.

5.2 A minimal grid

Generation is stochastic. The same inputs can produce noticeably different songs. If your interface doesn't expose a fixed random seed, you can't remove that variance, only sample it. That's why every cell in the grid gets repeated generations.

FactorLevelsWhy
Target durationyour target, plus one shorter and one longer (e.g. 2:00 / 2:30 / 3:00)shows whether failures depend on the cap
Lyric variantfull draft vs. budget-fitted trimtests the arithmetic from Section 3
Repeats3 per cellseparates a systematic failure from bad luck

That's 3 × 2 × 3 = 18 generations, so the canary run is far smaller than a batch of fifty or a hundred. Check your own plan's credit cost per generation before you start. We don't quote prices here because they change. If 18 is too many, drop the longer duration level and keep three repeats. Cut levels before you cut repeats.

5.3 Structure markers

Suno users commonly put bracketed section labels in the lyrics, such as [Intro], [Verse 1], [Chorus], [Bridge], [Outro] and sometimes [End]. These metatags are hints, not a contract. Whether and how well they're followed is exactly what H3 and H4 test. Use the same markers in every variant so they stay a constant rather than a variable.

[Intro]

[Verse 1]
Line one of the first verse
Line two of the first verse
Line three of the first verse
Line four of the first verse

[Chorus]
Chorus line one
Chorus line two
Chorus line three
Chorus line four

[Verse 2]
...

[Chorus]
...

[Bridge]
Bridge line one
Bridge line two

[Final Chorus]
...

[Outro]

[End]

Write a distinctive word or phrase into each section, such as a place name in the bridge. Then you can tell by ear whether that section was sung, without relying on the model's labels.

5.4 Name files so the log can be joined

canary/
  lyrics_full.txt
  lyrics_trim.txt
  style_prompt.txt
  audio/
    d150_trim_r1.mp3
    d150_trim_r2.mp3
    d150_trim_r3.mp3
    d150_full_r1.mp3
    ...
  log.csv

The pattern is d<target seconds>_<lyric variant>_r<repeat>. When you download from Suno, rename each file right away. Two clips with similar titles are easy to mix up later.

6. Pass/fail criteria, written before listening

Write the criteria down before you hear anything. Otherwise you'll quietly lower the bar for a track you happen to like.

HypothesisMeasureExample pass rule (adjust to your project)
H1 lengthfile duration (ffprobe)within ±10% of target
H2 tempoestimated BPM + listeningwithin ±8% of the requested BPM and vocals not audibly rushed
H3 structuresection checklistevery marked section present, in order, with its marker word audible
H4 endingtail level + listeningno cut mid-phrase; final chord or tail resolves

A configuration passes only if at least 3 of 3 repeats pass all four checks, or 2 of 3 if your project tolerates manual rejection during the batch. Decide which rule applies before you start.

7. Measure: commands and a script

7.1 Setup

python3 -m venv .venv
. .venv/bin/activate
pip install librosa soundfile numpy
# ffprobe comes with ffmpeg; install it with your OS package manager

7.2 Exact duration

for f in canary/audio/*.mp3; do
  printf "%s,%s\n" "$(basename "$f")" \
    "$(ffprobe -v error -show_entries format=duration -of csv=p=0 "$f")"
done

7.3 Tempo and ending heuristics

Save this as measure.py. It reports duration, an estimated tempo, and two tail-level numbers that flag likely hard cuts. Automatic tempo estimation often lands at half or double the true tempo, and it can be misled by syncopation. Treat these numbers as hints and confirm them by ear or with tap tempo.

#!/usr/bin/env python3
"""Rough per-file metrics for a Suno canary run. Heuristics, not ground truth."""
import csv
import sys

import librosa
import numpy as np


def db(x, ref):
    return 20 * np.log10(max(x, 1e-12) / max(ref, 1e-12))


def analyze(path):
    y, sr = librosa.load(path, sr=None, mono=True)
    duration = len(y) / sr
    tempo, _ = librosa.beat.beat_track(y=y, sr=sr)
    tempo = float(np.atleast_1d(tempo)[0])

    rms = librosa.feature.rms(y=y)[0]
    times = librosa.times_like(rms, sr=sr)
    peak = float(rms.max())
    tail = rms[times >= duration - 3.0]
    tail_mean_db = db(float(tail.mean()), peak)   # last 3 s vs loudest frame
    last_frame_db = db(float(rms[-1]), peak)      # very last frame vs loudest frame

    # Heuristic: a natural ending usually decays; a cut often stays loud to the last frame.
    suspect_cut = last_frame_db > -20.0
    return {
        "file": path.split("/")[-1],
        "duration_s": round(duration, 2),
        "tempo_est": round(tempo, 1),
        "tail3s_db": round(tail_mean_db, 1),
        "last_frame_db": round(last_frame_db, 1),
        "suspect_cut": suspect_cut,
    }


def main(paths):
    w = csv.DictWriter(sys.stdout, fieldnames=["file", "duration_s", "tempo_est",
                                               "tail3s_db", "last_frame_db", "suspect_cut"])
    w.writeheader()
    for p in paths:
        w.writerow(analyze(p))


if __name__ == "__main__":
    main(sys.argv[1:])
python3 measure.py canary/audio/*.mp3 > canary/metrics.csv

The −20 dB threshold is a starting point we picked for this protocol, not a calibrated value. Calibrate it on your own material. Find one track you're sure ends cleanly and one you're sure is cut, see where they land, and set the threshold between them.

7.4 The listening pass

The script can't tell whether the bridge was sung or whether the vocals sound rushed. Listen to each file once, start to finish, and fill in the log. Don't skip around: section loss is easy to miss if you jump straight to the end.

file,target_s,lyric_variant,duration_s,tempo_est,tempo_by_ear,rushed,intro,verse1,chorus1,verse2,chorus2,bridge,final_chorus,outro,ending,suspect_cut,pass,notes
d150_trim_r1.mp3,150,trim,,,,,,,,,,,,,,,,
d150_trim_r2.mp3,150,trim,,,,,,,,,,,,,,,,
d150_full_r1.mp3,150,full,,,,,,,,,,,,,,,,

Use y/n for section columns, and one of resolved, fade_ok, fade_mid_line, hard_cut or loop_stop for ending. These fixed values make the log easy to summarize.

7.5 Summarize by configuration

#!/usr/bin/env python3
"""Pass rate per (target, lyric_variant) from the filled-in log."""
import csv
from collections import defaultdict

cells = defaultdict(lambda: [0, 0])
with open("canary/log.csv", newline="") as fh:
    for row in csv.DictReader(fh):
        key = (row["target_s"], row["lyric_variant"])
        cells[key][1] += 1
        cells[key][0] += row["pass"].strip().lower() == "y"

for (target, variant), (ok, n) in sorted(cells.items()):
    print(f"target={target:>4}s  lyrics={variant:<5}  pass {ok}/{n}")

8. Read the results and decide

Most canary runs end in one of these patterns. Each one points to a different fix:

Calibrate the model of your song. Using passing tracks, compute the real bars_per_line: (vocal-time-by-ear ÷ sung lines) ÷ seconds_per_bar. Put that value back into song_budget.py for the next project. That's how the arithmetic becomes yours instead of a generic guess.

9. From canary to batch

  1. Lock the winning configuration: duration, lyric template (line counts per section), BPM hint, structure markers and style prompt.
  2. Write a lyric template with line counts per section. New lyrics for the batch must match it, for example "verse = 4 lines, chorus = 4 lines, bridge = 2 lines". Changing line counts invalidates the canary result.
  3. Keep the pass criteria from Section 6 and apply them to batch output too. Spot-check at least the first few files and a random sample after that.
  4. Set a stop rule: if the batch pass rate falls clearly below the canary rate (for example, more than one failure in the first five), stop and investigate before spending the rest.
  5. Re-run the canary when anything upstream changes: model version, style prompt, genre, tempo range, or a visible change in the duration control.

That last rule matters most. A canary result applies to the configuration you tested, not the next model version. If output changes after an update, check whether the cause is the model or your own configuration before you rewrite the prompts. We wrote a separate guide on telling a model regression from a configuration change.

10. Failure cases and what they usually mean

SymptomLikely causeFirst thing to try
Bridge or final chorus missingplanned length well above targettrim lyrics to fit budget
Vocals rushed, syllables crammedtoo many words per bar for the tempofewer words per line, or fewer lines
Track stops loud, mid-phraseno time left for an endingexplicit outro + headroom
Fade-out over a sung linelast section overrunsshorter final section, instrumental tail
Chorus repeated more than writtenlyrics too short for durationadd a verse or shorten target
Long instrumental loopslyrics too short, or instrumental bars over-requestedreduce instrumental markers
Estimated BPM half or double the requestbeat tracker octave errorcheck by ear before concluding drift
Sections in wrong ordermarkers loosely followedconsistent marker names; distinctive words per section

11. Mistakes that invalidate a canary run

12. Limitations

13. One-page checklist

  1. Count sung lines, including repeats.
  2. Run song_budget.py against your target. Aim for within ±10%.
  3. Prepare a full and a trimmed lyric variant with identical markers and a distinctive word in each section.
  4. Freeze the model, style prompt and toggles, and write them down.
  5. Write the pass criteria before generating.
  6. Generate the grid with 3 repeats per cell and rename the files immediately.
  7. Run ffprobe and measure.py, then do a full listening pass into log.csv.
  8. Summarize the pass rate per cell and pick a configuration that passes cleanly, not one on the edge.
  9. Calibrate bars_per_line from passing tracks.
  10. Batch with a fixed lyric template, spot checks and a stop rule. Re-run the canary run after any upstream change.

Where to go next

For more reproducible workflows that test agents and generators before you scale them, see the Agent Lab Journal guides. Definitions of canary test, BPM, metatags and batch generation are in the glossary.

We publish what works for us—and implement the same solutions for your business. We design AI automation, Telegram bots, chats, and AI agents for real-world processes. Discuss your project →