Cloning a track and creating one from scratch are different paths. musicaclone keeps them separate and blocks the things that commonly go wrong along the way.

You provide a link (YouTube, Facebook, TikTok, anything yt-dlp can open). It downloads the audio, asks Gemini whether itβs actually music, extracts the style and lyrics, puts together a spec, shows it to youβand only then generates it in Suno via the Kie API.
Use the original audio as a reference (upload-cover), so the melody and structure remain. It's not "a similar song".
Not music? It stops. Lyrics cut off? It warns you. Source is a 29 s teaser? It shows the duration before you spend a credit.
Each track becomes a folder with spec.json: taskId recorded before polling, actual cost measured, and regen to adjust without reanalyzing.
The top path is cloning (starting from a link). The bottom path is creation (starting from an idea). Both pass through the same gate before spending a credit.
prep + cloneThe original audio becomes the reference in Kie. The weight --audio-weight determines how closely the new track sticks to the original: 0.9 is almost the same arrangement, 0.4 is only inspired by it.
cria + regenNo reference. You provide style tags and the lyrics; the model does the rest. This is the way to create a soundtrack, jingle, or original theme.
Nothing runs as a service. It's a bash script with curl, jq, yt-dlp and an analysis Python script.
yt-dlp for downloading, jq for state, ffmpeg for audio extraction.
# debian/ubuntu sudo apt install jq ffmpeg pipx install yt-dlp
It generates the music. Get it at kie.ai/api-key and put it in a .env β the script reads it at runtime; itβs never version-controlled.
# .env next to the script KIE_API_KEY=sk-...
It listens to the audio and tells you whether itβs music, what the style is, and what the lyrics are.
# same .env GOOGLE_API_KEY=...
Real commands. The only step that uses credits is step 3βand it only runs after you review step 2.
Resolves the link, downloads the audio, classifies and analyzes it. Doesn't use Kie credits.
bash musica.sh prep "https://youtube.com/watch?v=..." minha-faixa
Here you can check what it understood. If the duration is teaser-length or the lyrics are incomplete, stop and fix it first.
bash musica.sh spec minha-faixa --- viking-gods [clone] --- source: βViking Godsβ by Bob Dominator duration: 29.07s <- teaser, not the whole song mode: cover style: epic pop, cinematic, female lead vocals, nordic, driving rhythm complete lyrics? false (75 words)
Upload the reference, start the job, and record the taskId before any poll β if the connection drops, you can resume instead of paying again.
bash musica.sh clone minha-faixa --audio-weight 0.85
Suno URLs expire, so get downloads it immediately and also calculates the actual generation cost.
bash musica.sh status minha-faixa # PENDING β TEXT_SUCCESS β FIRST_SUCCESS β SUCCESS bash musica.sh get minha-faixa # faixa-1.mp3, faixa-2.mp3, credits spent: 12
No link. You write the style in tags and deliver the lyrics in a file with [Verse] / [Chorus].
bash musica.sh cria meu-tema \ --style "epic cinematic pop, powerful female belt vocals, war drums, 118bpm" \ --title "Se Paga" --voz f --letra letra.txt bash musica.sh regen meu-tema
Everything lives in the spec.json, so regen redo it with the adjustment without repeating the download or analysis.
| Flag | What it does | When to adjust |
|---|---|---|
| --style | sound tags (genre, instrument, vocals, tempo) | the timbre came out wrong |
| --negative | what to avoid | It came out acoustic, and you wanted something heavier |
| --voz m|f | vocal gender | vocal came in switched |
| --audio-weight | only in the clone: how much it draws from the reference | 0.85β0.95 sticks close to the original, 0.4 is just inspiration |
| --style-weight | adherence to the described style | ignored the tags |
| --weirdness | creative deviation | it was too obvious |
| --model | V4 β¦ V5_5 (default V5) | V5_5 is the only one with --duration |
| --instrumental | no vocals | soundtrack and BGM |
A good style is a short tag in English, separated by commas: nordic folk metal, epic, male gang vocals, war drums, low brass, 140bpm. A continuous phrase works worse.
The script records the balance before generating and writes the actual cost to spec.json. The number below comes from a real generation on 2026-08-05, not a third-party chart.
| Service | Price | Does it clone audio? |
|---|---|---|
| Suno via Kie (this project) | 12 credits per generation of 2 tracks β US$ 0,06 (~US$ 0,03/track) | Yes β upload-cover |
| Suno direct (subscription) | 5 credits/song; Pro US$ 10/month β 500 songs | Web only, no official API |
| ElevenLabs Music | ~US$ 0.30β0.40 per generated minute (3 min β US$ 1.00β1.20) | No |
| Udio | credits per generation; Standard US$ 10/month = 2,400 credits | No mature public API |
Summary: Kie wins here because it's the only one of the three with cloning from reference audio, and costs ~17β20Γ less than ElevenLabs per track.
Everything marked as ready was run end to end, including a real paid generation.