PTENES
CLI Β· audio Β· Suno via Kie

Send the link. It understands the music before trying again.

Cloning a track and creating one from scratch are different paths. musicaclone keeps them separate and blocks the things that commonly go wrong along the way.

musicaclone project cover
What it is

A one-command-per-step CLI, with a gate before costly operations

You provide a link (YouTube, Facebook, TikTok, anything yt-dlp can open). It downloads the audio, asks Gemini whether it’s actually music, extracts the style and lyrics, puts together a spec, shows it to youβ€”and only then generates it in Suno via the Kie API.

🎧 A true clone

Use the original audio as a reference (upload-cover), so the melody and structure remain. It's not "a similar song".

πŸ›‘ Sanity checks

Not music? It stops. Lyrics cut off? It warns you. Source is a 29 s teaser? It shows the duration before you spend a credit.

πŸ“ State on disk

Each track becomes a folder with spec.json: taskId recorded before polling, actual cost measured, and regen to adjust without reanalyzing.

How it works

Two paths, one gate in the middle

The top path is cloning (starting from a link). The bottom path is creation (starting from an idea). Both pass through the same gate before spending a credit.

linkβ†’ yt-dlp: resolves + downloadsβ†’ Gemini: is it music?β†’ style + lyricsβ†’ spec.jsonβ†’ πŸ›‘ your approvalβ†’ reference uploadβ†’ Suno (Kie)β†’ pollβ†’ 2 MP3 tracks

Clone β€” prep + clone

The original audio becomes the reference in Kie. The weight --audio-weight determines how closely the new track sticks to the original: 0.9 is almost the same arrangement, 0.4 is only inspired by it.

Creation β€” cria + regen

No reference. You provide style tags and the lyrics; the model does the rest. This is the way to create a soundtrack, jingle, or original theme.

Prerequisites

Three keys and two binaries

Nothing runs as a service. It's a bash script with curl, jq, yt-dlp and an analysis Python script.

Binaries

yt-dlp for downloading, jq for state, ffmpeg for audio extraction.

# debian/ubuntu
sudo apt install jq ffmpeg
pipx install yt-dlp

Kie key

It generates the music. Get it at kie.ai/api-key and put it in a .env β€” the script reads it at runtime; it’s never version-controlled.

# .env next to the script
KIE_API_KEY=sk-...

Gemini key

It listens to the audio and tells you whether it’s music, what the style is, and what the lyrics are.

# same .env
GOOGLE_API_KEY=...
User Guide Β· step by step

From link to MP3

Real commands. The only step that uses credits is step 3β€”and it only runs after you review step 2.

1

Prepare the track

Resolves the link, downloads the audio, classifies and analyzes it. Doesn't use Kie credits.

bash musica.sh prep "https://youtube.com/watch?v=..." minha-faixa
2

Check the gate

Here you can check what it understood. If the duration is teaser-length or the lyrics are incomplete, stop and fix it first.

bash musica.sh spec minha-faixa

--- viking-gods  [clone] ---
source:    β€œViking Gods” by Bob Dominator
duration:  29.07s          <- teaser, not the whole song
mode:     cover
style:    epic pop, cinematic, female lead vocals, nordic, driving rhythm
complete lyrics? false  (75 words)
3

Generate (this uses credits)

Upload the reference, start the job, and record the taskId before any poll β€” if the connection drops, you can resume instead of paying again.

bash musica.sh clone minha-faixa --audio-weight 0.85
4

Track and download

Suno URLs expire, so get downloads it immediately and also calculates the actual generation cost.

bash musica.sh status minha-faixa   # PENDING β†’ TEXT_SUCCESS β†’ FIRST_SUCCESS β†’ SUCCESS
bash musica.sh get minha-faixa      # faixa-1.mp3, faixa-2.mp3, credits spent: 12
5

Create from scratch (the other path)

No link. You write the style in tags and deliver the lyrics in a file with [Verse] / [Chorus].

bash musica.sh cria meu-tema \
  --style "epic cinematic pop, powerful female belt vocals, war drums, 118bpm" \
  --title "Se Paga" --voz f --letra letra.txt
bash musica.sh regen meu-tema
Adjustments

What you can tweak when the result comes out wrong

Everything lives in the spec.json, so regen redo it with the adjustment without repeating the download or analysis.

FlagWhat it doesWhen to adjust
--stylesound tags (genre, instrument, vocals, tempo)the timbre came out wrong
--negativewhat to avoidIt came out acoustic, and you wanted something heavier
--voz m|fvocal gendervocal came in switched
--audio-weightonly in the clone: how much it draws from the reference0.85–0.95 sticks close to the original, 0.4 is just inspiration
--style-weightadherence to the described styleignored the tags
--weirdnesscreative deviationit was too obvious
--modelV4 … V5_5 (default V5)V5_5 is the only one with --duration
--instrumentalno vocalssoundtrack and BGM

A good style is a short tag in English, separated by commas: nordic folk metal, epic, male gang vocals, war drums, low brass, 140bpm. A continuous phrase works worse.

Costs

Measured, not estimated

The script records the balance before generating and writes the actual cost to spec.json. The number below comes from a real generation on 2026-08-05, not a third-party chart.

ServicePriceDoes it clone audio?
Suno via Kie (this project)12 credits per generation of 2 tracks β‰ˆ US$ 0,06 (~US$ 0,03/track)Yes β€” upload-cover
Suno direct (subscription)5 credits/song; Pro US$ 10/month β‰ˆ 500 songsWeb only, no official API
ElevenLabs Music~US$ 0.30–0.40 per generated minute (3 min β‰ˆ US$ 1.00–1.20)No
Udiocredits per generation; Standard US$ 10/month = 2,400 creditsNo mature public API

Summary: Kie wins here because it's the only one of the three with cloning from reference audio, and costs ~17–20Γ— less than ElevenLabs per track.

Roadmap

What exists and what's coming

Everything marked as ready was run end to end, including a real paid generation.

Ready
Cloning and creating, with gatesprep, spec, clone, creates, regen, status, get, balance. Cost measured automatically.
Ready
Delivery via Telegramenviar.sh sends the MP3 directly to the bot.
Next
PersonaReuse the same voice across tracks. The endpoint still needs to be confirmed in the Kie API.
Next
Extend and stemsExtend a good track and separate the instrumental from the vocals.
Afterward
Music videoConnect the output to INEMA's video pipeline.