AI Music Generator: Songs That Come Back as Separate Stems
Describe the song in words. Get vocals, drums, bass and the rest on separate tracks, ready to mix.
Open BINAI- You give
- A style, a mood, a subject: words, not notation
- You get
- A multitrack song: drums, bass, vocals and the rest, each on its own
- Plan
- All paid plans (training your own voice from Business)
- Cost
- About 94 credits for a 3-minute song with its stems
How it works
- 1
Describe the song
Write the style, the mood and who sings. Add your own lyrics, a tempo or a length if you want, or leave them to the model.
- 2
BINAI writes and splits it
The song is generated as one mix, then taken apart into stems. Quick stems arrive first, better ones replace them in place.
- 3
Mix each track
Level, pan, EQ, reverb, echo, pitch and more on every stem, in the mixer below the player.
- 4
Download
Export the mix as MP3 or WAV at streaming loudness, or every track as its own file in one ZIP.
What it does
Stems that add up to the song
Vocals, drums, bass and other. With every track on, you hear the original mix, not an approximation of it.
Three models, three jobs
MiniMax 2.6 for full songs with natural vocals, Lyria 3 Pro for rich arrangements, ElevenLabs for an exact length under a video.
Your lyrics or theirs
Paste your own lyrics and BINAI marks verse and chorus for the model. Leave them out and the model writes them.
A real mixer per track
Gain, pan, mute, solo, three-band EQ, reverb, echo, compression, pitch up to 12 semitones, low cut, warmth, stereo width and de-ess.
Record and add your own audio
Record a take from the microphone or add an audio file as a new track, then split and loop on the timeline.
Streaming loudness
Downloads come out at -14 LUFS, the level Spotify and YouTube play at, unless you switch it off.
What does BINAI's AI music generator make?
BINAI's AI music generator writes a song, a jingle or a background track from a description in plain words, and gives it back as separate stems. You describe the style, the mood, the subject and who sings. You do not need to read music or know what a stem is. What comes back is a multitrack project: vocals, drums, bass and the rest of the instruments, each on its own track in a mixer.
That is the difference from most AI song tools, which hand you one mixed file. If the drums are too loud under your voiceover, a single file leaves you stuck. With stems you turn the drums down, or the vocals off, and keep the rest.
How does BINAI produce separate tracks?
Every text-to-music model available today, including MiniMax Music, Lyria and ElevenLabs Music, returns one mixed stereo file. None of them returns stems. So BINAI does it in two steps:
- Generate the song as a mix, with the model you picked.
- Take it apart with a source separator. A fast pass delivers quick stems first, and the editor opens on those so you can start mixing. A slower, higher-quality pass then replaces them in place.
The better stems are a hybrid: vocals and drums from one separator, bass from another, and "other" calculated so that the four tracks still add up to the original song. With every track on, the editor plays the mix the model wrote, not a rebuilt version of it.
Which models can I choose?
- MiniMax 2.6, the default: full songs with natural vocals and your own lyrics, up to 5 minutes.
- Lyria 3 Pro, Google's model: rich arrangements, up to 3 minutes, the fastest.
- ElevenLabs Music: the length is exact, up to 5 minutes, which makes it the one for a background track that must end with the video.
On top of the description you can set a genre, up to four moods, a tempo between 50 and 200 BPM, the energy, a key, the vocals (none, female, male, duet or choir), the language of the lyrics, up to six instruments, the length and what the music is for: a song, background, a jingle or a podcast. BINAI turns all of that into the model's input directly, with no extra AI step in between, so the same settings always produce the same request.
What can I do in the mixer?
Each track has its own level, pan, mute and solo, plus a low, mid and high EQ, reverb and echo, a compressor, pitch shift up to 12 semitones either way at the same speed, a low cut, warmth, stereo width and a de-esser for harsh "s" sounds. You can record a take from your microphone, add an audio file as a new track, split a clip at the playhead and loop a section. The voice tools include autotune and enhance for a sung track. The multitrack editor runs in the web app at app.binai.it.
When you are done, download the mix as MP3 (320 kbps) or WAV (24-bit), or every track as its own MP3 in one ZIP, lined up from 0:00 so it opens correctly in any other editor. By default downloads are levelled to -14 LUFS, the loudness Spotify and YouTube play at. The downloaded file works under any video, including the ones you make in BINAI's editor.
How much does it cost?
A 3-minute song with its stems costs about 94 credits on Lyria 3 Pro, the cheapest model, separation included. MiniMax costs a little more, and ElevenLabs is priced by the second, so long tracks on it cost more. Plans include 1,000 (Solo), 2,400 (Pro), 4,300 (Business) or 9,000 (Studio) credits a month, shared with every other tool. Music is on every paid plan; training a model of your own voice needs Business or Studio. See pricing.
What are the honest limits?
Stem separation is very good but not perfect: a little of one instrument can bleed into another track, most often into "other". The length you ask for is exact only on ElevenLabs. And BINAI's content check cannot listen to music. For a song it measures the audio and reads the lyrics, and it gives no scores, because scoring a hook nobody heard would be misleading.
Can I use AI music on YouTube and in ads?
YouTube allows AI music, including in monetized videos, if the track does not copy an existing song or imitate a real artist, and if you disclose realistic AI music when it is the main focus of the video. Purely AI-generated music may have no copyright at all, which means you may not be able to stop others reusing it. Read can you use AI music on YouTube and who owns AI-generated content before you build a channel or an ad campaign on it.
BINAI offers no celebrity or artist voices. A voice change to your own voice is possible from the Business plan, and it starts with a recording of a consent phrase that confirms the voice is yours.
Who it is for
Small businesses
A jingle, on-hold music or a background track for a shop video, with the vocals turned down or off.
Creators and podcasters
An intro, a background bed cut to the exact length of the video, or a full song for a Short.
Marketers
Music for ads in the right mood and length, with stems so the editor can lower the drums under a voiceover.
Frequently asked questions
What is the best AI music generator with stems?
Look for one that gives you the stems, not only a mixed file. No text-to-music model returns stems on its own, so BINAI generates the song and then splits it into vocals, drums, bass and other with a source separator, and opens those tracks in a mixer where each one can be adjusted and downloaded separately.
Can AI split a song into vocals and instruments?
Yes. BINAI separates every generated song into four stems: vocals, drums, bass and other. A fast separator delivers quick stems first so you can start mixing, then a slower, higher-quality pass replaces them in place. The four stems still add up to the original mix when every track is on.
How long can an AI generated song be?
It depends on the model. MiniMax 2.6 and ElevenLabs Music go up to 5 minutes, Lyria 3 Pro up to 3. With ElevenLabs the length is exact, which makes it the right choice for a soundtrack that must end with the video. With the other two, the length you ask for is a target.
Can I use AI music from BINAI on YouTube?
YouTube allows AI music, including in monetized videos, as long as the track does not copy an existing song or imitate a real artist, and you disclose realistic AI music when it is the main focus. Purely AI-generated music may have no copyright, so others could reuse it. Our guide explains the rules.
Can I make a song in a famous singer's voice?
No. BINAI offers no celebrity voices, and imitating a real artist is a legal and platform risk. You can pick female, male, duet, choir or no vocals, and from the Business plan you can train a model of your own voice, starting with a recorded consent phrase that proves the voice is yours.
How much does an AI song cost in BINAI?
A 3-minute song with its stems costs about 94 credits on the cheapest model, Lyria 3 Pro, separation included. MiniMax costs a little more and ElevenLabs is priced by the second. Plans include 1,000 to 9,000 credits a month, shared across all tools. Music needs a paid plan, because every generation calls a paid model.