Meloria.aiMeloria.aiUnlimited music
Meloria.ai · Developer API

One API key, eight AI models

Generate instrumental or vocal music, singing-photo music videos, multi-scene AI music videos (MV Pro), dance videos, fashion review videos, 2-minute talking videos, multi-episode podcasts and long text-to-speech MP3 — from your own app, with the same credits you already use on Meloria.

What you can build

🎵
MusicInstrumental or with lyrics, 30 s – 10 min, MP3 up to 320 kbps. Describe a style or send your own lyrics.
1 song credit / track
🎤
Music video (MV)One photo + one song → the character sings in sync. 9:16, 1:1, 4:5 or 16:9, up to 40 s, optional HD.
≈ 1 credit / second · min 6
🎬
MV ProA song + its lyrics → AI analyses the vocals, writes a scene-by-scene script, draws every scene and renders a full lip-synced music video (up to 5 min).
1 song credit / scene + ≈ 2 credits / second
💃
Dance videoOne photo + a dance reference video → the character dances the same moves, optionally to your own song.
≈ 1.3 credit / second · min 6
👗
Fashion reviewOne product photo + music → a virtual model shows the outfit or accessory to the beat. Great for shops.
≈ 1 credit / second · min 6
🗣️
Talking videoOne photo + text (30 languages) → the person reads it to camera, lip-synced, up to 2 minutes.
0.1 credit / s voice + ≈ 1 credit / s video
🎙️
PodcastA topic → AI writes and reads a multi-episode series (1–30 min per episode) in a chosen voice. MP3 per episode.
0.05 credit / second
🎧
Text to speechLong text (up to 30 minutes ≈ 25,000 characters) → one MP3, natural voices in 30 languages, optional voice clone.
0.05 credit / second

Prices use the same wallet as the web app: song credits for music and video credits (1 credit = 1,000₫ ≈ $0.04) for everything else. See Pricing & limits.

1. Get an API key

Sign in with your Meloria account, then create a key. Keep it secret (you can reveal it again here or on your account page). Creating a new key replaces the old one; revoking stops all API access immediately.

Sign in to manage your API key.

2. Quick start

Every generation is a job: create it with one POST, then poll GET /api/v1/jobs/{id} every few seconds until status is done. Results are plain MP3/MP4 URLs.

# 1) create a 60-second lo-fi track (instrumental)
curl -X POST https://meloria.ai/api/v1/music \
  -H "Authorization: Bearer mk_live_xxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxx" \
  -H "Content-Type: application/json" \
  -d '{"prompt":"calm lo-fi hip hop, warm piano, vinyl crackle, rain","instrumental":true,"duration":60}'

# → {"ok":true,"id":"j9f1c2…","model":"music","status":"queued","progress":0,…}

# 2) poll until done
curl https://meloria.ai/api/v1/jobs/j9f1c2… -H "Authorization: Bearer mk_live_…"

# → {"ok":true,"status":"done","progress":100,"result":{"tracks":[{"audio_url":"https://meloria.ai/mp3/music/….mp3","title":"…","duration":60}]}}

Python

import requests, time
H = {"Authorization": "Bearer mk_live_…"}
job = requests.post("https://meloria.ai/api/v1/music", headers=H,
                    json={"prompt": "uplifting pop, female vocal", "lyrics": "…", "duration": 120}).json()
while True:
    j = requests.get(f"https://meloria.ai/api/v1/jobs/{job['id']}", headers=H).json()
    if j["status"] in ("done", "error"): break
    time.sleep(5)
print(j)

3. Authentication

Send your key in the Authorization header on every request (or in X-API-Key). Keys start with mk_live_. Never embed a key in client-side code — call the API from your server.

Authorization: Bearer mk_live_xxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxx
# or
X-API-Key: mk_live_xxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxx

Base URL: https://meloria.ai/api/v1. Request bodies are JSON (Content-Type: application/json) or multipart/form-data when files are uploaded. All responses are JSON and include ok.

4. Jobs & polling

GET/api/v1/jobs/{id}
FieldDescription
statusqueued → running → done | error. Keep polling while queued/running (every 3–10 s; results are cached for 2 s).
progress0–100 (approximate).
phaseMulti-step models only: voice, scenes, render, post (talking video); planning, episode_N (podcast).
resultPresent when done. Shape depends on the model (see each model below). File URLs are public and stay online for at least 7 days — download what you need to keep.
costCredits actually charged (settled when the job finishes; failed jobs are refunded automatically).
errorError code when status = error.
GET/api/v1/jobs

Your 20 most recent API jobs (no result refresh — poll individual jobs for live status).

🎵 Music — instrumental or with lyrics

POST/api/v1/music
ParameterTypeDescription
promptstringStyle / mood / instruments description. Required unless lyrics is given.
lyricsstringYour lyrics (≤ 3,000 chars). Omit for instrumental. Use [Verse], [Chorus] tags for structure.
instrumentalbooltrue = no vocals (lyrics ignored). Default false.
durationintSeconds, 30–600. Omit or "auto" to follow the lyrics.
vocal_languagestringauto (default), vi, en, ja, ko, zh, es, fr…
vocal_genderstringauto | male | female
titlestringOptional title (≤ 80 chars).
nintNumber of variations, 1–2 (each costs 1 song). Default 1.
bitratestring128k | 192k | 256k | 320k — capped by your plan.
bpm, key_scale, seedOptional advanced controls (e.g. bpm: 90, key_scale: "A minor").

Result: result.tracks[] with audio_url, title, lyrics, duration, genre, bpm, instrumental.

curl -X POST https://meloria.ai/api/v1/music -H "Authorization: Bearer mk_live_…" -H "Content-Type: application/json" -d '{
  "prompt": "Vietnamese ballad, acoustic guitar, emotional female vocal",
  "lyrics": "[Verse]\nĐêm nay phố vắng…\n[Chorus]\n…",
  "vocal_language": "vi", "vocal_gender": "female", "duration": "auto", "bitrate": "320k"
}'

🎤 Music video (MV) — singing photo

POST/api/v1/mv multipart/form-data
ParameterTypeDescription
imagefileRequired. JPG/PNG ≤ 8 MB. Front-facing, face clearly visible (the mouth is animated).
audiofileRequired. MP3/M4A/WAV ≤ 30 MB. Up to 40 s is used (first 40 s of longer files).
image2fileOptional second photo → two scenes.
durationintSeconds 5–40. Ignored when the audio is ≤ 40 s (its real length is used).
aspectstring9:16 (default) | 1:1 | 4:5 | 16:9
hdbool1080p output (price × 1.7). Default 720p.
stylestringstudio (default) | stage | duet | mv | cartoon | anime | mascot | ink
promptstringExtra scene description (≤ 500 chars), e.g. lighting, camera, background.
negativestringThings to avoid (≤ 300 chars).
lipsyncboolfalse = move with the music without singing. Default true.
singerstringWhen several people are in the photo: left | center | right.

Result: result.video_url (MP4). The job may take 3–8 minutes.

curl -X POST https://meloria.ai/api/v1/mv -H "Authorization: Bearer mk_live_…" \
  -F "[email protected]" -F "[email protected]" -F "aspect=9:16" -F "style=stage" -F "hd=1"

🎬 MV Pro — multi-scene AI music video from a song

POST/api/v1/mvpro multipart/form-data
ParameterTypeDescription
audiofileThe song (MP3/WAV/M4A ≤ 60 MB, ≤ 5 minutes). Required unless track_id is given.
lyricsstringRequired with audio. The sung lyrics — scenes are cut and captioned from them.
track_idintInstead of uploading: id of a track you generated with /api/v1/music (lyrics are taken from it).
title, genrestringOptional metadata used by the script writer.
imagefileOptional reference photo of the singer (face + shoulders) — every singing scene is drawn with this character.
ideastringStory / setting hint for the script (≤ 600 chars), e.g. "rainy Hanoi at night, lost love".
stylestringreal (default) | anime | 3d | paint | cine
aspectstring16:9 (default) | 9:16 | 4:5
langstringLanguage for captions/synopsis (default en).

Phases: analyze (vocal/instrumental segmentation) → images (script + one AI image per scene; scenes_ready/scenes show progress) → render. Typically 10–25 scenes and 10–20 minutes for a 3-minute song.

Result: result.video_url (MP4), result.title, result.scenes, result.song_credits_used. Cost: 1 song credit per scene image + video credits ⌈seconds × rate × 2⌉ (min 12).

curl -X POST https://meloria.ai/api/v1/mvpro -H "Authorization: Bearer mk_live_…" \
  -F "[email protected]" -F "lyrics=<song.txt" -F "title=Đêm Hà Nội" -F "[email protected]" \
  -F "style=cine" -F "aspect=16:9" -F "idea=rainy city at night, lost love"

💃 Dance video

POST/api/v1/dance multipart/form-data
ParameterTypeDescription
imagefileRequired. Full-body or half-body photo of the character, ≤ 8 MB.
dance_videofileReference dance clip (MP4 ≤ 60 MB). The moves — and, without audio, the sound — are taken from it. Required unless template is set.
templatestringName of a built-in dance template instead of uploading a clip (as listed on the web app).
audiofileOptional song to dance to (replaces the clip's sound).
start, endfloatCut points in the reference clip (seconds). Max 40 s are rendered.
dancersstringone (default) | all
aspect, hdSame as MV.

Result: result.video_url (MP4). Price ≈ MV × 1.3.

👗 Fashion review

POST/api/v1/fashion multipart/form-data
ParameterTypeDescription
imagefileRequired. Product photo: a model wearing the outfit, or the item itself (clothes, bag, shoes, jewellery), ≤ 8 MB.
audiofileRequired. Background music (≤ 40 s used).
productstringclothes | acc (accessory). Auto-detected when omitted.
lipsyncboolDefault false for fashion (model poses without singing).
aspect, hd, duration, promptSame as MV.

Result: result.video_url (MP4) — ready for TikTok, Reels, Shopee.

🗣️ Talking video — a photo reads your text (≤ 2 min)

POST/api/v1/talk multipart/form-data
ParameterTypeDescription
imagefileRequired. Portrait photo, face visible, ≤ 8 MB.
textstringRequired. What to say, ≤ 3,000 chars. Speech longer than 120 s is rejected (too_long) — about 1,500 characters at normal speed.
langstringSpeech language: vi, en, ja, ko, zh, fr, de, es, pt, ru, id, th, hi, ar, it, nl, pl, tr… (30). Default en.
voicestringVoice id from GET /api/v1/voices?lang=…. Default: the language's default voice.
speedfloat0.6 – 1.6 (default 1.0)
aspectstring9:16 (default) | 1:1 | 4:5 | 16:9
stylestringreal (default) | anime | 3d | paint | cine
hdbool1080p (price × 1.7).

The job runs in phases: voice (text-to-speech) → scenes → render. Scenes longer than ~22 s are split automatically and joined seamlessly.

Result: result.video_url (MP4), result.audio_url (the narration MP3), result.duration.

curl -X POST https://meloria.ai/api/v1/talk -H "Authorization: Bearer mk_live_…" \
  -F "[email protected]" -F "lang=en" -F "aspect=16:9" -F "hd=1" \
  -F "text=Hi everyone, welcome to today's product update. In the next two minutes I'll show you…"

🎙️ Podcast — AI-written, AI-read series

POST/api/v1/podcast
ParameterTypeDescription
topicstringRequired. What the series is about (≤ 300 chars).
titlestringSeries title (≤ 80 chars, optional).
kindstringpodcast (default) | story | ghost | confess | philo | buddha | history | sleep | kids | motiv
episodesint1–10 episodes per request (default 1). Episodes are produced one after another so the story stays consistent.
minutesintTarget length per episode, 1–30 (default 5).
langstringWriting & speaking language (default en).
voicestringVoice id from GET /api/v1/voices?lang=…. Default: the language's default voice.

Result: result.title, result.episodes[] with n, title, status, audio_url (MP3), duration. While running, phase is planning then episode_N, and finished episodes already appear in result.

curl -X POST https://meloria.ai/api/v1/podcast -H "Authorization: Bearer mk_live_…" -H "Content-Type: application/json" \
  -d '{"topic":"Simple habits that improve sleep","kind":"podcast","episodes":3,"minutes":5,"lang":"en"}'

🎧 Text to speech — long text → MP3

POST/api/v1/tts
ParameterTypeDescription
textstringRequired. Up to ≈ 25,000 characters (≈ 30 minutes of speech).
langstringLanguage code (30 supported, default en).
voicestringVoice id from GET /api/v1/voices?lang=…. Default: the language's default voice.
speedfloat0.6 – 1.6 (default 1.0)
tonestringstory for a softer storytelling delivery; omit for neutral.

Result: result.audio_url (MP3, 1 file), result.duration. Files are kept 7 days.

curl -X POST https://meloria.ai/api/v1/tts -H "Authorization: Bearer mk_live_…" -H "Content-Type: application/json" \
  -d '{"text":"Chapter one. It was a bright cold day in April…","lang":"en","voice":"vx:en_f_story|Emma|f","speed":1.0,"tone":"story"}'

Account & voices

GET/api/v1/me

Your balances: songs_remaining, video_credits, video_rate (credits per second for video, depends on your last pack), plan.

GET/api/v1/voices?lang=vi

Voices for a language: voices[] = {id, name, gender}. Pass id as voice to talk / podcast / tts. Special id kid = child voice.

Errors

Errors return {"ok":false,"error":"code"} with an HTTP status:

HTTPerrorDescription
401invalid_api_keyMissing, malformed or revoked key.
402out_of_credits / no_video_creditsNot enough song credits / video credits. Top up at /pricing (web) or in the app.
403account_restrictedAccount locked for policy violations.
404job_not_foundUnknown job id (jobs are private to the key owner).
413*_too_largeUpload over the size limit.
422*_required, too_long, prompt_blocked, nsfw, face_too_small…Invalid input — the code says which field. Adult content and blocked topics are rejected and may lead to a strike.
429daily_limit, rate_limitedPlan limit for today reached, or too many requests per minute (music: 5/min).
502upstream_errorEngine temporarily unavailable — retry later; nothing was charged.

Pricing & limits

The API uses the same wallet as meloria.ai and the Meloria app — no separate API plan. Prices are the web prices; nothing extra for API access.

ModelCostLimits
Music1 song credit per track (from your plan quota or a song pack)30 s – 10 min · ≤ 2 variations per call · plan daily cap · 5 calls/min
Music video (MV)⌈seconds × rate⌉ video credits · min 6 · HD × 1.7 · first 2 videos ever = 1 credit each≤ 40 s · images ≤ 8 MB · audio ≤ 30 MB
MV Pro1 song credit per scene image + video ⌈seconds × rate × 2⌉ · min 12≤ 5 min · audio ≤ 60 MB · lyrics required
Dance videoMV price × 1.3≤ 40 s · reference clip ≤ 60 MB
Fashion reviewSame as MV≤ 40 s
Talking videoVoice 0.1 credit/s + video ⌈seconds × rate⌉ (min 6, HD × 1.7)≤ 120 s speech · text ≤ 3,000 chars
Podcast0.05 credit per second of audio, per episode (≈ 15 credits for a 5-minute episode)1–30 min/episode · ≤ 10 episodes per request
Text to speech0.05 credit per second (≈ 3 credits per minute)≤ 30 min (≈ 25,000 chars) per request

rate = your video price per second: 1.00 credit/s by default, 0.90 after a 99-credit pack, 0.70 after a 199-credit pack (see GET /api/v1/me). 1 credit = 1,000₫ (≈ $0.04); packs from $1.99 at /pricing.

Charges are held when a job starts and settled on completion; failed or timed-out jobs are refunded automatically. Generated files are yours to use commercially under the Terms; content must comply with our policies (no adult content, no impersonation of real people without consent).

Questions or need higher limits? Contact us.