# ZenityX Studio Agent API

Generate AI images, videos, speech and music (songs) on ZenityX Studio
(studio.zenityx.com), billed to the key owner's Studio credit balance.

Base URL: https://vaeubbfuezitvnakuiad.supabase.co/functions/v1/api-v1
MCP endpoint (same capabilities as tools): https://mcp.zenityx.com

## Authentication
Send the user's API key on every request (keys look like zxs_live_...):
    Authorization: Bearer zxs_live_xxx     (or header: x-api-key: zxs_live_xxx)
Keys are created at studio.zenityx.com -> profile menu -> API Keys.
Users often hand you the key as a zenityx.env file (contains
ZENITYX_API_KEY=zxs_live_... plus ZENITYX_API_BASE / ZENITYX_MCP_URL) —
read it from the path or attachment they provide; never echo the key back.
Accounts without API access get 403 (see the error table).

## Typical flow
1. GET /models                      -> pick a model_key and see prices
2. (optional) POST /files           -> upload a local reference image, get a URL
3. POST /images, /videos, /voice,  -> returns generation_id immediately (202)
   /avatar or /music
4. GET /generations/{id}            -> poll every 5-10s until status is
                                       "success" (media URL ready) or "failed"

## Endpoints

### GET /models
Active models. Fields: model_key, model_name, model_type
(image|video|upscale|voice|avatar|music), description, credit_cost,
pricing_rules. Music models also carry music_guide (modes + required/optional
fields, limits, credits_per_job, section tags). How pricing works:
- images: credit_cost is the base (1K) price. If the model ALSO has
  pricing_rules keyed by resolution (e.g. {"1K":2,"2K":3,"4K":4}), the price
  of the resolution you request applies.
- videos: pricing_rules is keyed by combination, e.g. "720p_8_sound" = 720p,
  8s, with audio. A mode prefix like "fast_480p" or "pro_1080p" is selected
  with the "size" parameter of POST /videos.

### GET /credits
-> {"credits": <number>, "api_key": {"daily_credit_limit", "spent_today",
    "remaining_today", "resets_at"}}
credits = the key owner's balance. api_key (API keys only — absent on OAuth
sessions, which have no per-key cap) = THIS key's own daily cap: a job priced
above remaining_today is refused with 429 daily_limit_exceeded (nothing
charged) until resets_at (midnight Asia/Bangkok). daily_credit_limit null =
no cap. Check it before a batch of jobs.
billing (team accounts) = which wallet pays for THIS caller: {"payer":
"personal"} or {"payer": "team", "team", "role", "quota": {"limit_5h",
"used_5h", "remaining_5h", "resets_5h_at", "limit_week", "used_week",
"remaining_week", "resets_week_at"}, "team_balance"?}. With payer "team" the
jobs are paid from the team's credit pool within this member's quota (null
limit = unlimited) and "credits" (the personal wallet) is NOT used;
team_balance is shown to the team owner / managers only. A job above the
remaining quota is refused with 429 team_quota (nothing charged).

### POST /files
Upload a reference file. THREE ways — prefer the first:
0. NO KEY AT HAND (e.g. OAuth session)? POST https://vaeubbfuezitvnakuiad.supabase.co/functions/v1/api-v1/files/upload-url (empty
   body, authenticated) → returns a one-time upload command ("how") you can
   run as-is: curl --fail-with-body --show-error --silent --request POST      '<upload_url>' --data-binary '@/path/to/file'
   SENSITIVE URL: valid ~10 min, this account only, no Authorization header
   needed — don't log or share it. The response names the right param next.
1. RAW BINARY with your API key (the file never passes through your context):
       curl -s -X POST https://vaeubbfuezitvnakuiad.supabase.co/functions/v1/api-v1/files          -H "Authorization: Bearer $ZENITYX_API_KEY"          --data-binary @/path/to/photo.jpg
2. JSON {"content_base64": "<base64, no data: prefix>", "filename": "optional.png"}
   — only for small files (base64 through your context is ~1.5 chars/token;
   downscale to ~500px first if you must use this). Base64 path caps at 8MB.
Size limits by media kind (raw path): image 20MB · audio 25MB · video 100MB
(video refs are for avatar motion-control; per-model fit like dimensions and
durations is validated at generation time). Magic-byte validated.
PREP TIP: convert PNG references to JPEG before uploading (sips -s format
jpeg in.png --out ref.jpg / magick in.png ref.jpg) — models ignore
transparency and raw PNGs intermittently trigger provider errors.
-> 201 {"url": "https://...", ...} — pass this url in image_urls.
TIP: a previous ZenityX generation needs NO upload — its URL from
GET /generations works directly in image_urls.

### POST /images
    {"prompt": "...", "model_key": "nano-banana-2-lite",
     "aspect_ratio": "1:1", "resolution": "1K", "output_format": "png",
     "image_urls": ["https://..."]}
Only prompt is required. Default model: nano-banana-2-lite (1 credit).
BATCH: send "prompts": ["...", ...] (2-8) instead of prompt — ONE request
submits all variations (identical other params). idempotency_key REQUIRED
(per-item keys derived as key#1..#N); max_credits = TOTAL budget split evenly.
Returns submitted[] with the generation_ids — then poll them together with
GET /generations?ids=... until every item is terminal.
"google_search" is no longer available: it belonged to nano-banana-2, which is
retired. Every image model rejects it with a clear 400 (nothing charged).
Model routing (covers almost every real job — prefer these unless the user
asks for something specific):
- nano-banana-2-lite (1cr, 1K, fastest) — drafts, experiments, bulk variants,
  mood boards. The default.
- nano-banana-2-1 (2cr, up to 4K, up to 10 refs) — THE workhorse: use it for
  almost any real deliverable (photos, products, people, scenes, edits). Best
  all-round value; when unsure, pick this. Set resolution 2K/4K for final use
  (price rises with resolution: 1K=2, 2K=3, 4K=4). Google's successor to
  nano-banana-2: better quality, prompt adherence, text in images and
  character consistency at the SAME price.
- nano-banana-2 — RETIRED (2026-10-09): requests that still send it run on
  nano-banana-2-1 at the same price (the 202 carries model_key + notice).
  Send nano-banana-2-1.
- nano-banana-pro (5cr) — flagship, only when maximum quality clearly matters.
- gpt-image-2-5 (2cr, up to 4K; quality=flare default | sunburst = more
  polished at the SAME price; 1:1 at 4K allowed; 27:16/16:27/9:8/8:9 are 1K
  only; up to 16 refs) — THE specialist: reach for it for GRAPHICS,
  INFOGRAPHICS, diagrams, UI mockups, posters, and images with a lot of
  correct text. Use resolution 4K for these (crisp text/lines). It supersedes
  gpt-image-2 (legacy — still runs, but prefer 2.5: better on every axis at
  the same price). Not needed for ordinary photos — nano-banana-2-1 is better
  there.
Other models (seedream, flux, grok, recraft-vector) are situational — see each
model's description in GET /models. Full catalog + prices: GET /models.
image_urls (optional, max 10) enables image-to-image / editing; use URLs from
POST /files or from previous generations. quality (optional): only recraft-vector
(normal|pro — changes the price) and gpt-image-2-5 (flare|sunburst — same price);
any other value there is a free 400 invalid_quality, other models ignore it.
-> 202 {"generation_id", "estimated_credits", "poll": "/functions/v1/api-v1/generations/<id>"}
"poll" (every create endpoint) is a path from the HOST root — it already
contains /functions/v1/api-v1, so don't append it to Base URL. Simplest: GET
https://vaeubbfuezitvnakuiad.supabase.co/functions/v1/api-v1/generations/<generation_id>.
A 202 with "confirmation": "pending" (any create endpoint) means our generator's
answer was lost after the job may have been charged: poll that id like any
other — never resubmit. If it fails it is refunded automatically.
Lite/default models usually finish in under a minute; premium image models
(nano-banana-pro, gpt-image-2) typically take 1-2 minutes and can reach 3-5 —
keep polling, do not resubmit.

### POST /videos
    {"prompt": "...", "model_key": "kling-3.0/video", "duration": 5,
     "resolution": "720p", "aspect_ratio": "16:9", "sound": true,
     "image_urls": ["https://..."]}
Only prompt is required. sound DEFAULTS TO TRUE (most users want audio) —
send "sound": false explicitly to save credits on models that price audio.
Videos are EXPENSIVE (2 to 100+ credits depending on
model/duration/resolution/sound) — check /models pricing_rules or /credits
first when unsure.
UPLOAD SLOTS — two different things, don't mix them up:
- image_urls = FIRST/LAST FRAME: [first] or [first, last] (max 2; enforced);
  the video starts (and optionally ends) exactly on these images.
- reference_image_urls = REFERENCE mode: subject/style references the model
  composes from freely. Per-model caps: seedance-2-5 30, rest of the
  seedance-2 family 9, wan/3-0-video 10, minimax/h3 9,
  veo-3-1 family 3, gemini-omni-video 7, cinematic-2.0 4; others reject with a
  clear 400.
- image_urls and reference_image_urls are MUTUALLY EXCLUSIVE — sending both
  is a 400. Pick the slot that matches the intent.
- gemini-omni-video runs on a 7-unit quota per request: each image costs 1
  unit and a video costs 2 (one video max), so 7 images alone or 5 with a
  video. Over-quota requests are rejected, not trimmed.
- reference_video_urls (Seedance 2 family 3 / seedance-2-5 10 / wan/3-0-video
  5 / minimax/h3 3): video references — upload them first (POST /files); links on
  any other host are a free 400 ref_video_url_refused. These are BILLED: every family meters
  (reference + output) seconds, so a 10s clip attached to a 6s render prices
  16s. wan/3-0-video additionally caps reference + output at 30s. The clips'
  TOTAL length is capped too: seedance-2 / seedance-2-fast 15s, seedance-2-5
  30s (input_slots.reference_videos_total_seconds_max) — over it is a free 400.
  We measure each clip ourselves (MP4/MOV); a fragmented clip (ffmpeg
  frag_keyframe / -frag_duration, browser MediaRecorder) counts to its last
  fragment, and one we can't read to the end is a free 422
  ref_video_duration_unknown — export it with ffmpeg -movflags +faststart.
  reference_audio_urls (same per-model caps): audio guidance, e.g. a
  POST /voice result. bytedance/seedance-2-5: EACH audio clip must be 2-30s
  (input_slots.reference_audio_seconds_min/max) — outside it is a free 400
  ref_audio_duration_out_of_range naming the clip. wan/3-0-video: each clip
  at least 1s AND all clips together at most 15s
  (input_slots.reference_audio_seconds_min/max +
  reference_audio_total_seconds_max) — otherwise the same free 400, naming
  the short clip or the measured total.
- bytedance/seedance-2 reference images: each one's width/height must be
  between 0.4 and 2.5, i.e. 2:5 to 5:2
  (input_slots.reference_image_aspect_ratio_min/max). A taller or wider one
  (long phone screenshot, panorama) is a free 400 image_aspect_out_of_range
  naming its index — crop it first.
- edit_mode: true (bytedance/seedance-2-5 only): EDIT one existing clip —
  exactly ONE reference_video_urls entry (4-30s), no image_urls; write the
  prompt as the change ("Change the jacket to red. Keep everything else the
  same."). Output keeps the clip's length and aspect ratio, so duration /
  aspect_ratio are ignored. Price = rate x (clip + clip) seconds; a shorter
  render is refunded automatically (ledger type refund_partial), never charged
  above the quote.
- Machine-readable: every video model in GET /models carries
  param_guide.input_slots (e.g. first_last_frame_max / reference_images_max /
  reference_videos_max / reference_audio_max) — trust it over prose.
Video model routing (pick by the job; other models are situational):
- Thai speech / dialogue in the video → gemini-omni-video (best), or the
  veo-3-1 family (veo-3-1-lite is the budget pick). These are the ONLY models
  that speak Thai — do NOT use kling/seedance when the user needs Thai audio.
- footage, product, high-resolution, fashion, people/portrait → kling-3.0/video.
- complex / cinematic / multi-reference work (smartest) → bytedance/seedance-2.
- 2K output with built-in stereo sound, or editing existing footage (swap a
  subject, change wardrobe/dialogue, transfer motion) → minimax/h3. Reference
  mode needs duration >= 6. For a single 4-30s clip, bytedance/seedance-2-5
  with edit_mode=true also edits.
- budget footage without Thai speech → bytedance/seedance-1.5-pro.
- long clips (2-30s) on a budget with many references (10 images / 5 videos /
  5 audios) → wan/3-0-video: 2/4/8 cr/s at 480p/720p/1080p, about a quarter
  of seedance-2-5. quality=prime is the faster endpoint at ~1.5x, same output.
Video params are model-specific. DON'T guess — GET /models returns a
`param_guide` for every video model telling you exactly which params affect its
price and the valid combinations. Read it first, then send params that match one
`valid_combos` entry:
- size (mode): fast|normal|std|pro|standard|prime — when combos start with one
  (e.g. "std_nosound", "prime_720p"). On wan/3-0-video prime = the faster
  endpoint at ~1.5x the price, same output quality; default standard.
- resolution: 480p|720p|768p|1080p|2k|4k — when combos contain one (e.g. "480p_8_sound")
- duration: whole seconds — when combos contain a number (e.g. "5","10","480p_4")
- sound: true|false — when combos end in _sound / _nosound
If the params you send don't form a priced combination (e.g. a per-second model
with duration omitted), you get a 400 with the model's valid_combos — NO credits
spent, just retry with a valid combination. An out-of-range duration returns 400
with allowed_durations.
-> 202 {"generation_id", "estimated_credits", "poll": "..."}

### POST /upscale
    {"model_key": "ultimate-image-upscaler", "source_url": "https://...",
     "resolution": "4k", "output_format": "png"}
Upscale an existing image or video to higher resolution (model_type 'upscale'
in GET /models). source_url is the media to upscale (from a previous generation
or POST /files). Image upscalers (input_type image, e.g. ultimate-image-
upscaler) charge a fixed price. VIDEO upscalers bill per second on the length
WE measure (see Duration billing), so the video must be a ZenityX file (POST
/files, or a previous generation's URL) — a link on any other host is a free
400 media_url_refused. resolution
must be one the model lists (400 with allowed_resolutions otherwise).
output_format (image upscale only): png (default) | jpeg | webp — "jpg" is
accepted as jpeg; anything else is a 400 before any charge. Scope: an
image upscale needs the 'images' scope, a video upscale needs 'videos'.
-> 202 {"generation_id", "estimated_credits", "poll": "..."}

### GET /voices
-> {"voices": [{"voice_id", "name", "description", "use_case", "gender",
    "age", "accent", "thai_verified", "aliases": [...]}, ...],
    "default_voice_id": "..."}
The Thai voice catalog (all Thai-verified for Eleven v3). CALL THIS FIRST when
picking a voice — names are curated data and DO change. Use `voice_id` (always
safe) or any listed alias. use_case is one of: narration · conversational ·
social_media · characters_voice_acting · informative_educational · meditation.

### POST /voice
    {"text": "สวัสดีค่ะ ยินดีต้อนรับสู่ ZenityX",
     "voice": "วาริน", "language_code": "auto", "stability": 0.5}
Direct ElevenLabs Eleven v3 text-to-speech / dialogue; there is no v2 fallback.
Thai works out of the box — leave language_code "auto" or pass "th". Simple
case: send `text` + a `voice` from GET /voices (a voice_id or a listed alias
such as "วาริน" / "Sarah"). Multi-speaker: send
`dialogues: [{"text": "...", "voice": "วาริน"}, {"text": "...", "voice":
"ภูผา"}]` with at most 10 unique voices. Omitting voice uses the default
voice_id from GET /voices. An unknown voice returns 400 naming the discovery
route; an ambiguous alias returns 400 listing the matching voice_ids. Max 5000
characters total. Cost uses the active model's explicit per-1000-character
pricing and fails closed if pricing is unavailable. Needs the 'videos' scope.
-> 202 {"generation_id", "estimated_credits", "poll": "..."} — the finished
audio URL comes back in the generation's image_url.

### POST /avatar
    {"model_key": "infinitetalk", "image_url": "https://...",
     "audio_url": "https://...", "duration_seconds": 10, "resolution": "480p"}
Make a person in a photo talk (lipsync) or copy a motion. Each avatar model
(model_type 'avatar' in GET /models) needs specific inputs:
- LIPSYNC (infinitetalk, ltx-lipsync): image_url + audio_url → the person
  speaks the audio. Pairs with POST /voice (generate speech, feed its URL here).
  Audio: infinitetalk up to 600s (10 min), ltx-lipsync up to 20s.
- MOTION (kling-2.6/motion-control, kling-3.0/motion-control, wan-animate):
  image_url + video_url → the subject copies the motion video. kling 3-30s,
  wan-animate 3-120s.
- video-translate: video_url (3-120s) + output_language, optional quality: "speed"
  (default) or "precision" (sharper lip-sync, e.g. visible teeth; twice the
  speed price) — GET /models modes/default_mode, prices in pricing_rules
  (480p = speed, precision_480p = precision). Other avatar models take no
  quality (400 invalid_quality).
Each model's limits are in GET /models as duration_limits {media, min_seconds,
max_seconds}: max_seconds applies to the billed (rounded-up) seconds, so 20.02s
of audio is 21s and over a 20s limit; min_seconds to the measured length. Out of
range = free 400 media_duration_out_of_range with the limits and the measure.
Billed per second on the length WE measure of the audio (lipsync) or the video
(motion/translate) — see Duration billing. That file must be a ZenityX file
(POST /files, or a POST /voice / previous generation URL); a link on any other
host is a free 400 media_url_refused. The photo may be any https link.
resolution must be one the model lists. Missing an input returns a 400 that
names it. Scope: 'videos'.
Local photo/audio/video files: upload via POST /files first (raw binary or
/files/upload-url), then use the returned url here.
-> 202 {"generation_id", "estimated_credits", "poll": "..."} — result video in image_url.
Avatar jobs render much slower than real time, in proportion to the length:
infinitetalk takes about 15-45s per second of audio at 720p (7-17s at 480p), so
1 min of audio is ~15-45 min and 10 min is ~2.5-7.5 hours. The job stays
"generating" the whole time — that is normal: poll every few minutes (not every
5-10s) and never resubmit. If the provider fails it, it is refunded automatically.

### POST /music
    {"mode": "idea",
     "idea": "เพลงป๊อปไทยสดใส เล่าเรื่องนักเรียนที่เพิ่งแต่งเพลงแรกด้วย AI อบอุ่น มีกำลังใจ",
     "style": "Thai pop, bright piano, acoustic guitar, warm female vocal",
     "idempotency_key": "<uuid>", "max_credits": 5}
Songs with the Suno V6 family. ONE job = 2 songs (two takes of the same
request), each with its own MP3, cover image, title and lyrics. Price per job
(for the pair): music_guide.credits_per_job in GET /models (type music) — the
same on every music model. model_key (optional): suno/v6 (default, most natural
vocals) · suno/v6-mini (faster, for trying ideas) · suno/v6-wild (experimental).
Pick ONE mode — the same three ways as the Music Studio page:
- "idea" — describe the song in plain words: mood, story, instruments, and the
  language to sing in (Thai works). The AI writes the lyrics, melody,
  arrangement and title, and picks the length (usually 2-4 minutes).
  Required: idea (≤3000 chars). Optional: style (≤1000), instrumental: true
  (no vocals). NOT used in this mode (400 if sent): title, lyrics,
  duration_seconds, vocal_gender, exclude_styles — say those things in idea/style.
- "lyrics" — the AI sings the user's own words. Required: title (≤80), lyrics
  (≤5000), rights_confirmed: true. Optional: style, exclude_styles (≤1000,
  what to avoid, e.g. "heavy metal, autotune"), vocal_gender "female"|"male",
  duration_seconds (10-360, default 180 = 3 min).
  Lyrics = only the sung words plus section tags, each tag on its own line:
      [Verse]
      สมุดการบ้านวางอยู่ข้างหน้าต่าง
      ดินสอขีดคำที่คิดมาทั้งวัน
      [Chorus]
      นี่คือเพลงแรกของฉัน
  Tags: [Intro] [Verse] [Pre-Chorus] [Chorus] [Bridge] [Outro]. The sound
  (genre, mood, instruments, voice, tempo) goes in style, not in lyrics.
  Set rights_confirmed: true ONLY after the user confirms they wrote the lyrics
  or have the right to use them.
- "instrumental" — no vocals (background for clips, presentations, podcasts).
  Required: title, style (e.g. "lo-fi hip hop, soft electric piano, rain
  ambience, 80 bpm"). Optional: exclude_styles, duration_seconds.
style reads best as comma-separated English tags. NEVER use artist or band
names, or lyrics of existing songs, in any field: the provider refuses the job
— it fails about a minute later and is refunded, but the user waited for
nothing. Describe the sound instead ("90s Thai rock ballad, raspy male vocal").
Every mistake in the body is a free 400 that names the field (codes
bad_request / unknown_field) — nothing is charged. Scope: 'videos'.
-> 202 {"generation_id", "status": "generating", "mode", "estimated_credits",
"songs": 2, "poll"}. Songs usually take 1-3 minutes: poll GET /generations/{id}
every 10-15 s. On "success" the generation has tracks: [{"index", "title",
"audio_url" (MP3), "cover_url", "duration_seconds", "style", "lyrics"}, ...]
(2 songs); image_url / image_urls hold the MP3 links too. Give the user BOTH
songs — they were made and paid for together.
duration_seconds is a target, not a promise: most takes land within a few
seconds of it, but now and then one take runs much shorter or longer (asked 60
→ got 60 and 162). Read tracks[].duration_seconds and tell the user which take
fits the length they wanted.

### GET /generations/{id}
-> {"generation": {"status": "generating"|"finalizing"|"success"|"failed", "image_url":
"<media URL when success — also used for videos>", "image_urls": [<all files
when a generation returns several>], "actual_credit_cost",
"error_message", ...}}
"finalizing" (with media_pending: true) means the job is DONE and charged and its
files are being copied to permanent storage — keep polling (seconds, rarely a few
minutes); it is not a failure, never resubmit. Only "success" carries the URLs.
Poll every 5-10s. Lite image models usually finish in under a minute; premium
image models typically 1-2 minutes (up to 3-5); videos 1-5 minutes (up to 10+);
avatar/lipsync jobs scale with the audio length and can take hours (see POST
/avatar); music 1-3 minutes (a music result also has tracks[] — see POST /music).

### GET /packages
Credit top-up packages: {id, name, credits, bonus_credits, total_credits,
price_thb, is_popular}. Bigger packages include FREE bonus credits — the
ladder matters: e.g. 1,000 THB -> +50 free, 5,000 THB -> +350 free,
10,000 THB -> +1,000 free (10%).

### POST /topup
    {"package_id": "<id from GET /packages>"}
-> 201 {"purchase_id", "amount_thb", "qr_image_url", "expires_at", ...}
A caller billed to a team (GET /credits → billing.payer "team") gets 409
team_billing: team credits are added by ZenityX for the team owner, never here.
Creates a Thai PromptPay QR that credits the key owner's account. Show the QR
(image or link) to the user; they scan it with any Thai banking app. Credits
apply automatically on payment — poll GET /credits every ~10s, then retry the
original request.
Agent etiquette for top-ups:
- Suggest the smallest package that covers the shortfall, AND always show the
  full ladder with what the next size up adds for FREE — the USER decides.
- Never call POST /topup before the user explicitly agrees to a package.
- Card payment is web-only (studio.zenityx.com) — never handle card data.

### GET /generations?limit=10
Recent generations submitted through this API (newest first, max 20).
This is the RECOVERY endpoint: if a generate request times out on your side,
the job and its charge still exist on the server — find it here (match by
prompt/created_at) and resume polling its id. Never resubmit first; a
resubmit creates a second job and charges again.
BULK POLL: GET /generations?ids=<uuid,uuid,...> (max 10) returns exactly
those generations in one roundtrip — poll a whole batch with a single
request instead of one call per id.

### Kling 3.0 quality tiers
"quality" (= "size") on kling-3.0/video takes std|pro|4k and picks BOTH the
price tier and the rendered output quality (std→720p, pro→1080p, 4k→4K).
Do NOT send resolution for kling — it has no effect there.

### Duration billing (avatar + video upscale)
The server MEASURES the real duration of the billed media (audio for lipsync,
video for motion/translate/upscale) and bills on that, rounded up to whole
seconds (5.04 s bills 6) — duration_seconds you send is informational only (echoed back as requested_duration_seconds next to
detected/billable). Each model has a minimum billed length (5 s on the
WaveSpeed lipsync/animate/translate and video upscale models: a 3 s clip bills
5 s). estimated_credits and billable_duration_seconds already include it —
estimated_credits is exactly what the job is charged, from the same price
table. Unverifiable media (unsupported container, unreachable
URL) is rejected with code duration_unverifiable — upload via POST /files
first (mp3/wav/mp4/mov are parseable). So is a file whose length can only be
estimated from its header (streamed WAV) or an MP3 whose frames can't all be
counted to the end of the file (junk between frames, audio inside a trailing
tag, VBR without a Xing frame, over 25MB) — re-export it as MP4/M4A/WAV or a
constant-bitrate MP3. A fragmented
MP4/M4A (ffmpeg frag_keyframe / -frag_duration, browser MediaRecorder) is
measured fragment by fragment; one too long to read that way is refused the
same way — export it unfragmented (ffmpeg -movflags +faststart). Nothing is
charged on a duration_unverifiable.
The billed file must be on ZenityX storage (https://img.zenityx.xyz/… — the url
POST /files returns, or a generation result): we measure it, and the provider must render the
same bytes. Any other link is refused before we read it — 400
media_url_refused with the field to fix (audio_url / video_url / source_url)
and charged_credits 0. Upload it with POST /files and retry.

### Billing truth per generation
GET /generations[/:id] returns charged_credits / refunded_credits /
net_credits / billing_status (charged|refunded|reserved|not_charged).
FAILED generations are auto-refunded: their net_credits is 0 — ignore the
legacy actual_credit_cost field on failed rows (it records the gross debit
that was already returned).

## Billing safety (all generate/upscale endpoints)
Every billable POST (/images /videos /voice /avatar /upscale /music) accepts:
- "idempotency_key": your unique string per creative intent (e.g. a UUID).
  Retrying with the same key + same payload returns the ORIGINAL generation
  (200, "replayed": true) — no second charge. Same key + different payload
  -> 409 idempotency_conflict. ALWAYS set it when you might retry. The
  replay answers before any reference-file check or pricing, so a retry
  never re-measures your files.
- "max_credits": hard budget cap — if the estimated cost exceeds it the call
  is rejected BEFORE any charge (400 over_budget, includes estimated_credits).

## Errors
JSON: {"error": "<human message>", "code": "<machine code>"}
| HTTP | code                 | meaning / what to do                          |
|------|----------------------|-----------------------------------------------|
| 401  | auth_error           | missing/invalid/revoked key                    |
| 403  | auth_error           | account lacks API access (closed beta)         |
| 400  | insufficient_credits | not enough credits — GET /packages, offer POST /topup (PromptPay QR in chat) |
| 400  | generation_rejected  | generation rejected for another reason (see error text) |
| 503  | provider_unavailable | ZenityX's upstream AI provider is temporarily down — NOT the user's credits, nothing charged; retry after retry_after_sec, never offer a top-up for this |
| 500  | pricing_unavailable  | no price is configured for this model + setting (a ZenityX config problem) — nothing charged; try another resolution/model, don't retry in a loop (a temporary price-check failure is a 503 internal with retry_after_sec — retry that one) |
| 400  | over_budget          | estimated cost > your max_credits — nothing charged |
| 409  | idempotency_conflict | same idempotency_key reused with a DIFFERENT payload — pick a new key |
| 429  | rate_limit           | too many requests — wait retry_after_sec       |
| 429  | team_quota           | team accounts: this member's 5-hour / weekly quota (quota_window "5h" / "week") or the team's weekly cap ("team_week") is used up — nothing charged; wait retry_after_sec when present (absent = the job is bigger than the limit: waiting won't help), or the user asks their team owner / manager to reset or raise it. Never offer a top-up |
| 402  | team_pool_insufficient | team accounts: the team's credit pool can't cover the job — nothing charged; team credits are added by ZenityX for the team owner, POST /topup does not help. Tell the user to ask their team owner |
| 403  | team_inactive        | team accounts: this account is not an active member of the team the key / session is billed to, or the team is suspended — nothing charged |
| 409  | team_billing         | POST /topup from a caller billed to a team — top-ups are personal; nothing charged |
| 429  | daily_limit_exceeded | this key's daily credit cap reached — nothing charged; it resets at midnight Asia/Bangkok: retry_after_sec / resets_at say when (GET /credits → api_key.remaining_today shows it before you submit) |
| 400  | unknown_model        | bad model_key — call GET /models               |
| 400  | unknown_field        | POST /music got a field it doesn't use — the response lists allowed_fields; put genre/mood/instruments/voice in style. Nothing charged |
| 403  | music_not_open       | Music isn't open for this account yet. Nothing charged |
| 503  | model_unavailable    | this music model is under maintenance — use another music model or retry after retry_after_sec. Nothing charged |
| 400  | invalid_aspect_ratio | this model doesn't take that ratio — response lists allowed_aspect_ratios; each model's list is in GET /models (aspect_ratios) |
| 400  | image_aspect_out_of_range | seedance-2: a reference image is taller/wider than 2:5…5:2 — response names the index; crop it. Nothing charged |
| 400  | ref_audio_duration_out_of_range | reference audio length: seedance-2-5 each clip 2-30s; wan/3-0-video each clip ≥1s and ≤15s in total — response names the index (or the total) + measured seconds. Nothing charged |
| 400  | ref_video_url_refused | a reference_video_urls clip is not a ZenityX file — upload it (POST /files) and retry. Nothing charged |
| 400  | media_url_refused    | /avatar or video /upscale: the per-second-billed file (response `field`: audio_url / video_url / source_url) is not a ZenityX file — upload it (POST /files) and retry. Nothing charged |
| 400  | invalid_quality      | quality is not one of the model's modes (response `allowed_qualities`). POST /avatar: only video-translate has them (speed / precision). POST /images: recraft-vector (normal / pro) and gpt-image-2-5 (flare / sunburst). Nothing charged |
| 400  | media_duration_out_of_range | POST /avatar: the audio/video is longer (billed seconds) or shorter than the model's duration_limits in GET /models — response has field, min_seconds, max_seconds, detected_duration_seconds. Trim it or pick another model. Nothing charged |
| 400  | duration_unverifiable | we couldn't measure the billed file's length (or only estimate it) — re-export as MP4/M4A/WAV/CBR MP3, upload, retry. Nothing charged |
| 400  | invalid_resolution   | this IMAGE model doesn't render that resolution (e.g. nano-banana-2-lite is 1K-only) — response lists allowed_resolutions; each model's list is in GET /models (resolutions) |
| 413/415 | too_large / unsupported_type | upload rejected                     |

## Limits
30 generations / 5 min and 200 / hour per account. Uploads: 100 files / 5 min
AND 2GB cumulative / day per account (resets midnight UTC = 07:00 Asia/Bangkok;
a 429 carries `reason`: `rate_limit_5min` or `daily_quota_2gb` +
`retry_after_sec` counting down to the actual reset). Upload once and REUSE the
returned url across generations — re-uploading the same file burns quota.
Size caps by media kind on the raw/tokened paths (image 20MB · audio 25MB ·
video 100MB) — the 8MB cap applies ONLY to the JSON base64 path. Failed
generations are auto-refunded by the platform.

## Remote MCP (agents)
MCP endpoint: https://mcp.zenityx.com (Streamable HTTP). Two ways to connect:
- OAuth (Claude "Add custom connector", and clients with MCP OAuth): paste
  https://mcp.zenityx.com, log in with the ZenityX Studio account, approve. No key needed.
- API key (Codex/config-based clients): send header x-api-key: zxs_live_...
  (or Authorization: Bearer). Prefer env-backed config over a literal key —
  tools like `codex mcp list --json` print config values unredacted:
      [mcp_servers.zenityx-studio]
      url = "https://mcp.zenityx.com"
      env_http_headers = { "x-api-key" = "ZENITYX_API_KEY" }
  then `source zenityx.env` (the file downloaded when the key was created).
After a tool-schema release (serverInfo.version changes) reconnect / start a
new task — long-lived sessions keep the old tool catalog cached and will not
see newly added parameters like max_credits until they reconnect.

## Rules for agents
- Never invent model_key values — always pick from GET /models.
- Warn the user before expensive video generations; state estimated_credits.
- Music: never put artist/band names or lyrics of existing songs in any field;
  ask the user before setting rights_confirmed; deliver both songs of a job.
- If a request times out on your side, check GET /generations before anything
  else — the job usually completed and resubmitting would charge again.
- The API key is a secret. Never print or echo it back in full.
