AI model catalog: image, video, audio and chat APIs
Browse every image, video, audio and chat model on Unicorn API with its exact price in credits, inputs, examples and a playground. One key calls them all.
- GPT Image 2.5 Flare — image · OpenAI · 6 credits per image at 1K, 10 at 2K, 16 at 4K
OpenAI GPT Image 2.5 Flare: the faster GPT Image 2.5 variant for everyday generation and edits at 1K, 2K or 4K, with up to 16 reference images. - GPT Image 2.5 Sunburst — image · OpenAI · 6 credits per image at 1K, 10 at 2K, 16 at 4K
OpenAI GPT Image 2.5 Sunburst: the premium GPT Image 2.5 variant — tighter control and more polished output, slower than Flare. Same price. - GPT Image 2 — image · OpenAI · 5 credits per image at 1K, 8 at 2K, 15 at 4K
OpenAI GPT Image 2: the latest OpenAI image model with the widest aspect-ratio range at 1K, 2K or 4K, and up to 16 reference images. - GPT Image 1.5 — image · OpenAI · 3 credits per image at low quality, 4 at medium, 15 at high
OpenAI GPT Image 1.5: reliable prompt adherence and legible text at low, medium or high quality, with optional transparent background and up to 4 reference images. - Nano Banana 2 Lite — image · Google · 4 credits per image
Google Nano Banana 2 Lite: the cheapest Gemini image model — fast 1K images and edits with up to 10 reference images. - Grok Imagine 2 — image · Grok · 4 credits per image
xAI Grok Imagine 2: vivid, stylised images and edits with up to 5 reference images. - Grok Imagine — image · Grok · 6 credits per call; each call returns 6 images (2 with reference images)
xAI Grok Imagine: returns 6 variations per text prompt (2 when reference images are attached) for one flat price. - Nano Banana 2 — image · Google · 9 credits per image at 1K, 12 at 2K, 17 at 4K
Google Nano Banana 2: the current Gemini image model — strong realism, text and edits at 1K, 2K or 4K, with up to 14 reference images. - Nano Banana Pro — image · Google · 18 credits per image at 1K or 2K, 33 at 4K
Google Nano Banana Pro: high-fidelity images with accurate text rendering at 1K, 2K or 4K, with up to 8 reference images. - Nano Banana — image · Google · 6 credits per image
Google Nano Banana (Gemini image): quick, natural-looking images and edits with up to 3 reference images. - Seedream 5 Pro — image · ByteDance · 6 credits per image at 1K, 10 at 2K
ByteDance Seedream 5 Pro: the top Seedream tier for polished illustration and product shots at 1K or 2K, with up to 10 reference images. - Seedream 5 Lite — image · ByteDance · 6 credits per image at 2K, 8 at 3K
ByteDance Seedream 5 Lite: fast everyday generation at 2K or 3K, with up to 14 reference images. - Seedream 4.5 — image · ByteDance · 6 credits per image (2K or 4K)
ByteDance Seedream 4.5 at 2K or 4K. Accepts up to 14 reference images for edits and consistent characters. - Wan 2.7 Image — image · Wan · 5 credits per image (1K or 2K)
Alibaba Wan 2.7 Image: flat-priced 1K or 2K generation and editing with up to 9 reference images. - Wan 2.7 Image Pro — image · Wan · 10 credits per image (1K, 2K or 4K)
Alibaba Wan 2.7 Image Pro: the higher-quality Wan tier with 4K output (text-to-image), flat-priced, with up to 9 reference images. - Flux 2 Pro — image · Black Forest Labs · 6 credits per image; +3 per reference image
Black Forest Labs Flux 2 Pro: photoreal detail and strong prompt following at 1K or 2K, with up to 8 reference images. - Z-Image — image · Qwen · 3 credits per image
Fast, low-cost text-to-image model from Tongyi. Good for quick drafts and large batches. - Grok Imagine Video 1.5 VIP — video · Grok · 2 credits per second at 480p, 3 at 720p, 5 at 1080p
Grok Imagine Video 1.5 at the lowest per-second rate in the catalog. Image-to-video only: animate a start frame at 480p, 720p or 1080p for 1-15 seconds. - Grok Imagine Video 1.5 — video · Grok · 3 credits per second at 480p, 4 at 720p, 7 at 1080p
xAI Grok Imagine Video 1.5: very low cost per second with strong motion quality. Image-to-video only: animate a start frame at 480p, 720p or 1080p for 1-15 seconds. Can be busy at peak times. - Grok Imagine Video 1.0 — video · Grok · 3 credits per second at 480p, 4 at 720p, 7 at 1080p
xAI Grok Imagine Video 1.0: text-to-video or image-to-video at 480p, 720p or 1080p for 6-30 seconds, at the same low per-second rate as 1.5. - Wan 3.0 — video · Wan · 6 credits per second at 480p, 12 at 720p, 24 at 1080p; audio included
Alibaba Wan 3.0: text-to-video or image-to-video up to 30 seconds at 480p, 720p or 1080p with native synced audio at no extra cost. - Wan 2.7 — video · Wan · 10 credits per second at 720p, 15 at 1080p
Alibaba Wan 2.7: image-to-video only. Animate a start frame for 2-15 seconds at 720p or 1080p. - Seedance 2.0 Fast — video · ByteDance · 8 credits per second at 480p, 18 at 720p; audio included
ByteDance Seedance 2.0 Fast: the quicker, cheaper Seedance 2.0 for text-to-video or image-to-video at 480p or 720p, 4-15 seconds, with native audio. - Seedance 2.0 — video · ByteDance · 10 credits per second at 480p, 22 at 720p; audio included
ByteDance Seedance 2.0: high-quality text-to-video or image-to-video at 480p or 720p, 4-15 seconds, with native audio. - Seedance 2.5 — video · ByteDance · 19.6 credits per second at 480p, 44.1 at 720p, 79.8 at 1080p (rounded up on the total); audio included
ByteDance Seedance 2.5: text-to-video or image-to-video up to 30 seconds at 480p, 720p or 1080p with native audio. Image-to-video output follows the start frame (adaptive aspect). Optional web search grounds the prompt. - Wan 2.2 Fast — video · Wan · 14 credits per 81 frames at 720p (8 at 480p); 18 / 10 with interpolation; scaled by frame count
Alibaba Wan 2.2 Fast: image-to-video only. Length is set in frames (81-121) and playback FPS (5-30); optional frame interpolation for smoother motion. - Seedance 1.5 Pro — video · ByteDance · 3 credits per second (6 with audio) + 3 credits per video
ByteDance Seedance 1.5 Pro: budget text-to-video or image-to-video for 4-12 seconds, with an optional AI audio track. - Kling 3.0 — video · Kling · 9 credits per second in std quality, 11 in pro
Kuaishou Kling 3.0: cinematic text-to-video or image-to-video, with native sound, subjects, storyboards and start/end frames. - Kling 3.0 Omni — video · Kling · 9.8 credits per second at 720p (12.6 with audio), 12.6 at 1080p (16.1 with audio), 46.9 at 4k (rounded up on the total)
Create, animate, reference or transform video with Kling 3.0 Omni. Includes start/end frames, characters, native audio and custom storyboards, up to 4K. - Kling 2.5 Turbo Pro — video · Kling · 40 credits for a 5 second clip, 80 for 10 seconds
Kuaishou Kling 2.5 Turbo Pro: fast, fluid text-to-video or image-to-video in 5 or 10 second clips at a flat price. - 4o Image — image · OpenAI · 6 credits per image
- Ai Music API Add Instrumental — audio · Suno · 12 credits per track
- Ai Music API Boost Music Style — audio · Suno · 1 credit per track
- Ai Music API Create Music Video — video · Suno · 2 credits per video
- Ai Music API Extend Music — audio · Suno · 12 credits per track
- Ai Music API Generate MIDI from Audio — audio · Suno · 1 credit per track
- Ai Music API Generate Music Cover — image · Suno · 1 credit per image
- Ai Music API Generate Persona — audio · Suno · 1 credit per track
- Ai Music API Get TimeStamped Lyrics — audio · Suno · 1 credit per track
- Ai Music API Mashup — audio · Suno · 12 credits per track
- Ai Music API Separate Vocals — audio · Suno · 10 credits per track separate_vocal, 20 split_stem_advanced, 50 split_stem
- Ai Music API Sounds — audio · Suno · 3 credits per track
- Ai Music API Upload And Convert to WAV Format — audio · Suno · 1 credit per track
- Ai Music API Upload And Cover Audio — audio · Suno · 12 credits per track
- AI music generate — audio · Suno · 12 credits per track
- Claude Fable 5 — chat · Anthropic · 800 credits / M input tokens, 4000 / M output
- Claude HaiKu 4.5 — chat · Anthropic · 55 credits / M input tokens, 285 / M output
- Claude Opus 4.5 — chat · Anthropic · 285 credits / M input tokens, 1430 / M output
- Claude Opus 4.6 — chat · Anthropic · 285 credits / M input tokens, 1430 / M output
- Claude Opus 4.7 — chat · Anthropic · 285 credits / M input tokens, 1430 / M output
- Claude Opus 4.8 — chat · Anthropic · 400 credits / M input tokens, 2000 / M output
- Claude Opus 5 — chat · Anthropic · 400 credits / M input tokens, 2000 / M output
- Claude Sonnet 4.6 — chat · Anthropic · 170 credits / M input tokens, 855 / M output
- Claude Sonnet 5 — chat · Anthropic · 170 credits / M input tokens, 855 / M output
- Claude Sonnet 5.5 — chat · Anthropic · 160 credits / M input tokens, 800 / M output
- DeepSeek V4.1 Flash — chat · DeepSeek · 24 credits / M input tokens, 95 / M output
- Elevenlabs Text to Dialogue V3 — audio · Elevenlabs · 14 credits per request
- Elevenlabs Text to Speech Multilingual V2 — audio · Elevenlabs · 12 credits per 1,000 characters
- Elevenlabs Text to Speech Turbo V2.5 — audio · Elevenlabs · 6 credits per 1,000 characters
- Flux 2 Flex Image To Image — image · Black Forest Labs · 14 credits per image 1K, 24 2K
- Flux 2 Flex Text To Image — image · Black Forest Labs · 14 credits per image 1K, 24 2K
- Gemini 2.5 Pro — chat · Google · 76 credits / M input tokens, 600 / M output
- Gemini 3 Pro — chat · Google · 100 credits / M input tokens, 700 / M output
- Gemini 3.5 Flash — chat · Google · 90 credits / M input tokens, 540 / M output
- Gemini 3.5 Flash OpenAI — chat · Google · 90 credits / M input tokens, 540 / M output
- Gemini 3.6 Flash — chat · Google · 45 credits / M input tokens, 225 / M output
- Gemini 3.6 Flash OpenAI — chat · Google · 45 credits / M input tokens, 225 / M output
- Gemini 3.7 Flash — chat · Google · 45 credits / M input tokens, 225 / M output
- Gemini 3.7 Flash OpenAI — chat · Google · 45 credits / M input tokens, 225 / M output
- Gemini 3.8 Flash — chat · Google · 45 credits / M input tokens, 225 / M output
- Gemini 3.8 Flash OpenAI — chat · Google · 45 credits / M input tokens, 225 / M output
- Gemini Omni Flash 1.1 — video · Google · 63 credits per video 4 1080p, 63 4 720p, 63 4 360p, 84 6 1080p, 84 6 720p, 84 6 360p, and 10 more
- Gemini Omni Video — video · Google · 63 credits per video 4 1080p without video list, 63 4 720p without video list, 84 6 1080p without video list, 84 6 720p without video list, 105 8 1080p without video list, 105 8 720p without video list, and 9 more
- Google Imagen 4 Ultra-Text to Image — image · Google · 12 credits per image
Google Imagen 4 Ultra is the latest breakthrough in text-to-image generation, offering unmatched precision, speed, and photorealistic quality. Developed by Google DeepMind, this model transforms detailed text prompts int - Google Imagen 4-Text to Image — image · Google · 8 credits per image
Google Imagen 4, developed by Google DeepMind and released at Google I/O 2025, is an advanced text-to-image generation model. It converts textual prompts into photorealistic, high-quality images with exceptional detail, - GPT 5.2 — chat · OpenAI · 87.5 credits / M input tokens, 700 / M output
- GPT 5.5 — chat · OpenAI · 280 credits / M input tokens, 1680 / M output
- GPT 5.6 Luna — chat · OpenAI · 11.2 credits / M input tokens, 67.2 / M output
- GPT 5.6 Sol — chat · OpenAI · 280 credits / M input tokens, 1680 / M output
- GPT 5.6 Terra — chat · OpenAI · 112 credits / M input tokens, 672 / M output
- GPT 6 Astra — chat · OpenAI · 560 credits / M input tokens, 2800 / M output
- GPT 6 Luna — chat · OpenAI · 6 credits / M input tokens, 30 / M output
- GPT 6 Sol — chat · OpenAI · 120 credits / M input tokens, 600 / M output
- GPT 6.1 Sol — chat · OpenAI · 120 credits / M input tokens, 600 / M output
- Grok 4.5 — chat · Grok · 160 credits / M input tokens, 480 / M output
- Grok 4.6 — chat · Grok · 160 credits / M input tokens, 480 / M output
- Grok 4.7 — chat · Grok · 160 credits / M input tokens, 480 / M output
- Grok Imagine Extend — video · Grok · 15 credits per video
- Grok Imagine Upscale — video · Grok · 10 credits per video 720p, 20 1080p
- Hailuo 02 Image to Video Pro — video · Hailuo · 57 credits per video
- Hailuo 02 Text to Video Pro — video · Hailuo · 57 credits per video
- Hailuo 02 Text to Video Standard — video · Hailuo · 30 credits per video 6, 50 10
- Hailuo 2.3 Image to Video Pro — video · Hailuo · 45 credits per video 6 768P, 80 6 1080P, 90 10 768P
- Hailuo 2.3 Image to Video Standard — video · Hailuo · 30 credits per video 6 768P, 50 10 768P, 50 6 1080P
- HappyHorse 1.1 Image To Video — video · Alibaba · 22.5 credits per second 720p, 29 1080p
- HappyHorse 1.1 Reference To Video — video · Alibaba · 22.5 credits per second 720p, 29 1080p
- HappyHorse 1.1 Text To Video — video · Alibaba · 22.5 credits per second 720p, 29 1080p
- HappyHorse Image To Video — video · Alibaba · 28 credits per second 720p, 48 1080p
- HappyHorse Reference To Video — video · Alibaba · 28 credits per second 720p, 48 1080p
- HappyHorse Text To Video — video · Alibaba · 28 credits per second 720p, 48 1080p
- HappyHorse Video Edit — video · Alibaba · 28 credits per second 720p, 48 1080p
- Ideogram Character Base — image · Ideogram · 12 credits per image TURBO, 18 BALANCED, 24 QUALITY
- ideogram character edit — image · Ideogram · 12 credits per image TURBO, 18 BALANCED, 24 QUALITY
- ideogram character remix — image · Ideogram · 12 credits per image TURBO, 18 BALANCED, 24 QUALITY
- Ideogram V3 Edit Image-Image to Image — image · Ideogram · 4 credits per image TURBO, 7 BALANCED, 10 QUALITY
Ideogram V3 Edit API enables precise mask-based image editing, letting you change selected areas while keeping the rest untouched. Ideal for background replacement, object edits, and detail enhancements with flexible ren - Ideogram V3 Remix — image · Ideogram · 4 credits per image TURBO, 7 BALANCED, 10 QUALITY
Ideogram 3.0 Remix is an advanced image-to-image AI model designed for prompt-based visual transformation, enabling users to generate new variations from existing images with high fidelity. - Ideogram V3 Text to Image — image · Ideogram · 4 credits per image TURBO, 7 BALANCED, 10 QUALITY
Ideogram V3 API is a new text-to-image model from Ideogram, featuring improved realism, style control, and precise text rendering. It supports Turbo, Default, and Quality modes to suit different creative needs. - Imagen 4 Fast-Text to Image — image · Google · 4 credits per image
- Kimi K3 — chat · Moonshot · 480 credits / M input tokens, 2400 / M output
- Kling 2.1 Master — video · Kling · 160 credits per video 5, 320 10
Kling 2.1 Master API is the premium endpoint of the Kling 2.1 model developed by Kuaishou. It enables text-to-video and image-to-video generation with cinematic quality, motion fluidity, and prompt precision in 1080p res - Kling 2.1 Pro — video · Kling · 50 credits per video 5, 100 10
The Kling 2.1 Pro is the latest model in the Kling AI series, specializing in image-to-video generation. With enhanced motion simulation, high-quality 1080p output, and improved semantic understanding, it is perfect for - Kling 2.1 Standard — video · Kling · 25 credits per video 5, 50 10
Kling 2.1 Standard is a high-performance image-to-video model from Kling AI (Kuaishou), offering fast, cinematic AI video generation with 720P quality for creators and developers. - Kling 2.6 Image To Video — video · Kling · 110 credits per video without sound 10, 110 with sound 5, 220 with sound 10
- Kling 2.6 Motion Control — video · Kling · 11 credits per second 720p, 18 1080p
- Kling 2.6 Text To Video — video · Kling · 55 credits per video without sound 5, 110 without sound 10, 110 with sound 5, 220 with sound 10
- Kling 3.0 Motion Control — video · Kling · 20 credits per second
- Kling v2.1 Master — video · Kling · 160 credits per video 5, 320 10
Unleash Cinematic Magic: Turn Static Images into Dynamic Masterpieces with Kling v2.1 Master. - MiniMax H3 Image to Video — video · Hailuo · 8 credits per second 768P, 13 2K
- MiniMax H3 Reference to Video — video · Hailuo · 8 credits per second 768P, 13 2K
- MiniMax H3 Text to Video — video · Hailuo · 8 credits per second 768P, 13 2K
- OmniHuman 1.5 — video · ByteDance · 27 credits per second
- PixVerse V6 Image To Video — video · Pixverse · 4 credits per second 360p without generate audio switch, 5.6 540p without generate audio switch, 5.6 360p with generate audio switch, 7.2 720p without generate audio switch, 7.2 540p with generate audio switch, 9.6 720p with generate audio switch, and 2 more
- PixVerse V6 Video Fusion — video · Pixverse · 4 credits per second 360p without generate audio switch, 5.6 540p without generate audio switch, 5.6 360p with generate audio switch, 7.2 720p without generate audio switch, 7.2 540p with generate audio switch, 9.6 720p with generate audio switch, and 2 more
- Qwen Image Edit — image · Qwen · 5 credits per image
Qwen-Image-Edit is an open-source image editing model based on Qwen-Image, supporting semantic and appearance editing with precise, visually coherent results. It also handles bilingual (Chinese and English) text editing - Qwen2 Text to Image — image · Qwen · 6 credits per image
- Qwen3 Image to Image — image · Qwen · 5 credits per image 2K, 5 1K
- Qwen3 Pro Image to Image — image · Qwen · 7 credits per image 1K, 12 2K
- Qwen3 Pro Text to Image — image · Qwen · 7 credits per image 1K, 12 2K
- Qwen3 Text to Image — image · Qwen · 5 credits per image 2K, 5 1K
- Recraft Crisp Upscale — image · Recraft · 1 credit per image
- Recraft Remove Background — image · Recraft · 1 credit per image
- Runway — video · Runway · 12 credits per video
Empower Your Vision: Runway API Unleashes Cinematic AI. - Seedance 2.0 Mini — video · ByteDance · 2.4 credits per second 480p with reference video urls, 3.8 480p without reference video urls, 5 720p with reference video urls, 8.2 720p without reference video urls
- Seedream 5.0 Flash Image to Image — image · ByteDance · 4 credits per image 2K, 4 1.5K, 4 1K
- Seedream 5.0 Flash Layer Decomposition — image · ByteDance · 4 credits per image 2K, 4 1.5K, 4 1K
- Seedream 5.0 Flash Text To Image — image · ByteDance · 4 credits per image 2K, 4 1.5K, 4 1K
- topaz Upscale Image — image · Topaz · 10 credits per image 2, 20 4
- Veo 1080p video — video · Google · 5 credits per video
- Veo Extend — video · Google · 30 credits per video
- Wan 2.2 A14B Text to Video Turbo — video · Wan · 40 credits per video 480p, 80 720p
Wan-2.2 Text-to-Video A14B Turbo is an advanced AI video model that transforms text prompts into high-quality cinematic clips with smooth motion, rich details, and accurate scene alignment. Designed for developers and cr - Wan 2.5 Image To Video — video · Wan · 60 credits per video 5 720p, 100 5 1080p, 120 10 720p, 200 10 1080p
- Wan 2.5 Text To Video — video · Wan · 60 credits per video 5 720p, 100 5 1080p, 120 10 720p, 200 10 1080p
- Wan 2.6 Image To Video — video · Wan · 70 credits per video 5 720p, 105 5 1080p, 140 10 720p, 210 10 1080p, 210 15 720p, 315 15 1080p
- Wan 2.6 Text To Video — video · Wan · 70 credits per video 5 720p, 105 5 1080p, 140 10 720p, 210 10 1080p, 210 15 720p, 315 15 1080p
- Wan 2.6 Video To Video — video · Wan · 70 credits per video 5 720p, 105 5 1080p, 140 10 720p, 210 10 1080p
- Wan 2.7 Reference to Video — video · Wan · 16 credits per second 720p, 24 1080p
- Wan 2.7 Text to Video — video · Wan · 16 credits per second 720p, 24 1080p
- Wan 3.0 Video Prime — video · Wan · 12.2 credits per second 480P, 25.2 720P, 50.4 1080P