| | 0.0012 / second | Upscale images with Aura SR, billed by inference time. |
| | CosyVoice 3.5 Flash: 0.144 / 10k chars- CosyVoice 3.5 Flash: 0.144 / 10k chars
- CosyVoice 3 Flash: 0.18 / 10k chars
- CosyVoice 3.5 Plus: 0.27 / 10k chars
- CosyVoice 3 Plus or 2: 0.36 / 10k chars
| Generate multilingual speech with CosyVoice models. |
| | 0.043 / image | Generate images from text prompts with Doubao Seedream. |
| | 0.34 / hour | Transcribe speech audio into text. |
| | 0.77 / 10k chars | Generate speech audio from text with Doubao voices. |
| | 0.075 / image | Generate illustrations for text-to-EPUB workflows. |
| | 0.048 / image | Edit images with Flux Pro Kontext. |
| | Low: 0.012 / image- Low: 0.012 / image
- Medium: 0.072 / image
- High: 0.264 / image
- Low Large: 0.024 / image
- Medium Large: 0.144 / image
- High Large: 0.528 / image
| Generate and edit images with GPT Image 2. |
| | Read: 0.05 / M tokens- Read: 0.05 / M tokens
- Search: 0.05 / M tokens
| Read web pages and search web content for agents. |
| | 9.9 / 1,000 pages | Translate and optionally colorize manga pages from a ZIP archive. |
| | Nano Banana: 0.047 / image- Nano Banana: 0.047 / image
- Nano Banana Edit: 0.047 / image
- Nano Banana Pro 1K or 2K: 0.18 / image
- Nano Banana Pro 4K: 0.36 / image
- Nano Banana 2 0.5K: 0.072 / image
- Nano Banana 2 1K: 0.096 / image
- Nano Banana 2 2K: 0.144 / image
- Nano Banana 2 4K: 0.192 / image
| Generate and edit images with Nano Banana models. |
| | 0.3 / image | Recognize text in images and optionally correct detected text. |
| | Next Generation: 0.144 / 10k chars- Next Generation: 0.144 / 10k chars
- Classic: 1.09 / 10k chars
| Generate speech from text with Leina voice services. |
| | Generate: 0.04 / image- Generate: 0.04 / image
- Edit: 0.04 / image
| Generate or edit images through the OpenAI-compatible image API. |
| | ~2.5 / 1,000 pages | Convert PDF documents into EPUB files. |
| | ~2.5 / 1,000 pages | Convert PDF documents into Markdown files. |
| | Voice Cloning: 0.0018 / voice- Voice Cloning: 0.0018 / voice
- Voice Design: 0.035 / voice
| Create custom voices by cloning or describing a voice. |
| | Input: 0.1 / M tokens- Input: 0.1 / M tokens
- Output: 0.17 / M tokens
| Extract and analyze information from documents with Qwen. |
| | Qwen Image 2.0: 0.034 / image- Qwen Image 2.0: 0.034 / image
- Qwen Image 2.0 Pro: 0.085 / image
| Generate images with Qwen Image models. |
| | Qwen Image Edit Plus: 0.034 / image- Qwen Image Edit Plus: 0.034 / image
- Qwen Image Edit: 0.051 / image
- Qwen Image Edit Max: 0.085 / image
| Edit images with Qwen image models. |
| | 0.06 / image | Separate an image into multiple visual layers. |
| | 0.0005 / image | Translate text in images while preserving the visual layout. |
| | 0.000039 / second | Transcribe audio files into text with Qwen ASR. |
| | 0.144 / 10k chars | Generate speech audio from text with Qwen voices. |
| | 0.021 / image | Cut out subjects and remove image backgrounds. |
| | Fast Image to Video: 3.96 / M tokens- Fast Image to Video: 3.96 / M tokens
- Fast Text to Video: 6.66 / M tokens
- Image to Video 720p: 5.04 / M tokens
- Text to Video 720p: 8.28 / M tokens
- Image to Video 1080p: 5.58 / M tokens
- Text to Video 1080p: 9.18 / M tokens
| Generate videos from prompts and reference media with Seedance 2.0. |
| | Text to Video 720p: 0.36 / second- Text to Video 720p: 0.36 / second
- Text to Video 1080p: 0.6 / second
- Image to Video 720p: 0.36 / second
- Image to Video 1080p: 0.6 / second
| Generate video clips from text prompts or a source image. |
| | Free | Store temporary files for free. Files may be removed periodically; each file can be up to 500 MB. |
| | 0.009 / image | Compress PNG, JPEG, and WebP images. |
| | 0.035 / image | Generate and edit images with Wan models. |
| | Keyframe Wan 2.2 Flash 480p: 0.017 / second- Keyframe Wan 2.2 Flash 480p: 0.017 / second
- Keyframe Wan 2.2 Flash 720p: 0.034 / second
- Keyframe Wan 2.2 Flash 1080p: 0.085 / second
- Keyframe Wan 2.1 Plus 720p: 0.119 / second
- Image to Video Wan 2.6 Flash 720p Silent: 0.0261 / second
- Image to Video Wan 2.6 Flash 1080p Silent: 0.0435 / second
- Image to Video Wan 2.6 Flash 720p Audio: 0.0522 / second
- Image to Video Wan 2.6 Flash 1080p Audio: 0.087 / second
- Image to Video Wan 2.6 720p: 0.1044 / second
- Image to Video Wan 2.6 1080p: 0.174 / second
- Reference to Video Wan 2.6 Flash 720p Silent: 0.0261 / second
- Reference to Video Wan 2.6 Flash 1080p Silent: 0.0435 / second
- Reference to Video Wan 2.6 Flash 720p Audio: 0.0522 / second
- Reference to Video Wan 2.6 Flash 1080p Audio: 0.087 / second
- Reference to Video Wan 2.6 720p: 0.1044 / second
- Reference to Video Wan 2.6 1080p: 0.174 / second
- Text to Video Wan 2.6 720p: 0.1044 / second
- Text to Video Wan 2.6 1080p: 0.174 / second
| Generate videos from text, images, keyframes, or reference material with Wan models. |