All models
Save 91%421 models across multiple vendors and groups
Input: $0.1290 / 1M Tokens
Output: $1.0750 / 1M Tokens
Gemini 3.5 Flash?Lite is a cost?effective multimodal model that enables low?cost sub?agent task execution and document parsing. It supports text, image, video, audio, and PDF inputs, and is designed for high?volume agent workflows, straightforward data extraction, and applications where latency and API costs are key constraints.
Input: $0.1875 / 1M Tokens
Output: $0.3750 / 1M Tokens
Our state-of-the-art flagship model is setting industry benchmarks in non-hallucination rates, agent?tool invocation, and instruction?following capabilities.
Input: $4.2000 / 1M Tokens
Output: $21.0000 / 1M Tokens
Kimi K3 is Kimi’s most powerful flagship model to date, with 2.8 trillion parameters. It is built on the KDA hybrid linear attention mechanism (Kimi Delta Attention) and attention residual techniques, natively supports visual understanding, and features a 1-million-token context window. As the world’s first open-source model at the 3-trillion-parameter scale, it is designed for cutting-edge intelligent applications such as long?sequence programming, knowledge work, and reasoning.
Input: $0.3000 / 1M Tokens
Output: $1.5000 / 1M Tokens
Claude Sonnet 5 natively supports a 1-million-token context window (1 million tokens is both the default and the maximum; no smaller?context variants are available), a 128K?token maximum output length, adaptive reasoning, and the same tools and platform capabilities as Claude Sonnet 4.6—except for the Priority Tier feature, which is not supported in Claude Sonnet 5.
Input: $1.5000 / 1M Tokens
Output: $7.5000 / 1M Tokens
Claude?fable?5 is a high?performance, publicly available large language model launched by Anthropic, featuring ultra?long context windows, multimodal understanding, advanced reasoning capabilities, and enterprise?grade knowledge work expertise—while also incorporating robust safety safeguards.
Input: $1.2600 / 1M Tokens
Output: $6.3000 / 1M Tokens
A new generation of production?ready large models, comprehensively upgraded in coding, agent capabilities, and multimodal performance, delivers enhanced autonomous planning, long?chain execution, and dynamic repair to tackle real?world enterprise?level complex tasks.
Input: $1.2600 / 1M Tokens
Output: $6.3000 / 1M Tokens
Seed?Evolving is a model in the Seed series, designed for agent? and coding?oriented scenarios. It excels at complex task orchestration, long?term planning, code generation, and tool invocation. By using the unified Model ID `doubao?seed?evolving`, you can continuously benefit from the latest model capabilities without needing to manage version updates.
$0.0271 / 次
Nano Banana Lite is designed to be the efficiency expert in the image-generation lineup, delivering ultra-low-latency, cost-effective image generation and editing. It supports only 1K?resolution image generation.
Input: $0.1500 / 1M Tokens
Output: $0.9000 / 1M Tokens
Gemini 3.5 Flash has officially launched (GA), delivering stable performance and enabling large-scale deployment in production environments. As our most advanced Flash model, it consistently delivers industry-leading capabilities at scale across agent execution, code generation, and long?term tasks.
Input: $0.1500 / 1M Tokens
Output: $0.7500 / 1M Tokens
Gemini 3.6 Flash consistently delivers cutting-edge, model?level intelligence, optimized to handle real?world tasks faster and at lower cost. Designed for the age of intelligent agents, it excels at code generation, agent execution, and spatial reasoning. This model is particularly effective in rapid agent workflows that involve complex coding cycles and iterations.
Input: $0.0180 / 1M Tokens
Output: $0.1080 / 1M Tokens
GPT?5.6 Luna is optimized for low?cost, high?volume processing scenarios, aligning with the lightweight model tier of the earlier GPT?5 Nano and supporting all suffix variants—?high, ?xhigh, ?low, ?medium, ?max, and ?ultra.
Input: $0.4500 / 1M Tokens
Output: $2.7000 / 1M Tokens
GPT?5.6 Sol is the flagship base model in the GPT?5.6 series, corresponding to the original GPT?5 baseline with no suffix. By aliasing gpt?5.6, requests are routed to GPT?5.6 Sol; the series supports all suffix variants—?high, ?xhigh, ?low, ?medium, ?max, and ?ultra.
Input: $0.5500 / 1M Tokens
Output: $3.3000 / 1M Tokens
GPT Image 2 is our state-of-the-art image-generation model, delivering fast, high-quality image generation and editing. It supports flexible image resolutions and high-fidelity image inputs: the Official tier and above enable 1K, 2K, and 4K outputs, while the Codex?exclusive and default tiers are limited to 1K.
Input: $0.3000 / 1M Tokens
Output: $0.9000 / 1M Tokens
Grok 4.5 is a cutting-edge model built by SpaceX AI specifically for coding, agent?driven tasks, and knowledge work. It was trained at SpaceX AI’s data center in Memphis, using a new dataset that spans the fields of science, engineering, and mathematics.
$0.0026 / 次
Kling-3.0-Turbo is a lightweight, ultra-fast video-generation model launched by Kuaishou’s Xingling. While ensuring stable and coherent human?motion animation, it dramatically accelerates rendering speed and reduces inference costs, making it ideal for the automated production of large-scale short?video content.
$0.0390 / 次
The Qwen-Image?2.0 series accelerated?version models seamlessly integrate image generation and image editing, offering advanced text?to?image capabilities with support for 1,000?token prompts, exquisitely realistic textures, meticulous rendering of photorealistic scenes, and enhanced semantic alignment. The accelerated version strikes an optimal balance between model performance and quality.
Input: $1.0800 / 1M Tokens
Output: $3.2400 / 1M Tokens
The Max model, the largest and most capable variant in the Qwen3.7 series, currently offers a pure?text?only interface for users to try out. Qwen3.7 is a next?generation flagship model designed for the agent?centric era, with its core strengths lying in the breadth and depth of its agent?level capabilities: it excels at programming, office automation and productivity tasks, and long?term autonomous execution.
$0.0975 / 次
Wanxiang 2.7 is a flagship model for image generation and editing, supporting text-to-image, text?to?image?group, image?to?image?group, image editing, multi?image reference generation, and interactive editing. It delivers superior performance in text rendering, subject consistency, and adherence to complex prompts.
Input: $0.7500 / 1M Tokens
Output: $3.7500 / 1M Tokens
Claude?Opus?4?8 is currently Anthropic’s most powerful “tool capable of handling lengthy, complex tasks on its own,” making it particularly well suited for developers tackling large-scale projects, building AI agents, or working in scenarios with extremely high demands for quality and autonomy.
Input: $0.7500 / 1M Tokens
Output: $3.7500 / 1M Tokens
Claude Opus 5, the default model for Claude Max and the most powerful model for Claude Pro, is a thoughtful, proactive, cost-efficient agentic model that delivers nearly the cutting-edge intelligence of flagship Claude Fable 5 at half the cost and sets new state-of-the-art standards for coding and professional knowledge work.