Skip to content
  • Models
  • Rankings
  • Ori
ElevenLabs launch offer: every ElevenLabs model is 50% off through October 19, 2026. See ElevenLabs models
OpenRouterOpenRouter
© 2026 OpenRouter, Inc

Product

  • Chat
  • Rankings
  • Benchmarks
  • Apps
  • Discover
  • Models
  • Ori
  • Collections
  • Providers
  • Tools
  • Pricing
  • Business
  • Enterprise
  • Labs

Company

  • About
  • Blog
  • Careers
    Hiring
  • Privacy
  • Terms of Service
  • Trust Center
  • Support
  • Works With OR
  • Data
  • Brand

Developer

  • Documentation
  • API Reference
  • Developer Platform
  • Status
  • AI Site Map

Connect

  • Discord
  • GitHub
  • LinkedIn
  • X
  • YouTube
Favicon for Alibaba

Alibaba Cloud Int.

Browse models provided by Alibaba Cloud Int. (Terms of Service)

39 models

Tokens processed on OpenRouter

  • Favicon for z-ai
    Z.ai: GLM 5.3 PrimeGLM 5.3 Prime

    GLM-5.3-Prime is the high-speed variant of Z.ai's GLM-5.3, inheriting its full capabilities while delivering 1.5–2× the output throughput through inference acceleration. It supports text input and output with a 1M-token context window and up to 128K output tokens, and targets coding and agentic workloads, including long-horizon multi-turn agent orchestration, real-time conversation, and streaming code generation. Reasoning is always on and cannot be disabled. Reasoning efforts low, high, and max are supported; max is the default.

    by z-aiSep 23, 20261M context$2.80/M input tokens$8.80/M output tokens
  • Favicon for qwen
    Qwen: Qwen3.8 Max PrimeQwen3.8 Max Prime

    Qwen3.8 Max Prime is a higher-throughput variant of Qwen3.8 Max from Alibaba's Qwen team, served as a separate SKU at a higher price point. It accepts text, image, and video input and returns text, with a 1M-token context window and reasoning enabled by default. Tool calling, structured outputs, and configurable reasoning effort are supported, matching Qwen3.8 Max.

    by qwenSep 23, 20261M context$4/M input tokens$12/M output tokens
  • Favicon for qwen
    Qwen: Qwen3.8 Omni FlashQwen3.8 Omni Flash

    Qwen3.8 Omni Flash is an omni-modal reasoning model from Alibaba, the first Qwen model built around agentic capabilities with native audio-video understanding. It is suited for audio-video analysis and summarization, video editing and production workflows, audio-video dialogue, coding, knowledge work, and GUI interaction, and it is particularly strong at long-form multimedia tasks that combine speech, sound, and visual context. It also supports two-channel and four-channel spatial audio understanding.

    by qwenSep 21, 20261M context$0.15/M input tokens$0.47/M output tokens
  • Favicon for deepseek
    DeepSeek: DeepSeek V4.1 FlashDeepSeek V4.1 Flash

    DeepSeek V4.1 Flash is a sparse mixture-of-experts model from DeepSeek, and the first built on the company's Causal Encoder-Decoder (CED) architecture. It activates 8B parameters on input and 16B on output from a 552B-parameter backbone, an asymmetric split that keeps per-token compute low relative to the model's total size. Image understanding is native to the architecture, with visual and text embeddings trained jointly from the start of pre-training rather than added afterward as in the earlier experimental V4 Flash Vision Exp. It is suited for coding, terminal, and computer-use agents, along with long-horizon tasks that must run to completion across many steps and long-context analysis. Compressed KV caching cuts cache memory to roughly a quarter of the previous Flash generation, significantly reducing costs on agentic workloads. DeepSeek positions it as the cost-efficient tier of the V4.1 family and reports that it exceeds V4 Pro on performance, speed, and task completion time.

    by deepseekSep 10, 20261.05M context$0.15/M input tokens$0.60/M output tokens
  • Favicon for qwen
    Qwen: Qwen3.8 Max (0902)Qwen3.8 Max (0902)

    Qwen3.8 Max 0902 is an updated snapshot of Qwen3.8 Max from Alibaba's Qwen team. It is a 2.4-trillion-parameter mixture-of-experts model that accepts text, image, and video input and returns text, with a 1M-token context window and reasoning enabled by default. This snapshot is post-trained for coding and agentic work, including multi-step software projects, multi-tool orchestration, and long-horizon task execution. It also targets chart reasoning, document parsing, and multimodal understanding over long documents and extended video. Tool calling, structured outputs, and configurable reasoning effort are supported.

    by qwenSep 3, 20261M context$2/M input tokens$6/M output tokens
  • Favicon for alibaba
    Alibaba: Wan 3.0 PrimeWan 3.0 Prime

    Wan 3.0 Prime is a fast-mode variant of Wan 3.0 from Alibaba. It supports text-to-video and first-frame image-to-video generation.

    by alibabaAug 27, 2026from $0.068/second
  • Favicon for qwen
    Qwen: Qwen3.8 FlashQwen3.8 Flash

    Qwen3.8 Flash is a multimodal reasoning model from Alibaba. It is suited for coding assistance, agentic workflows, visual understanding, document and codebase analysis, desktop interaction, chart analysis, and long-video analysis.

    by qwenAug 26, 20261M context$0.15/M input tokens$0.47/M output tokens
  • Favicon for alibaba
    Alibaba: Wan 3.0Wan 3.0
    15% off

    Wan 3.0 is a video generation model from Alibaba for text-to-video, image-to-video, and reference-guided video generation. It produces 480p, 720p, or 1080p video with durations from 2 to 30 seconds.

    by alibabaAug 24, 2026from $0.0425/second
  • Favicon for z-ai
    Z.ai: GLM 5.3GLM 5.3

    GLM-5.3 is a large-scale reasoning model from Z.ai, built for complex software engineering and long-horizon agent tasks. It supports text input and output with a 1M-token context window, and improves on GLM-5.2 in coding and in the balance between performance and token efficiency. Reasoning is always on and cannot be disabled. Reasoning efforts low, high, and max are supported; max is the default.

    by z-aiAug 18, 20261.05M context$1.19/M input tokens$3.74/M output tokens
  • Favicon for qwen
    Qwen: Qwen3.8 27BQwen3.8 27B

    Qwen3.8 27B is an open-weight dense vision-language model from Qwen. It is suited for coding, professional workflows, research, multimodal interaction, and long-running agent tasks, with flexible thinking that can be enabled or disabled.

    by qwenAug 14, 2026262K context$0.425/M input tokens$2.55/M output tokens
  • Favicon for qwen
    Qwen: Qwen3.8 2.4T A95BQwen3.8 2.4T A95B

    Qwen3.8 2.4T A95B is an open-weight sparse mixture-of-experts model from Qwen and the open-weight variant of Qwen3.8 Max, with 95 billion active parameters out of 2.4 trillion total. It is suited for coding, research, complex reasoning, and agentic workflows.

    by qwenAug 12, 20261M context$2/M input tokens$6/M output tokens
  • Favicon for deepseek
    DeepSeek: DeepSeek V4 Pro 0813DeepSeek V4 Pro 0813

    DeepSeek V4 Pro 0813 is a large-scale mixture-of-experts model from DeepSeek. This is the GA release of DeepSeek V4 Pro.

    by deepseekAug 12, 20261.05M context$0.5808/M input tokens$1.742/M output tokens
  • Favicon for qwen
    Qwen: Qwen Image 3 ProQwen Image 3 Pro

    Qwen Image 3 Pro is an image generation and editing model from Qwen. It supports precise rendering of text and details as small as 10px, along with richer world knowledge compared to previous generations.

    by qwenAug 5, 2026from $0.04/image
  • Favicon for qwen
    Qwen: Qwen Image 3Qwen Image 3

    Qwen Image 3 is a unified image generation and editing model from Qwen. It supports precise rendering of text and details as small as 10px, along with a richer world knowledge base than previous generations.

    by qwenAug 5, 2026from $0.03/image
  • Favicon for deepseek
    DeepSeek: DeepSeek V4 Flash 0731DeepSeek V4 Flash 0731

    DeepSeek V4 Flash 0731 is a sparse mixture-of-experts model from DeepSeek, with 13B active parameters out of 284B total. This re-post-trained revision is suited for coding, reasoning, and agent workflows. This is the GA release of DeepSeek V4 Flash.

    by deepseekJul 31, 20261.05M context$0.176/M input tokens$0.528/M output tokens
  • Favicon for qwen
    Qwen: Qwen3.7 FlashQwen3.7 Flash

    Qwen3.7 Flash is a vision-language reasoning model from Alibaba. It is suited for multimodal agents, visual coding, search, and computer interaction, with strengths in object recognition, spatial understanding, and real-world visual perception.

    by qwenJul 27, 20261M context$0.03/M input tokens$0.13/M output tokens
  • Favicon for qwen
    Qwen: Qwen-Audio-3.0-TTS FlashQwen-Audio-3.0-TTS Flash

    Qwen-Audio-3.0-TTS Flash is Alibaba's fast, cost-efficient text-to-speech model, generating spoken audio from text via the DashScope Speech Synthesizer API.

    by qwenJul 23, 2026$15/M characters
  • Favicon for qwen
    Qwen: Qwen-Audio-3.0-TTS PlusQwen-Audio-3.0-TTS Plus

    Qwen-Audio-3.0-TTS Plus is Alibaba's higher-quality text-to-speech model, generating spoken audio from text via the DashScope Speech Synthesizer API.

    by qwenJul 23, 2026$20/M characters
  • Favicon for moonshotai
    MoonshotAI: Kimi K3Kimi K3

    Kimi K3 is a 2.8T parameter open-weight multimodal reasoning model from Moonshot AI. It is suited for complex coding, knowledge work, and long-horizon agentic workflows, and is particularly strong at navigating large repositories, using tools, debugging, and iterating against images, logs, tests, and runtime feedback. Its architecture uses KDA and Attention Residuals for computational efficiency.

    by moonshotaiJul 16, 20261.05M context$3.45/M input tokens$17.25/M output tokens
  • Favicon for alibaba
    Alibaba: HappyHorse 1.1HappyHorse 1.1

    HappyHorse 1.1 is a video generation model from Alibaba. It generates short videos from a text prompt, a single starting image, or a set of reference images, with output up to 1080p and durations of 3 to 15 seconds. It is suited for creative content, social media clips, and image-driven animation, and improves on the prior version with stronger prompt adherence, smoother motion, and more consistent characters across frames.

    by alibabaJun 24, 2026from $0.0988/second
  • Favicon for alibaba
    Alibaba: HappyHorse 1.0HappyHorse 1.0

    HappyHorse 1.0 is a video generation model from Alibaba. It generates short videos from a text prompt, a single starting image, or a set of reference images, with output up to 1080p and durations of 3 to 15 seconds. It is suited for creative content, social media clips, and image-driven animation across a range of aspect ratios.

    by alibabaJun 24, 2026from $0.0988/second
  • Favicon for z-ai
    Z.ai: GLM 5.2GLM 5.2

    GLM 5.2 is a large-scale reasoning model from Z.ai. It supports text input and output with a 1M-token context window, and is suited for long-horizon agent workflows, project-level software engineering, and complex multi-step automation. Reasoning efforts high and xhigh are supported; xhigh maps to max reasoning. It is particularly strong at coding and tool use across long-running tasks, able to maintain engineering context and follow standards consistently through a full development workflow, from requirements to multi-platform deployment, in a single task.

    by z-aiJun 16, 20261.05M context$0.966/M input tokens$3.036/M output tokens
  • Favicon for moonshotai
    MoonshotAI: Kimi K2.7 CodeKimi K2.7 Code

    MoonshotAI: Kimi K2.7 Code is a coding-focused model in Moonshot AI's Kimi K2 family, built to complete end-to-end programming tasks reliably over long contexts. It uses a native multimodal mixture-of-experts architecture that accepts text and image input, and it always operates in a thinking mode, preserving full reasoning content across multi-turn conversations. With a 256K-token context window, it targets long-horizon coding, agentic task decomposition, and multi-turn dialogue. The model activates 32B parameters out of roughly 1T total.

    by moonshotaiJun 12, 2026262K context$0.95/M input tokens$4/M output tokens
  • Favicon for qwen
    Qwen: Qwen3.7 PlusQwen3.7 Plus

    Qwen3.7-Plus is a cost-effective model in Alibaba's Qwen3.7 series. It supports text and image input with text output, building on the series' text capabilities with a comprehensive upgrade to its vision-language abilities while retaining full-stack, agent-level intelligence for coding, tool use, and productivity workflows. Its distinguishing trait is multi-modal interactive hybrid agent capability: it can perceive real-world scenes, read screens and interact with GUIs, generate code from visual references, and perform end-to-end navigation within mobile apps.

    by qwenJun 3, 20261M context$0.32/M input tokens$1.28/M output tokens
  • Favicon for qwen
    Qwen: Qwen3.7 MaxQwen3.7 Max

    Qwen3.7-Max is the flagship model in Alibaba's Qwen3.7 series. It supports text input and output and is designed for agent-centric workloads, with particular strengths in coding, office and productivity tasks, and long-horizon autonomous execution. The model offers notable gains in coding and agentic performance over prior Qwen generations and supports explicit prompt caching for efficient repeated context use.

    by qwenMay 21, 20261M context$1.475/M input tokens$4.425/M output tokens
  • Favicon for qwen
    Qwen: Qwen3.5 Plus 2026-04-20Qwen3.5 Plus 2026-04-20

    Qwen3.5 Plus (April 2026) is a large-scale multimodal language model from Alibaba. It accepts text, image, and video input and produces text output, with a 1M token context window. This is an updated version of Qwen3.5 Plus with tiered pricing above 256K tokens.

    by qwenApr 27, 20261M context$0.30/M input tokens$1.80/M output tokens
  • Favicon for qwen
    Qwen: Qwen3.6 FlashQwen3.6 Flash

    Qwen3.6 Flash is a fast, efficient language model from Alibaba's Qwen 3.6 series. It supports text, image, and video input with a 1M token context window. Tiered pricing kicks in above 256K tokens. Prompt caching is supported, with both explicit cache read and cache creation pricing.

    by qwenApr 27, 20261M context$0.1875/M input tokens$1.125/M output tokens
  • Favicon for qwen
    Qwen: Qwen3.6 27BQwen3.6 27B

    Qwen3.6 27B is a dense 27-billion-parameter language model from the Qwen Team at Alibaba, released in April 2026. It features hybrid multimodal capabilities — accepting text, image, and video inputs — and supports a 262,144-token context window. The model is designed for agentic coding and reasoning tasks, with particular strength in repository-level code comprehension, front-end development workflows, and multi-step problem solving. It includes a built-in thinking mode for extended reasoning and preserves thinking context across conversation history. Qwen3.6 27B supports 201 languages and dialects and is released under the Apache 2.0 license.

    by qwenApr 27, 2026262K context$0.45/M input tokens$2.70/M output tokens
  • Favicon for deepseek
    DeepSeek: DeepSeek V4 Pro 0423DeepSeek V4 Pro 0423

    DeepSeek V4 Pro is a large-scale Mixture-of-Experts model from DeepSeek with 1.6T total parameters and 49B activated parameters, supporting a 1M-token context window. It is designed for advanced reasoning, coding, and long-horizon agent workflows, with strong performance across knowledge, math, and software engineering benchmarks. Built on the same architecture as DeepSeek V4 Flash, it introduces a hybrid attention system for efficient long-context processing. Reasoning efforts high and xhigh are supported; xhigh maps to max reasoning. It is well suited for complex workloads such as full-codebase analysis, multi-step automation, and large-scale information synthesis, where both capability and efficiency are critical

    by deepseekApr 24, 20261.05M context$1.416/M input tokens$2.832/M output tokens
  • Favicon for z-ai
    Z.ai: GLM 5.1GLM 5.1

    GLM-5.1 delivers a major leap in coding capability, with particularly significant gains in handling long-horizon tasks. Unlike previous models built around minute-level interactions, GLM-5.1 can work independently and continuously on a single task for more than 8 hours, autonomously planning, executing, and improving itself throughout the process, ultimately delivering complete, engineering-grade results.

    by z-aiApr 7, 2026203K context$1.33/M input tokens$4.18/M output tokens
  • Favicon for qwen
    Qwen: Qwen3.6 PlusQwen3.6 Plus

    Qwen 3.6 Plus builds on a hybrid architecture that combines efficient linear attention with sparse mixture-of-experts routing, enabling strong scalability and high-performance inference. Compared to the 3.5 series, it delivers major gains in agentic coding, front-end development, and overall reasoning, with a significantly improved “vibe coding” experience. The model excels at complex tasks such as 3D scenes, games, and repository-level problem solving, achieving a 78.8 score on SWE-bench Verified. It represents a substantial leap in both pure-text and multimodal capabilities, performing at the level of leading state-of-the-art models.

    by qwenApr 2, 20261M context$0.325/M input tokens$1.95/M output tokens
  • Favicon for qwen
    Qwen: Qwen3.5-35B-A3BQwen3.5-35B-A3B

    The Qwen3.5 Series 35B-A3B is a native vision-language model designed with a hybrid architecture that integrates linear attention mechanisms and a sparse mixture-of-experts model, achieving higher inference efficiency. Its overall performance is comparable to that of the Qwen3.5-27B.

    by qwenFeb 25, 2026256K context$0.1625/M input tokens$1.30/M output tokens
  • Favicon for qwen
    Qwen: Qwen3.5-27BQwen3.5-27B

    The Qwen3.5 27B native vision-language Dense model incorporates a linear attention mechanism, delivering fast response times while balancing inference speed and performance. Its overall capabilities are comparable to those of the Qwen3.5-122B-A10B.

    by qwenFeb 25, 2026256K context$0.195/M input tokens$1.56/M output tokens
  • Favicon for qwen
    Qwen: Qwen3.5-122B-A10BQwen3.5-122B-A10B

    The Qwen3.5 122B-A10B native vision-language model is built on a hybrid architecture that integrates a linear attention mechanism with a sparse mixture-of-experts model, achieving higher inference efficiency. In terms of overall performance, this model is second only to Qwen3.5-397B-A17B. Its text capabilities significantly outperform those of Qwen3-235B-2507, and its visual capabilities surpass those of Qwen3-VL-235B.

    by qwenFeb 25, 2026262K context$0.26/M input tokens$2.08/M output tokens
  • Favicon for qwen
    Qwen: Qwen3.5-FlashQwen3.5-Flash

    The Qwen3.5 native vision-language Flash models are built on a hybrid architecture that integrates a linear attention mechanism with a sparse mixture-of-experts model, achieving higher inference efficiency. Compared to the 3 series, these models deliver a leap forward in performance for both pure text and multimodal tasks, offering fast response times while balancing inference speed and overall performance.

    by qwenFeb 25, 20261M context$0.065/M input tokens$0.26/M output tokens
  • Favicon for qwen
    Qwen: Qwen3.5 Plus 2026-02-15Qwen3.5 Plus 2026-02-15

    The Qwen3.5 native vision-language series Plus models are built on a hybrid architecture that integrates linear attention mechanisms with sparse mixture-of-experts models, achieving higher inference efficiency. In a variety of task evaluations, the 3.5 series consistently demonstrates performance on par with state-of-the-art leading models. Compared to the 3 series, these models show a leap forward in both pure-text and multimodal capabilities.

    by qwenFeb 16, 20261M context$0.26/M input tokens$1.56/M output tokens
  • Favicon for qwen
    Qwen: Qwen3.5 397B A17BQwen3.5 397B A17B

    The Qwen3.5 series 397B-A17B native vision-language model is built on a hybrid architecture that integrates a linear attention mechanism with a sparse mixture-of-experts model, achieving higher inference efficiency. It delivers state-of-the-art performance comparable to leading-edge models across a wide range of tasks, including language understanding, logical reasoning, code generation, agent-based tasks, image understanding, video understanding, and graphical user interface (GUI) interactions. With its robust code-generation and agent capabilities, the model exhibits strong generalization across diverse agent.

    by qwenFeb 16, 2026256K context$0.39/M input tokens$2.34/M output tokens
  • Favicon for qwen
    Qwen: Qwen3 Coder FlashQwen3 Coder Flash

    Qwen3 Coder Flash is Alibaba's fast and cost efficient version of their proprietary Qwen3 Coder Plus. It is a powerful coding agent model specializing in autonomous programming via tool calling and environment interaction, combining coding proficiency with versatile general-purpose abilities.

    by qwenSep 17, 2025128K context$0.195/M input tokens$0.975/M output tokens
  • Favicon for qwen
    Qwen: Qwen-PlusQwen-Plus

    Qwen-Plus, based on the Qwen2.5 foundation model, is a 131K context model with a balanced performance, speed, and cost combination.

    by qwenFeb 1, 2025131K context$0.26/M input tokens$0.78/M output tokens