AI Models

412 models Free & Paid Cập nhật: 4 giờ trước

Qwen3.8 27B is an open-weight dense vision-language model from Qwen. It is suited for coding, professional workflows, research, multimodal interaction, and long-running agent tasks, with flexible thinking that can be...

by |Th8 2026 |262K context |$0.4500/M input |$3.20/M output
262K tokens

Dots3-Note Preview is an open-weight mixture-of-experts model from Dots Studio, with 16B active parameters out of 280B total. It is the lightest model in the Dots 3 family and is...

by |Th8 2026 |512K context |Miễn phí input |Miễn phí output
512K tokens

Gemini 3.7 Flash is a multimodal model from Google for fast agentic workflows, coding, and complex multi-step reasoning. It is designed for tasks that require responsive performance and reliable multi-step...

by |Th8 2026 |1M context |$0.1875/M input |$0.9375/M output
1M tokens

Gemini 3.7 Flash is a multimodal model from Google for fast agentic workflows, coding, and complex multi-step reasoning. It is designed for tasks that require responsive performance and reliable multi-step...

by |Th8 2026 |1M context |$0.3750/M input |$1.88/M output
1M tokens

Seed 2.1 Turbo is a multimodal model from ByteDance Seed for coding and long-horizon agent workflows. It is suited for end-to-end software delivery, multi-step task execution, and understanding visual and...

by |Th8 2026 |262K context |$0.5000/M input |$2.50/M output
262K tokens

Qwen3.8 2.4T A95B is an open-weight sparse mixture-of-experts model from Qwen and the open-weight variant of [Qwen3.8 Max](/qwen/qwen3.8-max), with 95 billion active parameters out of 2.4 trillion total. It is...

by |Th8 2026 |1M context |$2.00/M input |$6.00/M output
1M tokens

Seed 2.0 Code is a model from ByteDance Seed optimized for agentic coding. It is suited for frontend development, multilingual programming tasks, and coding-agent workflows in tools such as Claude...

by |Th8 2026 |262K context |$0.5000/M input |$3.00/M output
262K tokens

DeepSeek V4 Pro 0813 is a large-scale mixture-of-experts model from DeepSeek. This is the GA release of DeepSeek V4 Pro.

by |Th8 2026 |1M context |$1.32/M input |$3.96/M output
1M tokens

Grok 4.6 is SpaceXAI's smartest model with frontier performance on coding, knowledge work, and STEM.

by |Th8 2026 |500K context |$2.00/M input |$6.00/M output
500K tokens

NVIDIA Nemotron 3.5 Lightning is an open mixture-of-experts model from NVIDIA, with 3B active parameters out of 30B total. It is suited for high-throughput agentic workloads and specialized tasks that...

by |Th8 2026 |1M context |$0.0800/M input |$0.2000/M output
1M tokens

NVIDIA Nemotron 3.5 Lightning is an open mixture-of-experts model from NVIDIA, with 3B active parameters out of 30B total. It is suited for high-throughput agentic workloads and specialized tasks that...

by |Th8 2026 |1M context |Miễn phí input |Miễn phí output
1M tokens

Sakana Namazu is a Japanese-specialized reasoning model from Sakana AI, based on Kimi K2.6 with additional training for Japanese language and business contexts. It is suited for Japanese instruction following,...

by |Th8 2026 |262K context |$0.9500/M input |$4.00/M output
262K tokens

Solar Pro 4 is Upstage's cost-efficient large language model, featuring a 524K context window. It is built for long-horizon tasks and agentic workflows, with strong capabilities in office productivity, document-intensive...

by |Th8 2026 |524K context |$0.0300/M input |$0.1200/M output
524K tokens

Muse Glimmer 30B is a dense, open-weight multimodal model from Meta Superintelligence Labs, distilled from Muse Spark and optimized for autonomous agents on consumer hardware. It is suited for long-horizon...

by |Th8 2026 |131K context |$0.3500/M input |$1.50/M output
131K tokens

Muse Spark 1.2 is a reasoning model from Meta, designed for complex agentic tasks. It accepts text, images, video, audio, and PDF documents, returns text, and offers a 1M-token context...

by |Th8 2026 |1M context |$1.25/M input |$4.25/M output
1M tokens

Qwen3.8 Max is the flagship model in Alibaba's Qwen3.8 series, the general-availability successor to the Qwen3.8 Max Preview. It is a multimodal reasoning model intended for complex reasoning, visual understanding,...

by |Th8 2026 |1M context |$2.00/M input |$6.00/M output
1M tokens

This model always redirects to the latest model in the DeepSeek V4 Flash family.

by |Th8 2026 |1.3M context |$0.0786/M input |$0.1572/M output
1.3M tokens

DeepSeek V4 Flash 0731 is a sparse mixture-of-experts model from DeepSeek, with 13B active parameters out of 284B total. This re-post-trained revision is suited for coding, reasoning, and agent workflows....

by |Th7 2026 |1.3M context |$0.1400/M input |$0.2800/M output
1.3M tokens

Inkling Small is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 12B active parameters out of 276B total. It is positioned as the smaller, more efficient member of...

by |Th7 2026 |524K context |$0.4500/M input |$1.20/M output
524K tokens

Qwen3.7 Flash is a vision-language reasoning model from Alibaba. It is suited for multimodal agents, visual coding, search, and computer interaction, with strengths in object recognition, spatial understanding, and real-world...

by |Th7 2026 |1M context |$0.0300/M input |$0.1300/M output
1M tokens

Fast-mode variant of [Opus 5](/anthropic/claude-opus-5) - identical capabilities with higher output speed at 2x pricing relative to regular Opus 5. Learn more in Anthropic's docs: https://platform.claude.com/docs/en/build-with-claude/fast-mode

by |Th7 2026 |1M context |$10.00/M input |$50.00/M output
1M tokens

Claude Opus 5 is Anthropic’s flagship model for demanding reasoning, coding, and long-horizon agentic work. It is particularly strong at end-to-end software tasks, code review and bug finding, visual analysis...

by |Th7 2026 |1M context |$2.50/M input |$12.50/M output
1M tokens

Claude Opus 5 is Anthropic’s flagship model for demanding reasoning, coding, and long-horizon agentic work. It is particularly strong at end-to-end software tasks, code review and bug finding, visual analysis...

by |Th7 2026 |1M context |$5.00/M input |$25.00/M output
1M tokens

*Ling-3.0-flash* is a *124B-parameter Mixture-of-Experts (MoE) model*, with approximately *5.1B parameters activated per token*. The model is designed with *token efficiency and production-scale agentic inference* as key priorities, enabling developers...

by |Th7 2026 |262K context |$0.0210/M input |$0.0630/M output
262K tokens

Laguna S 2.1 is the latest coding agent model from [Poolside](). Laguna S 2.1 is a 118B total parameter model with 8B active parameters, scoring 70.2% on Terminal-Bench 2.1 and...

by |Th7 2026 |1M context |$0.0900/M input |$0.1800/M output
1M tokens

Laguna S 2.1 is the latest coding agent model from [Poolside](). Laguna S 2.1 is a 118B total parameter model with 8B active parameters, scoring 70.2% on Terminal-Bench 2.1 and...

by |Th7 2026 |262K context |Miễn phí input |Miễn phí output
262K tokens

Gemini 3.6 Flash is a high-efficiency model from Google for coding, agentic workflows, and web and app development. It is designed to produce polished outputs with fewer unnecessary edits and...

by |Th7 2026 |1M context |$0.7500/M input |$3.75/M output
1M tokens

Gemini 3.6 Flash is a high-efficiency model from Google for coding, agentic workflows, and web and app development. It is designed to produce polished outputs with fewer unnecessary edits and...

by |Th7 2026 |1M context |$0.3750/M input |$1.88/M output
1M tokens

Gemini 3.5 Flash Lite is a high-efficiency model from Google with upgraded agentic capabilities. It is suited for subagents that execute focused tasks within complex, multi-agent workflows.

by |Th7 2026 |1M context |$0.3000/M input |$2.50/M output
1M tokens

Gemini 3.5 Flash Lite is a high-efficiency model from Google with upgraded agentic capabilities. It is suited for subagents that execute focused tasks within complex, multi-agent workflows.

by |Th7 2026 |1M context |$0.1500/M input |$1.25/M output
1M tokens

LongCat 2.0 is a sparse mixture-of-experts language model from Meituan, with 48B active parameters out of 1.6T total. It is suited for coding, repository-level changes, long-horizon problem solving, and agentic...

by |Th7 2026 |1M context |$0.3000/M input |$1.20/M output
1M tokens

Inkling is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 41B active parameters out of 975B total. It is designed for general-purpose reasoning, coding, agentic and tool-use systems,...

by |Th7 2026 |1M context |$0.9500/M input |$4.05/M output
1M tokens

Inkling is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 41B active parameters out of 975B total. It is designed for general-purpose reasoning, coding, agentic and tool-use systems,...

by |Th7 2026 |524K context |$0.5000/M input |$2.03/M output
524K tokens

Auto Router (Beta) is a task-aware router from OpenRouter. It classifies each request, then routes it the [most popular model](/rankings#task-spend) for that task based on aggregate spend, filtered by your...

by |Th7 2026 |2M context |Miễn phí input |Miễn phí output
2M tokens

Kimi K3 is a 2.8T parameter open-weight multimodal reasoning model from Moonshot AI. It is suited for complex coding, knowledge work, and long-horizon agentic workflows, and is particularly strong at...

by |Th7 2026 |1M context |$3.00/M input |$15.00/M output
1M tokens

Muse Spark 1.1 is a multimodal reasoning model from Meta, built for agentic tasks. It accepts text, images, video, audio, and PDF documents and returns text, with a 1M-token context...

by |Th7 2026 |1M context |$1.25/M input |$4.25/M output
1M tokens

KAT-Coder-Air V2.5 is a flagship-level Agentic Coding model that can directly hand over an entire issue or an entire business workflow to it, allowing it to autonomously locate and make...

by |Th7 2026 |256K context |$0.1500/M input |$0.6000/M output
256K tokens

KAT-Coder-Pro V2.5 is a flagship-level Agentic Coding model that can directly hand over an entire issue or an entire business workflow to it, allowing it to autonomously locate and make...

by |Th7 2026 |256K context |$0.7400/M input |$2.96/M output
256K tokens

GPT-5.6 Luna Pro is the same underlying model as [GPT-5.6 Luna](https://openrouter.ai/openai/gpt-5.6-luna), served with `reasoning.mode` set to `pro` for higher-quality responses on complex tasks. Learn more in OpenAI's docs: https://developers.openai.com/api/docs/guides/reasoning#reasoning-mode

by |Th7 2026 |1.1M context |$0.1000/M input |$0.6000/M output
1.1M tokens

GPT-5.6 Luna Pro is the same underlying model as [GPT-5.6 Luna](https://openrouter.ai/openai/gpt-5.6-luna), served with `reasoning.mode` set to `pro` for higher-quality responses on complex tasks. Learn more in OpenAI's docs: https://developers.openai.com/api/docs/guides/reasoning#reasoning-mode

by |Th7 2026 |1.1M context |$0.2000/M input |$1.20/M output
1.1M tokens

GPT-5.6 Luna is a fast, cost-efficient model in OpenAI's GPT-5.6 series. It is suited for high-volume, latency-sensitive tasks such as chat, classification, and lightweight agentic workflows, providing capable reasoning for...

by |Th7 2026 |1.1M context |$0.1000/M input |$0.6000/M output
1.1M tokens

GPT-5.6 Luna is a fast, cost-efficient model in OpenAI's GPT-5.6 series. It is suited for high-volume, latency-sensitive tasks such as chat, classification, and lightweight agentic workflows, providing capable reasoning for...

by |Th7 2026 |1.1M context |$0.2000/M input |$1.20/M output
1.1M tokens

GPT-5.6 Terra Pro is the same underlying model as [GPT-5.6 Terra](https://openrouter.ai/openai/gpt-5.6-terra), served with `reasoning.mode` set to `pro` for higher-quality responses on complex tasks. Learn more in OpenAI's docs: https://developers.openai.com/api/docs/guides/reasoning#reasoning-mode

by |Th7 2026 |1.1M context |$2.00/M input |$12.00/M output
1.1M tokens

GPT-5.6 Terra Pro is the same underlying model as [GPT-5.6 Terra](https://openrouter.ai/openai/gpt-5.6-terra), served with `reasoning.mode` set to `pro` for higher-quality responses on complex tasks. Learn more in OpenAI's docs: https://developers.openai.com/api/docs/guides/reasoning#reasoning-mode

by |Th7 2026 |1.1M context |$1.00/M input |$6.00/M output
1.1M tokens

GPT-5.6 Terra is a balanced model in OpenAI's GPT-5.6 series, positioned between the flagship Sol tier and the cost-efficient Luna tier. It is suited for everyday coding, reasoning, and agentic...

by |Th7 2026 |1.1M context |$1.00/M input |$6.00/M output
1.1M tokens

GPT-5.6 Terra is a balanced model in OpenAI's GPT-5.6 series, positioned between the flagship Sol tier and the cost-efficient Luna tier. It is suited for everyday coding, reasoning, and agentic...

by |Th7 2026 |1.1M context |$2.00/M input |$12.00/M output
1.1M tokens

GPT-5.6 Sol Pro is the same underlying model as [GPT-5.6 Sol](https://openrouter.ai/openai/gpt-5.6-sol), served with `reasoning.mode` set to `pro` for higher-quality responses on complex tasks. Learn more in OpenAI's docs: https://developers.openai.com/api/docs/guides/reasoning#reasoning-mode

by |Th7 2026 |1.1M context |$1.25/M input |$7.50/M output
1.1M tokens

GPT-5.6 Sol Pro is the same underlying model as [GPT-5.6 Sol](https://openrouter.ai/openai/gpt-5.6-sol), served with `reasoning.mode` set to `pro` for higher-quality responses on complex tasks. Learn more in OpenAI's docs: https://developers.openai.com/api/docs/guides/reasoning#reasoning-mode

by |Th7 2026 |1.1M context |$2.50/M input |$15.00/M output
1.1M tokens

GPT-5.6 Sol is the flagship model in OpenAI's GPT-5.6 series. It is suited for complex reasoning, coding, and agentic workflows, and is particularly strong at command-line and multi-step coding tasks...

by |Th7 2026 |1.1M context |$2.50/M input |$15.00/M output
1.1M tokens

GPT-5.6 Sol is the flagship model in OpenAI's GPT-5.6 series. It is suited for complex reasoning, coding, and agentic workflows, and is particularly strong at command-line and multi-step coding tasks...

by |Th7 2026 |1.1M context |$1.25/M input |$7.50/M output
1.1M tokens