AI Models

466 models Free & Paid Update: 2 hours trước

Nex-N2.5 is an agentic model built to turn goals into working, verified outcomes. Its core strength is agentic coding within a visual feedback loop: it can explore codebases, implement multi-file...

by |Sep 2026 |262K context |$0.0250/M input |$0.1000/M output
262K tokens ⓘ

Nex-N2.5 is an agentic model built to turn goals into working, verified outcomes. Its core strength is agentic coding within a visual feedback loop: it can explore codebases, implement multi-file...

by |Sep 2026 |262K context |$0.0750/M input |$0.2500/M output
262K tokens ⓘ

GPT-6 Astra is OpenAI's flagship model for demanding end-to-end work. It is suited for advanced analysis, software engineering, deep research, scientific work, and document creation, with particular strengths in long-horizon...

by |Sep 2026 |1.1M context |$10.00/M input |$50.00/M output
1.1M tokens ⓘ

GPT-6 Astra is OpenAI's flagship model for demanding end-to-end work. It is suited for advanced analysis, software engineering, deep research, scientific work, and document creation, with particular strengths in long-horizon...

by |Sep 2026 |1.1M context |$5.00/M input |$25.00/M output
1.1M tokens ⓘ

GPT-6 Astra Pro is the same underlying model as [GPT-6 Astra](https://openrouter.ai/openai/gpt-6-astra), served with `reasoning.mode` set to `pro` for higher-quality responses on complex tasks. **Cost note:** pro mode spends far more...

by |Sep 2026 |1.1M context |$10.00/M input |$50.00/M output
1.1M tokens ⓘ

GPT-6 Astra Pro is the same underlying model as [GPT-6 Astra](https://openrouter.ai/openai/gpt-6-astra), served with `reasoning.mode` set to `pro` for higher-quality responses on complex tasks. **Cost note:** pro mode spends far more...

by |Sep 2026 |1.1M context |$5.00/M input |$25.00/M output
1.1M tokens ⓘ

Ling 3.0 Flash Sante is a health and medicine-focused mixture-of-experts model from InclusionAI, built on Ling 3.0 Flash with 5.1B active parameters out of 124B total. It is designed for...

by |Sep 2026 |262K context |Miễn phí input |Miễn phí output
262K tokens ⓘ

Qwen3.8 Max 0902 is an updated snapshot of Qwen3.8 Max from Alibaba's Qwen team. It is a 2.4-trillion-parameter mixture-of-experts model that accepts text, image, and video input and returns text,...

by |Sep 2026 |1M context |$2.00/M input |$6.00/M output
1M tokens ⓘ

Muse Spark 1.3 Contributor is the cost-efficient contributor tier of Meta’s multimodal reasoning model for experimentation, learning, and early-stage agentic, multi-agent, and coding workflows. It is designed to track information...

by |Sep 2026 |1M context |$0.1000/M input |$0.2000/M output
1M tokens ⓘ

Muse Spark 1.3 is a multimodal reasoning model from Meta for long-running agentic, multi-agent, and coding workflows. It is designed to keep track of information across extended tasks, work through...

by |Sep 2026 |1M context |$1.25/M input |$4.25/M output
1M tokens ⓘ

Gemini 3.8 Flash is Google's most intelligent Flash model with significant gains from 3.7 Flash across software engineering, agentic tasks, and multi-step reasoning.

by |Sep 2026 |1M context |$0.7500/M input |$3.75/M output
1M tokens ⓘ

Gemini 3.8 Flash is Google's most intelligent Flash model with significant gains from 3.7 Flash across software engineering, agentic tasks, and multi-step reasoning.

by |Sep 2026 |1M context |$0.3750/M input |$1.88/M output
1M tokens ⓘ

Claude Fable 5.1 improves on Claude Fable 5 across the board, with the biggest gains in agentic coding, long-running agentic workflows, and knowledge work: long code refactors, front-end and visual...

by |Sep 2026 |1M context |$10.00/M input |$50.00/M output
1M tokens ⓘ

Claude Fable 5.1 improves on Claude Fable 5 across the board, with the biggest gains in agentic coding, long-running agentic workflows, and knowledge work: long code refactors, front-end and visual...

by |Sep 2026 |1M context |$5.00/M input |$25.00/M output
1M tokens ⓘ

Granite 4.2 8B is a dense reasoning model from IBM. It is suited for mathematics, code generation, multilingual dialogue, and agentic workflows that need multi-step reasoning. It supports full, low-effort,...

by |Aug 2026 |131K context |$0.0600/M input |$0.2500/M output
131K tokens ⓘ

Tencent: Hy4 preview is a mixture-of-experts model from Tencent, with 49B active parameters out of 770B total. It is designed for coding agents, complex tool-use workflows, and productivity tasks that...

by |Aug 2026 |1M context |$0.7506/M input |$2.25/M output
1M tokens ⓘ

Ling 3.0 Flash Fin is a finance-focused mixture-of-experts model from InclusionAI, built on Ling 3.0 Flash with 5.1B active parameters out of 124B total. It is designed for real-world investment...

by |Aug 2026 |262K context |$0.0420/M input |$0.1232/M output
262K tokens ⓘ

This model always redirects to the latest model in the GLM Flash family.

by |Aug 2026 |1M context |$0.0263/M input |$0.9287/M output
1M tokens ⓘ

Qwen3.8 Flash is a multimodal reasoning model from Alibaba. It is suited for coding assistance, agentic workflows, visual understanding, document and codebase analysis, desktop interaction, chart analysis, and long-video analysis.

by |Aug 2026 |1M context |$0.1500/M input |$0.4700/M output
1M tokens ⓘ

GLM-5.3-Flash is a native multimodal model from Z.ai. It is suited for efficient coding and long-horizon agent tasks. Its hybrid sparse and linear attention architecture maintains accurate long-context behavior while...

by |Aug 2026 |1M context |$0.1500/M input |$0.5000/M output
1M tokens ⓘ

GLM-5.3-Flash is a native multimodal model from Z.ai. It is suited for efficient coding and long-horizon agent tasks. Its hybrid sparse and linear attention architecture maintains accurate long-context behavior while...

by |Aug 2026 |1M context |$0.0600/M input |$0.2000/M output
1M tokens ⓘ

Muse Spark 1.2 contributor tier is a reasoning model from Meta designed for developers who want to start building at an even lower cost. It’s meaningfully cheaper than Muse Spark...

by |Aug 2026 |1M context |$0.1000/M input |$0.2000/M output
1M tokens ⓘ

DeepSeek V4 Flash Vision Exp is an experimental vision-enabled version of [DeepSeek V4 Flash 0731](https://openrouter.ai/deepseek/deepseek-v4-flash-0731) from DeepSeek, adding image understanding while matching the base model on text capabilities including agents,...

by |Aug 2026 |1M context |$0.2156/M input |$0.6468/M output
1M tokens ⓘ

Hy-MT2-1.8B is a compact 1.8B-parameter translation model from Tencent. It supports 33 language pairs and five Chinese dialect and minority-language pairs, with workflows for structured, delimiter-based, contextual, glossary-based, and style-guided...

by |Aug 2026 |8K context |$0.0440/M input |$0.1770/M output

Hy-MT2-30B-A3B is Tencent's flagship translation model in the Hy-MT2 family. It supports 33 language pairs and five Chinese dialect and minority-language pairs, with workflows for structured, delimiter-based, contextual, glossary-based, and...

by |Aug 2026 |8K context |$0.0740/M input |$0.2950/M output

This model always redirects to the latest GLM model from Z.ai.

by |Aug 2026 |1M context |$0.1200/M input |$4.00/M output
1M tokens ⓘ

Hy-MT2-7B is a 7B-parameter translation model from Tencent. It supports 33 language pairs and five Chinese dialect and minority-language pairs, with workflows for structured, delimiter-based, contextual, glossary-based, and style-guided translation.

by |Aug 2026 |8K context |$0.0740/M input |$0.2950/M output

GLM-5.3 is a large-scale reasoning model from Z.ai, built for complex software engineering and long-horizon agent tasks. It supports text input and output with a 1M-token context window, and improves...

by |Aug 2026 |1M context |$0.4500/M input |$2.00/M output
1M tokens ⓘ

GLM-5.3 is a large-scale reasoning model from Z.ai, built for complex software engineering and long-horizon agent tasks. It supports text input and output with a 1M-token context window, and improves...

by |Aug 2026 |1M context |$1.40/M input |$4.40/M output
1M tokens ⓘ

Qwen3.8 27B is an open-weight dense vision-language model from Qwen. It is suited for coding, professional workflows, research, multimodal interaction, and long-running agent tasks, with flexible thinking that can be...

by |Aug 2026 |262K context |Miễn phí input |Miễn phí output
262K tokens ⓘ

Qwen3.8 27B is an open-weight dense vision-language model from Qwen. It is suited for coding, professional workflows, research, multimodal interaction, and long-running agent tasks, with flexible thinking that can be...

by |Aug 2026 |1M context |$0.4200/M input |$3.00/M output
1M tokens ⓘ

Dots3-Note Preview is an open-weight mixture-of-experts model from Dots Studio, with 16B active parameters out of 280B total. It is the lightest model in the Dots 3 family and is...

by |Aug 2026 |512K context |Miễn phí input |Miễn phí output
512K tokens ⓘ

Gemini 3.7 Flash is a multimodal model from Google for fast agentic workflows, coding, and complex multi-step reasoning. It is designed for tasks that require responsive performance and reliable multi-step...

by |Aug 2026 |1M context |$0.7500/M input |$3.75/M output
1M tokens ⓘ

Gemini 3.7 Flash is a multimodal model from Google for fast agentic workflows, coding, and complex multi-step reasoning. It is designed for tasks that require responsive performance and reliable multi-step...

by |Aug 2026 |1M context |$0.3750/M input |$1.88/M output
1M tokens ⓘ

Seed 2.1 Turbo is a multimodal model from ByteDance Seed for coding and long-horizon agent workflows. It is suited for end-to-end software delivery, multi-step task execution, and understanding visual and...

by |Aug 2026 |262K context |$0.5000/M input |$2.50/M output
262K tokens ⓘ

Qwen3.8 2.4T A95B is an open-weight sparse mixture-of-experts model from Qwen and the open-weight variant of [Qwen3.8 Max](/qwen/qwen3.8-max), with 95 billion active parameters out of 2.4 trillion total. It is...

by |Aug 2026 |1M context |$2.00/M input |$6.00/M output
1M tokens ⓘ

Seed 2.0 Code is a model from ByteDance Seed optimized for agentic coding. It is suited for frontend development, multilingual programming tasks, and coding-agent workflows in tools such as Claude...

by |Aug 2026 |262K context |$0.5000/M input |$3.00/M output
262K tokens ⓘ

DeepSeek V4 Pro 0813 is a large-scale mixture-of-experts model from DeepSeek. This is the GA release of DeepSeek V4 Pro.

by |Aug 2026 |1M context |$0.6600/M input |$1.98/M output
1M tokens ⓘ

Grok 4.6 is a model from SpaceXAI with frontier performance on coding, knowledge work, and STEM. It is succeeded by [Grok 4.7](/x-ai/grok-4.7).

by |Aug 2026 |500K context |$2.00/M input |$6.00/M output
500K tokens ⓘ

LFM2.5-2.6B is a compact reasoning model from Liquid AI. It is suited for agent workflows, data extraction, RAG, and long-context processing. Liquid advises against using it for agentic coding or...

by |Aug 2026 |66K context |Miễn phí input |Miễn phí output
66K tokens ⓘ

NVIDIA Nemotron 3.5 Lightning is an open mixture-of-experts model from NVIDIA, with 3B active parameters out of 30B total. It is suited for high-throughput agentic workloads and specialized tasks that...

by |Aug 2026 |262K context |$0.0600/M input |$0.1600/M output
262K tokens ⓘ

NVIDIA Nemotron 3.5 Lightning is an open mixture-of-experts model from NVIDIA, with 3B active parameters out of 30B total. It is suited for high-throughput agentic workloads and specialized tasks that...

by |Aug 2026 |1M context |Miễn phí input |Miễn phí output
1M tokens ⓘ

Sakana Namazu is a Japanese-specialized reasoning model from Sakana AI, based on Kimi K2.6 with additional training for Japanese language and business contexts. It is suited for Japanese instruction following,...

by |Aug 2026 |262K context |$0.9500/M input |$4.00/M output
262K tokens ⓘ

Solar Pro 4 is Upstage's cost-efficient large language model, featuring a 524K context window. It is built for long-horizon tasks and agentic workflows, with strong capabilities in office productivity, document-intensive...

by |Aug 2026 |524K context |$0.0900/M input |$0.3600/M output
524K tokens ⓘ

Muse Glimmer 30B is a dense, open-weight multimodal model from Meta Superintelligence Labs, distilled from Muse Spark and optimized for autonomous agents on consumer hardware. It is suited for long-horizon...

by |Aug 2026 |131K context |$0.3500/M input |$1.50/M output
131K tokens ⓘ

Muse Spark 1.2 is a reasoning model from Meta, designed for complex agentic tasks. It accepts text, images, video, and PDF documents, returns text, and offers a 1M-token context window....

by |Aug 2026 |1M context |$1.25/M input |$4.25/M output
1M tokens ⓘ

DeepSeek V4 Flash 0731 is a sparse mixture-of-experts model from DeepSeek, with 13B active parameters out of 284B total. This re-post-trained revision is suited for coding, reasoning, and agent workflows....

by |Jul 2026 |1M context |$0.0115/M input |$1.28/M output
1M tokens ⓘ

Inkling Small is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 12B active parameters out of 276B total. It is positioned as the smaller, more efficient member of...

by |Jul 2026 |524K context |$0.4500/M input |$1.20/M output
524K tokens ⓘ

Inkling Small is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 12B active parameters out of 276B total. It is positioned as the smaller, more efficient member of...

by |Jul 2026 |1M context |Miễn phí input |Miễn phí output
1M tokens ⓘ