AI Models

464 models Free & Paid Cập nhật: 3 giờ trước

Apodex 1.1 Mini is a reasoning-first model from Apodex, built for complex, long-horizon research and forecasting tasks. It works directly with files, data, code, and tools to produce verifiable results,...

by |Th10 2026 |262K context |Miễn phí input |Miễn phí output
262K tokens ⓘ

Pareto is a multimodal composite model built for research, coding, and agentic workflows, while delivering frontier-level performance across a broad range of general-purpose tasks. This is a preview of the...

by |Th10 2026 |1M context |$0.8000/M input |$3.20/M output
1M tokens ⓘ

GPT-6.1 Sol Pro is the same underlying model as [GPT-6.1 Sol](https://openrouter.ai/openai/gpt-6.1-sol), served with `reasoning.mode` set to `pro` for higher-quality responses on complex tasks. **Cost note:** pro mode spends far more...

by |Th9 2026 |1.1M context |$2.00/M input |$10.00/M output
1.1M tokens ⓘ

GPT-6.1 Sol is an upgrade to GPT-6 Sol from OpenAI, positioned below the flagship GPT-6 Astra in the GPT-6 series. It is suited for agentic coding, computer use, document-heavy professional...

by |Th9 2026 |1.1M context |$2.00/M input |$10.00/M output
1.1M tokens ⓘ

Claude Sonnet 5.5 is Anthropic's Sonnet-class model for well-scoped everyday work, succeeding Claude Sonnet 5 as a direct upgrade. It is especially strong at building features, fixing bugs, and producing...

by |Th9 2026 |1M context |$2.00/M input |$10.00/M output
1M tokens ⓘ

Claude Sonnet 5.5 is Anthropic's Sonnet-class model for well-scoped everyday work, succeeding Claude Sonnet 5 as a direct upgrade. It is especially strong at building features, fixing bugs, and producing...

by |Th9 2026 |1M context |$1.00/M input |$5.00/M output
1M tokens ⓘ

Jev Router picks the best model and reasoning effort for each request, balancing quality, speed, and cost. It runs on [Jev](https://openrouter.ai/~typesafe/jev-latest), TypeSafe's first System One model, and adapts as your...

by |Th9 2026 |1M context |Miễn phí input |Miễn phí output
1M tokens ⓘ

Perceptron Mk1.5 is Perceptron's embodied reasoning model for physical agents. It accepts text, image, video, and audio input, and answers with text plus optional structured annotations: points, boxes, polygons, tracks,...

by |Th9 2026 |37K context |$0.1500/M input |$1.50/M output

Ember-1 is a specialized reasoning model from Fireworks Research, built on [Kimi K3](https://openrouter.ai/moonshotai/kimi-k3). It is designed to make every token go further: it produces shorter reasoning traces, using roughly 40%...

by |Th9 2026 |1M context |$3.00/M input |$15.00/M output
1M tokens ⓘ

GLM-5.3-Prime is the high-speed variant of Z.ai's GLM-5.3, inheriting its full capabilities while delivering 1.5–2× the output throughput through inference acceleration. It supports text input and output with a 1M-token...

by |Th9 2026 |1M context |$2.80/M input |$8.80/M output
1M tokens ⓘ

Qwen3.8 Max Prime is a higher-throughput variant of Qwen3.8 Max from Alibaba's Qwen team, served as a separate SKU at a higher price point. It accepts text, image, and video...

by |Th9 2026 |1M context |$4.00/M input |$12.00/M output
1M tokens ⓘ

Space Bunny Alpha is an anonymous large model with blazing-fast inference, strong coding capabilities and native multimodal input support. It delivers adjustable reasoning effort, and a 1M-token context window. Space...

by |Th9 2026 |1M context |Miễn phí input |Miễn phí output
1M tokens ⓘ

Aion 3.5 Mini is a multi-model roleplaying and storytelling system from AionLabs, built on the GLM family of models. It is the smaller, lower-cost sibling of Aion 3.5 and uses...

by |Th9 2026 |262K context |$0.7000/M input |$1.40/M output
262K tokens ⓘ

Aion 3.5 is a multi-model roleplaying and storytelling system from AionLabs, built on the GLM family of models. It uses a collaborative generation process in which multiple specialized models each...

by |Th9 2026 |262K context |$3.00/M input |$6.00/M output
262K tokens ⓘ

Solar Mini 4 is Upstage's compact, cost-efficient language model, a 35B-parameter mixture-of-experts with 3B active parameters and a 524K context window. It is built for agentic use cases where response...

by |Th9 2026 |524K context |$0.0500/M input |$0.2000/M output
524K tokens ⓘ

Command A+ is Cohere's flagship model for enterprise agentic workflows. It accepts text and image inputs with a 192K context window, supports native tool calling with strict tool schemas, structured...

by |Th9 2026 |192K context |$0.3000/M input |$1.50/M output
192K tokens ⓘ

GPT-6 Luna Pro is the same underlying model as [GPT-6 Luna](https://openrouter.ai/openai/gpt-6-luna), served with `reasoning.mode` set to `pro` for higher-quality responses on complex tasks. **Cost note:** pro mode spends far more...

by |Th9 2026 |1.1M context |$0.1000/M input |$0.5000/M output
1.1M tokens ⓘ

GPT-6 Luna Pro is the same underlying model as [GPT-6 Luna](https://openrouter.ai/openai/gpt-6-luna), served with `reasoning.mode` set to `pro` for higher-quality responses on complex tasks. **Cost note:** pro mode spends far more...

by |Th9 2026 |1.1M context |$0.0500/M input |$0.2500/M output
1.1M tokens ⓘ

GPT-6 Luna is the fast, cost-efficient model in OpenAI's GPT-6 series, positioned below GPT-6 Sol. It is suited for high-volume and latency-sensitive workloads such as chat, classification, and lightweight agentic...

by |Th9 2026 |1.1M context |$0.1000/M input |$0.5000/M output
1.1M tokens ⓘ

GPT-6 Luna is the fast, cost-efficient model in OpenAI's GPT-6 series, positioned below GPT-6 Sol. It is suited for high-volume and latency-sensitive workloads such as chat, classification, and lightweight agentic...

by |Th9 2026 |1.1M context |$0.0500/M input |$0.2500/M output
1.1M tokens ⓘ

GPT-6 Sol Pro is the same underlying model as [GPT-6 Sol](https://openrouter.ai/openai/gpt-6-sol), served with `reasoning.mode` set to `pro` for higher-quality responses on complex tasks. **Cost note:** pro mode spends far more...

by |Th9 2026 |1.1M context |$1.00/M input |$5.00/M output
1.1M tokens ⓘ

GPT-6 Sol Pro is the same underlying model as [GPT-6 Sol](https://openrouter.ai/openai/gpt-6-sol), served with `reasoning.mode` set to `pro` for higher-quality responses on complex tasks. **Cost note:** pro mode spends far more...

by |Th9 2026 |1.1M context |$2.00/M input |$10.00/M output
1.1M tokens ⓘ

GPT-6 Sol is the cost-efficient high-end model in OpenAI's GPT-6 series, positioned below the flagship GPT-6 Astra and above the fast GPT-6 Luna tier. It is suited for demanding professional...

by |Th9 2026 |1.1M context |$2.00/M input |$10.00/M output
1.1M tokens ⓘ

GPT-6 Sol is the cost-efficient high-end model in OpenAI's GPT-6 series, positioned below the flagship GPT-6 Astra and above the fast GPT-6 Luna tier. It is suited for demanding professional...

by |Th9 2026 |1.1M context |$1.00/M input |$5.00/M output
1.1M tokens ⓘ

Claude Opus 5.5 is Anthropic's flagship model for demanding reasoning, coding, and long-horizon agentic work, succeeding Claude Opus 5. It is particularly strong at multi-step changes in large codebases, code...

by |Th9 2026 |1M context |$4.00/M input |$20.00/M output
1M tokens ⓘ

Claude Opus 5.5 is Anthropic's flagship model for demanding reasoning, coding, and long-horizon agentic work, succeeding Claude Opus 5. It is particularly strong at multi-step changes in large codebases, code...

by |Th9 2026 |1M context |$2.00/M input |$10.00/M output
1M tokens ⓘ

MiMo-V2.6-Pro-UltraSpeed is the fast speed edition of Xiaomi's flagship foundation model, MiMo-V2.6-Pro. Built from the same 1T MiMo-V2.6-Pro checkpoint, it matches the original model in quality while delivering roughly 10x...

by |Th9 2026 |1M context |$4.35/M input |$8.70/M output
1M tokens ⓘ

MiMo-V2.6-Flash is an open-source foundation model developed by Xiaomi. Built on a Mixture-of-Experts architecture with 309B total parameters and 15B activated per token, it employs a hybrid attention mechanism for...

by |Th9 2026 |1.1M context |$0.1400/M input |$0.2800/M output
1.1M tokens ⓘ

MiMo-V2.6-Pro is the flagship foundation model developed by Xiaomi. Built at a scale of over 1T parameters, it is designed to push the ceiling of capability for the most demanding...

by |Th9 2026 |1.1M context |$0.4350/M input |$0.8700/M output
1.1M tokens ⓘ

Grok 4.7 is SpaceXAI's flagship model for coding, agentic tasks, and knowledge work, succeeding Grok 4.6. It is particularly strong at long-running software engineering tasks, verifying its own work, and...

by |Th9 2026 |500K context |$2.00/M input |$6.00/M output
500K tokens ⓘ

Qwen3.8 Omni Flash is an omni-modal reasoning model from Alibaba, the first Qwen model built around agentic capabilities with native audio-video understanding. It is suited for audio-video analysis and summarization,...

by |Th9 2026 |1M context |$0.1500/M input |$0.4700/M output
1M tokens ⓘ

Bonsai 2 27B is a 27B-parameter reasoning model from PrismML derived from Qwen3.8-27B. It supports coding, mathematics, tool calling, and image understanding with a 262K-token context window. Ternary compression shrinks...

by |Th9 2026 |262K context |$0.0750/M input |$0.5000/M output
262K tokens ⓘ

GLM-5.3-FlashX is the high-speed variant of Z.ai's GLM-5.3-Flash, a native multimodal model delivering inference speeds of up to 200 tokens/s. Built on the same hybrid sparse and linear attention architecture...

by |Th9 2026 |1M context |$0.3700/M input |$1.25/M output
1M tokens ⓘ

Pareto is a multimodal composite model built for research, coding, and agentic workflows, while delivering frontier-level performance across a broad range of general-purpose tasks.

by |Th9 2026 |262K context |$2.50/M input |$7.50/M output
262K tokens ⓘ

This model always redirects to the latest model in the DeepSeek Pro family.

by |Th9 2026 |1M context |$0.1320/M input |$0.3960/M output
1M tokens ⓘ

Schematron V2 Turbo is a 3B-parameter HTML-to-JSON extraction model from Inference.net. It prioritizes throughput for high-volume extraction workloads. Extraction instructions must be supplied through a JSON schema in response_format rather...

by |Th9 2026 |128K context |$0.0300/M input |$0.1500/M output
128K tokens ⓘ

Schematron V2 Small is a 3B-parameter HTML-to-JSON extraction model from Inference.net. It prioritizes extraction quality for complex schemas and long pages. Extraction instructions must be supplied through a JSON schema...

by |Th9 2026 |128K context |$0.0500/M input |$0.2300/M output
128K tokens ⓘ

This model always redirects to the latest model in the GPT Astra family.

by |Th9 2026 |1.1M context |$10.00/M input |$50.00/M output
1.1M tokens ⓘ

This model always redirects to the latest model in the GPT Sol family.

by |Th9 2026 |1.1M context |$2.00/M input |$10.00/M output
1.1M tokens ⓘ

This model always redirects to the latest model in the GPT Terra family.

by |Th9 2026 |1.1M context |$2.00/M input |$12.00/M output
1.1M tokens ⓘ

This model always redirects to the latest model in the GPT Luna family.

by |Th9 2026 |1.1M context |$0.1000/M input |$0.5000/M output
1.1M tokens ⓘ

Fugu Ultra v2 is the higher-performance model in Sakana AI's Fugu family. Rather than a single monolithic model, Fugu is a learned multi-agent orchestration system: a language model trained to...

by |Th9 2026 |1M context |$5.00/M input |$30.00/M output
1M tokens ⓘ

Fugu Max is the cost-performance model in Sakana AI's Fugu family. Rather than a single monolithic model, Fugu is a learned multi-agent orchestration system: a language model trained to route...

by |Th9 2026 |1M context |$2.00/M input |$6.00/M output
1M tokens ⓘ

Ling 3.0 Flash VL builds on Ling 3.0 Flash (124B total / 5.5B active MoE from InclusionAI), further strengthening its language capabilities while adding native visual perception and advanced visual...

by |Th9 2026 |262K context |$0.0210/M input |$0.0616/M output
262K tokens ⓘ

DeepSeek V4.1 Flash is a sparse mixture-of-experts model from DeepSeek, and the first built on the company's Causal Encoder-Decoder (CED) architecture. It activates 8B parameters on input and 16B on...

by |Th9 2026 |1M context |$0.3000/M input |$1.20/M output
1M tokens ⓘ

DeepSeek V4.1 Flash is a sparse mixture-of-experts model from DeepSeek, and the first built on the company's Causal Encoder-Decoder (CED) architecture. It activates 8B parameters on input and 16B on...

by |Th9 2026 |1M context |$0.1120/M input |$0.3360/M output
1M tokens ⓘ

Mercury 2.5 is the fastest reasoning LLM, and the latest diffusion LLM (dLLM) from Inception. Instead of generating tokens sequentially, Mercury 2.5 produces and refines multiple tokens in parallel, achieving...

by |Th9 2026 |260K context |$0.0400/M input |$0.1500/M output
260K tokens ⓘ

Nex-N2.5 is an agentic model built to turn goals into working, verified outcomes. Its core strength is agentic coding within a visual feedback loop: it can explore codebases, implement multi-file...

by |Th9 2026 |262K context |$0.0250/M input |$0.1000/M output
262K tokens ⓘ

Nex-N2.5 is an agentic model built to turn goals into working, verified outcomes. Its core strength is agentic coding within a visual feedback loop: it can explore codebases, implement multi-file...

by |Th9 2026 |262K context |$0.0750/M input |$0.2500/M output
262K tokens ⓘ