AI Models

458 Modelle Free & Paid Cập nhật: 3 hours trước

Step 5 Preview is StepFun's flagship model for agentic work, built on a sparse Mixture-of-Experts architecture (27B active / 600B total parameters). It performs strongly in software engineering and professional...

by |Oct 2026 |1M context |$1.00/M input |$2.70/M output
1M tokens ⓘ

Claude Haiku 5.5 is Anthropic's small, fast model for high-volume, cost-sensitive work such as summarization, subagents, and browser use. It succeeds Claude Haiku 4.5 with stronger coding, computer use, and...

by |Oct 2026 |1M context |$0.1000/M input |$0.5000/M output
1M tokens ⓘ

Claude Haiku 5.5 is Anthropic's small, fast model for high-volume, cost-sensitive work such as summarization, subagents, and browser use. It succeeds Claude Haiku 4.5 with stronger coding, computer use, and...

by |Oct 2026 |1M context |$0.0500/M input |$0.2500/M output
1M tokens ⓘ

Nano Banana 2.1 (Gemini Nano Banana 2.1) is Google's image generation and editing model on the Flash tier, succeeding Nano Banana 2 and Nano Banana Pro. It improves product recontextualization,...

by |Oct 2026 |66K context |$1.50/M input |$7.50/M output
66K tokens ⓘ

Mistral Large 4 is a frontier multimodal (text and image input) model from Mistral AI built for reasoning, coding, and agentic workloads. It offers a 1M-token context window with up...

by |Oct 2026 |1M context |$0.6800/M input |$2.09/M output
1M tokens ⓘ

Ling 3.1 Flash is a hybrid reasoning mixture-of-experts model from inclusionAI, with 25B active parameters out of 560B total.

by |Oct 2026 |262K context |Miễn phí input |Miễn phí output
262K tokens ⓘ

Apodex 1.1 Mini is a reasoning-first model from Apodex, built for complex, long-horizon research and forecasting tasks. It works directly with files, data, code, and tools to produce verifiable results,...

by |Oct 2026 |262K context |Miễn phí input |Miễn phí output
262K tokens ⓘ

Pareto is a multimodal composite model built for research, coding, and agentic workflows, while delivering frontier-level performance across a broad range of general-purpose tasks. This is a preview of the...

by |Oct 2026 |1M context |$0.8000/M input |$3.20/M output
1M tokens ⓘ

GPT-6.1 Sol Pro is the same underlying model as [GPT-6.1 Sol](https://openrouter.ai/openai/gpt-6.1-sol), served with `reasoning.mode` set to `pro` for higher-quality responses on complex tasks. **Cost note:** pro mode spends far more...

by |Sep 2026 |1.1M context |$2.00/M input |$10.00/M output
1.1M tokens ⓘ

GPT-6.1 Sol Pro is the same underlying model as [GPT-6.1 Sol](https://openrouter.ai/openai/gpt-6.1-sol), served with `reasoning.mode` set to `pro` for higher-quality responses on complex tasks. **Cost note:** pro mode spends far more...

by |Sep 2026 |1.1M context |$1.00/M input |$5.00/M output
1.1M tokens ⓘ

GPT-6.1 Sol is an upgrade to GPT-6 Sol from OpenAI, positioned below the flagship GPT-6 Astra in the GPT-6 series. It is suited for agentic coding, computer use, document-heavy professional...

by |Sep 2026 |1.1M context |$2.00/M input |$10.00/M output
1.1M tokens ⓘ

GPT-6.1 Sol is an upgrade to GPT-6 Sol from OpenAI, positioned below the flagship GPT-6 Astra in the GPT-6 series. It is suited for agentic coding, computer use, document-heavy professional...

by |Sep 2026 |1.1M context |$1.00/M input |$5.00/M output
1.1M tokens ⓘ

Claude Sonnet 5.5 is Anthropic's Sonnet-class model for well-scoped everyday work, succeeding Claude Sonnet 5 as a direct upgrade. It is especially strong at building features, fixing bugs, and producing...

by |Sep 2026 |1M context |$2.00/M input |$10.00/M output
1M tokens ⓘ

Claude Sonnet 5.5 is Anthropic's Sonnet-class model for well-scoped everyday work, succeeding Claude Sonnet 5 as a direct upgrade. It is especially strong at building features, fixing bugs, and producing...

by |Sep 2026 |1M context |$1.00/M input |$5.00/M output
1M tokens ⓘ

Jev Router picks the best model and reasoning effort for each request, balancing quality, speed, and cost. It runs on [Jev](https://openrouter.ai/~typesafe/jev-latest), TypeSafe's first System One model, and adapts as your...

by |Sep 2026 |1M context |Miễn phí input |Miễn phí output
1M tokens ⓘ

Perceptron Mk1.5 is Perceptron's embodied reasoning model for physical agents. It accepts text, Bild, video, and audio input, and answers with text plus optional structured annotations: points, boxes, polygons, tracks,...

by |Sep 2026 |37K context |$0.1500/M input |$1.50/M output

Ember-1 is a specialized reasoning model from Fireworks Research, built on [Kimi K3](https://openrouter.ai/moonshotai/kimi-k3). It is designed to make every token go further: it produces shorter reasoning traces, using roughly 40%...

by |Sep 2026 |1M context |$3.00/M input |$15.00/M output
1M tokens ⓘ

GLM-5.3-Prime is the high-speed variant of Z.ai's GLM-5.3, inheriting its full capabilities while delivering 1.5–2× the output throughput through inference acceleration. It supports text input and output with a 1M-token...

by |Sep 2026 |1M context |$2.80/M input |$8.80/M output
1M tokens ⓘ

Qwen3.8 Max Prime is a higher-throughput variant of Qwen3.8 Max from Alibaba's Qwen team, served as a separate SKU at a higher price point. It accepts text, Bild, and video...

by |Sep 2026 |1M context |$4.00/M input |$12.00/M output
1M tokens ⓘ

Aion 3.5 Mini is a multi-model roleplaying and storytelling system from AionLabs, built on the GLM family of models. It is the smaller, lower-cost sibling of Aion 3.5 and uses...

by |Sep 2026 |262K context |$0.7000/M input |$1.40/M output
262K tokens ⓘ

Aion 3.5 is a multi-model roleplaying and storytelling system from AionLabs, built on the GLM family of models. It uses a collaborative generation process in which multiple specialized models each...

by |Sep 2026 |262K context |$3.00/M input |$6.00/M output
262K tokens ⓘ

Solar Mini 4 is Upstage's compact, cost-efficient language model, a 35B-parameter mixture-of-experts with 3B active parameters and a 524K context window. It is built for agentic use cases where response...

by |Sep 2026 |524K context |$0.0500/M input |$0.2000/M output
524K tokens ⓘ

Command A+ is Cohere's flagship model for enterprise agentic workflows. It accepts text and image inputs with a 192K context window, supports native tool calling with strict tool schemas, structured...

by |Sep 2026 |192K context |$0.3000/M input |$1.50/M output
192K tokens ⓘ

GPT-6 Luna Pro is the same underlying model as [GPT-6 Luna](https://openrouter.ai/openai/gpt-6-luna), served with `reasoning.mode` set to `pro` for higher-quality responses on complex tasks. **Cost note:** pro mode spends far more...

by |Sep 2026 |1.1M context |$0.1000/M input |$0.5000/M output
1.1M tokens ⓘ

GPT-6 Luna Pro is the same underlying model as [GPT-6 Luna](https://openrouter.ai/openai/gpt-6-luna), served with `reasoning.mode` set to `pro` for higher-quality responses on complex tasks. **Cost note:** pro mode spends far more...

by |Sep 2026 |1.1M context |$0.0500/M input |$0.2500/M output
1.1M tokens ⓘ

GPT-6 Luna is the fast, cost-efficient model in OpenAI's GPT-6 series, positioned below GPT-6 Sol. It is suited for high-volume and latency-sensitive workloads such as chat, classification, and lightweight agentic...

by |Sep 2026 |1.1M context |$0.1000/M input |$0.5000/M output
1.1M tokens ⓘ

GPT-6 Luna is the fast, cost-efficient model in OpenAI's GPT-6 series, positioned below GPT-6 Sol. It is suited for high-volume and latency-sensitive workloads such as chat, classification, and lightweight agentic...

by |Sep 2026 |1.1M context |$0.0500/M input |$0.2500/M output
1.1M tokens ⓘ

GPT-6 Sol Pro is the same underlying model as [GPT-6 Sol](https://openrouter.ai/openai/gpt-6-sol), served with `reasoning.mode` set to `pro` for higher-quality responses on complex tasks. **Cost note:** pro mode spends far more...

by |Sep 2026 |1.1M context |$2.00/M input |$10.00/M output
1.1M tokens ⓘ

GPT-6 Sol Pro is the same underlying model as [GPT-6 Sol](https://openrouter.ai/openai/gpt-6-sol), served with `reasoning.mode` set to `pro` for higher-quality responses on complex tasks. **Cost note:** pro mode spends far more...

by |Sep 2026 |1.1M context |$1.00/M input |$5.00/M output
1.1M tokens ⓘ

GPT-6 Sol is the cost-efficient high-end model in OpenAI's GPT-6 series, positioned below the flagship GPT-6 Astra and above the fast GPT-6 Luna tier. It is suited for demanding professional...

by |Sep 2026 |1.1M context |$2.00/M input |$10.00/M output
1.1M tokens ⓘ

GPT-6 Sol is the cost-efficient high-end model in OpenAI's GPT-6 series, positioned below the flagship GPT-6 Astra and above the fast GPT-6 Luna tier. It is suited for demanding professional...

by |Sep 2026 |1.1M context |$1.00/M input |$5.00/M output
1.1M tokens ⓘ

Claude Opus 5.5 is Anthropic's flagship model for demanding reasoning, coding, and long-horizon agentic work, succeeding Claude Opus 5. It is particularly strong at multi-step changes in large codebases, code...

by |Sep 2026 |1M context |$4.00/M input |$20.00/M output
1M tokens ⓘ

Claude Opus 5.5 is Anthropic's flagship model for demanding reasoning, coding, and long-horizon agentic work, succeeding Claude Opus 5. It is particularly strong at multi-step changes in large codebases, code...

by |Sep 2026 |1M context |$2.00/M input |$10.00/M output
1M tokens ⓘ

Switchyard is an open-source model router that switches between multiple models to optimize the cost of requests. By default it will use OpenRouter market data to select the most popular...

by |Sep 2026 |1M context |Miễn phí input |Miễn phí output
1M tokens ⓘ

MiMo-V2.6-Pro-UltraSpeed is the fast speed edition of Xiaomi's flagship foundation model, MiMo-V2.6-Pro. Built from the same 1T MiMo-V2.6-Pro checkpoint, it matches the original model in quality while delivering roughly 10x...

by |Sep 2026 |1M context |$4.35/M input |$8.70/M output
1M tokens ⓘ

MiMo-V2.6-Flash is an open-source foundation model developed by Xiaomi. Built on a Mixture-of-Experts architecture with 309B total parameters and 15B activated per token, it employs a hybrid attention mechanism for...

by |Sep 2026 |1.1M context |$0.1400/M input |$0.2800/M output
1.1M tokens ⓘ

MiMo-V2.6-Pro is the flagship foundation model developed by Xiaomi. Built at a scale of over 1T parameters, it is designed to push the ceiling of capability for the most demanding...

by |Sep 2026 |1.1M context |$0.4350/M input |$0.8700/M output
1.1M tokens ⓘ

Grok 4.7 is SpaceXAI's flagship model for coding, agentic tasks, and knowledge work, succeeding Grok 4.6. It is particularly strong at long-running software engineering tasks, verifying its own work, and...

by |Sep 2026 |500K context |$2.00/M input |$6.00/M output
500K tokens ⓘ

Qwen3.8 Omni Flash is an omni-modal reasoning model from Alibaba, the first Qwen model built around agentic capabilities with native audio-video understanding. It is suited for audio-video analysis and summarization,...

by |Sep 2026 |1M context |$0.1500/M input |$0.4700/M output
1M tokens ⓘ

Bonsai 2 27B is a 27B-parameter reasoning model from PrismML derived from Qwen3.8-27B. It supports coding, mathematics, tool calling, and image understanding with a 262K-token context window. Ternary compression shrinks...

by |Sep 2026 |262K context |$0.0750/M input |$0.5000/M output
262K tokens ⓘ

GLM-5.3-FlashX is the high-speed variant of Z.ai's GLM-5.3-Flash, a native multimodal model delivering inference speeds of up to 200 tokens/s. Built on the same hybrid sparse and linear attention architecture...

by |Sep 2026 |1M context |$0.3700/M input |$1.25/M output
1M tokens ⓘ

Pareto is a multimodal composite model built for research, coding, and agentic workflows, while delivering frontier-level performance across a broad range of general-purpose tasks.

by |Sep 2026 |262K context |$2.50/M input |$7.50/M output
262K tokens ⓘ

This model always redirects to the latest model in the DeepSeek Pro family.

by |Sep 2026 |1M context |$0.1426/M input |$0.4277/M output
1M tokens ⓘ

Schematron V2 Turbo is a 3B-parameter HTML-to-JSON extraction model from Inference.net. It prioritizes throughput for high-volume extraction workloads. Extraction instructions must be supplied through a JSON schema in response_format rather...

by |Sep 2026 |128K context |$0.0300/M input |$0.1500/M output
128K tokens ⓘ

Schematron V2 Small is a 3B-parameter HTML-to-JSON extraction model from Inference.net. It prioritizes extraction quality for complex schemas and long pages. Extraction instructions must be supplied through a JSON schema...

by |Sep 2026 |128K context |$0.0500/M input |$0.2300/M output
128K tokens ⓘ

This model always redirects to the latest model in the GPT Astra family.

by |Sep 2026 |1.1M context |$10.00/M input |$50.00/M output
1.1M tokens ⓘ

This model always redirects to the latest model in the GPT Sol family.

by |Sep 2026 |1.1M context |$2.00/M input |$10.00/M output
1.1M tokens ⓘ

This model always redirects to the latest model in the GPT Terra family.

by |Sep 2026 |1.1M context |$2.00/M input |$12.00/M output
1.1M tokens ⓘ

This model always redirects to the latest model in the GPT Luna family.

by |Sep 2026 |1.1M context |$0.1000/M input |$0.5000/M output
1.1M tokens ⓘ