AI Models

412 models Free & Paid 更新: 9 hours trước

Qwen3-Max is an updated release built on the Qwen3 series, offering major improvements in reasoning, instruction following, multilingual support, and long-tail knowledge coverage compared to the January 2025 version. It...

经过 |九月 2025 |262K 上下文 |$0.7800/米输入 |$3.90/米输出
262K 代币

Qwen3 Coder Plus is Alibaba's proprietary version of the Open Source Qwen3 Coder 480B A35B. It is a powerful coding agent model specializing in autonomous programming via tool calling and...

经过 |九月 2025 |1M 上下文 |$0.6500/米输入 |$3.25/米输出
1M代币

GPT-5-Codex is a specialized version of GPT-5 optimized for software engineering and coding workflows. It is designed for both interactive development sessions and long, independent execution of complex engineering tasks....

经过 |九月 2025 |400K 上下文 |$0.6250/米输入 |$5.00/米输出
400K 代币

DeepSeek-V3.1 Terminus is an update to [DeepSeek V3.1](/deepseek/deepseek-chat-v3.1) that maintains the model's original capabilities while addressing issues reported by users, including language consistency and agent capabilities, further optimizing the model's...

经过 |九月 2025 |164K 上下文 |$0.2700/米输入 |$1.00/米输出
164K 代币

Qwen3 Coder Flash is Alibaba's fast and cost efficient version of their proprietary Qwen3 Coder Plus. It is a powerful coding agent model specializing in autonomous programming via tool calling...

经过 |九月 2025 |1M 上下文 |$0.1950/米输入 |$0.9750/米输出
1M代币

Qwen3-Next-80B-A3B-Thinking is a reasoning-first chat model in the Qwen3-Next line that outputs structured “thinking” traces by default. It’s designed for hard multi-step problems; math proofs, code synthesis/debugging, 逻辑, and agentic...

经过 |九月 2025 |262K 上下文 |$0.1500/米输入 |$1.20/米输出
262K 代币

Qwen3-Next-80B-A3B-Instruct is an instruction-tuned chat model in the Qwen3-Next series optimized for fast, stable responses without “thinking” traces. It targets complex tasks across reasoning, code generation, knowledge QA, and multilingual...

经过 |九月 2025 |262K 上下文 |$0.1000/米输入 |$1.10/米输出
262K 代币

Qwen Plus 0728, based on the Qwen3 foundation model, is a 1 million context hybrid reasoning model with a balanced performance, speed, and cost combination.

经过 |九月 2025 |1M 上下文 |$0.2600/米输入 |$0.7800/米输出
1M代币

Qwen Plus 0728, based on the Qwen3 foundation model, is a 1 million context hybrid reasoning model with a balanced performance, speed, and cost combination.

经过 |九月 2025 |1M 上下文 |$0.2600/米输入 |$0.7800/米输出
1M代币

NVIDIA-Nemotron-Nano-9B-v2 is a large language model (法学硕士) trained from scratch by NVIDIA, and designed as a unified model for both reasoning and non-reasoning tasks. It responds to user queries and...

经过 |九月 2025 |128K 上下文 |自由输入 |自由输出
128K 代币

Kimi K2 0905 is the September update of [Kimi K2 0711](moonshotai/kimi-k2). It is a large-scale Mixture-of-Experts (教育部) language model developed by Moonshot AI, featuring 1 trillion total parameters with 32...

经过 |九月 2025 |262K 上下文 |$0.6000/米输入 |$2.50/米输出
262K 代币

Qwen3-30B-A3B-Thinking-2507 is a 30B parameter Mixture-of-Experts reasoning model optimized for complex tasks requiring extended multi-step thinking. The model is designed specifically for “thinking mode,” where internal reasoning traces are separated...

经过 |八月 2025 |82K 上下文 |$0.2000/米输入 |$2.40/米输出
82K 代币

Hermes 4 70B is a hybrid reasoning model from Nous Research, built on Meta-Llama-3.1-70B. It introduces the same hybrid mode as the larger 405B release, allowing the model to either...

经过 |八月 2025 |131K 上下文 |$0.1300/米输入 |$0.4000/米输出
131K 代币

Hermes 4 是Nous Research发布的基于Meta-Llama-3.1-405B构建的大规模推理模型. It introduces a hybrid reasoning mode, where the model can choose to deliberate internally with...

经过 |八月 2025 |131K 上下文 |$1.00/米输入 |$3.00/米输出
131K 代币

DeepSeek-V3.1 is a large hybrid reasoning model (671B parameters, 37B 活跃) that supports both thinking and non-thinking modes via prompt templates. It extends the DeepSeek-V3 base with a two-phase long-context...

经过 |八月 2025 |164K 上下文 |$0.2500/米输入 |$0.9500/米输出
164K 代币

米斯特拉尔介质 3.1 is an updated version of Mistral Medium 3, 这是一种高性能企业级语言模型,旨在以显着降低运营成本的方式提供前沿级功能. It balances...

经过 |八月 2025 |131K 上下文 |$0.4000/米输入 |$2.00/米输出
131K 代币

GLM-4.5V is a vision-language foundation model for multimodal agent applications. Built on a Mixture-of-Experts (教育部) architecture with 106B parameters and 12B activated parameters, it achieves state-of-the-art results in video understanding,...

经过 |八月 2025 |66K 上下文 |$0.6000/米输入 |$1.80/米输出
66K 代币

Jamba Large 1.7 is the latest model in the Jamba open family, offering improvements in grounding, instruction-following, 和整体效率. Built on a hybrid SSM-Transformer architecture with a 256K context...

经过 |八月 2025 |256K 上下文 |$2.00/米输入 |$8.00/米输出
256K 代币

GPT-5 is OpenAI’s most advanced model, offering major improvements in reasoning, code quality, and user experience. It is optimized for complex tasks that require step-by-step reasoning, instruction following, and accuracy...

经过 |八月 2025 |400K 上下文 |$0.6250/米输入 |$5.00/米输出
400K 代币

GPT-5 is OpenAI’s most advanced model, offering major improvements in reasoning, code quality, and user experience. It is optimized for complex tasks that require step-by-step reasoning, instruction following, and accuracy...

经过 |八月 2025 |400K 上下文 |$1.25/米输入 |$10.00/米输出
400K 代币

GPT-5 Mini is a compact version of GPT-5, designed to handle lighter-weight reasoning tasks. It provides the same instruction-following and safety-tuning benefits as GPT-5, but with reduced latency and cost....

经过 |八月 2025 |400K 上下文 |$0.1250/米输入 |$1.00/米输出
400K 代币

GPT-5 Mini is a compact version of GPT-5, designed to handle lighter-weight reasoning tasks. It provides the same instruction-following and safety-tuning benefits as GPT-5, but with reduced latency and cost....

经过 |八月 2025 |400K 上下文 |$0.2500/米输入 |$2.00/米输出
400K 代币

GPT-5-Nano is the smallest and fastest variant in the GPT-5 system, optimized for developer tools, rapid interactions, and ultra-low latency environments. While limited in reasoning depth compared to its larger...

经过 |八月 2025 |400K 上下文 |$0.0250/米输入 |$0.2000/米输出
400K 代币

GPT-5-Nano is the smallest and fastest variant in the GPT-5 system, optimized for developer tools, rapid interactions, and ultra-low latency environments. While limited in reasoning depth compared to its larger...

经过 |八月 2025 |400K 上下文 |$0.0500/米输入 |$0.4000/米输出
400K 代币

gpt-oss-120b is an open-weight, 117B-parameter Mixture-of-Experts (教育部) language model from OpenAI designed for high-reasoning, agentic, and general-purpose production use cases. It activates 5.1B parameters per forward pass and is optimized...

经过 |八月 2025 |131K 上下文 |$0.0300/米输入 |$0.1700/米输出
131K 代币

gpt-oss-20b is an open-weight 21B parameter model released by OpenAI under the Apache 2.0 license. 它使用专家组合 (教育部) architecture with 3.6B active parameters per forward pass, optimized for...

经过 |八月 2025 |131K 上下文 |自由输入 |自由输出
131K 代币

gpt-oss-20b is an open-weight 21B parameter model released by OpenAI under the Apache 2.0 license. 它使用专家组合 (教育部) architecture with 3.6B active parameters per forward pass, optimized for...

经过 |八月 2025 |131K 上下文 |$0.0300/米输入 |$0.1300/米输出
131K 代币

近距离工作 4.1 is an updated version of Anthropic’s flagship model, offering improved performance in coding, reasoning, and agentic tasks. It achieves 74.5% on SWE-bench Verified and shows notable gains...

经过 |八月 2025 |200K 上下文 |$7.50/米输入 |$37.50/米输出
200K 代币

近距离工作 4.1 is an updated version of Anthropic’s flagship model, offering improved performance in coding, reasoning, and agentic tasks. It achieves 74.5% on SWE-bench Verified and shows notable gains...

经过 |八月 2025 |200K 上下文 |$15.00/米输入 |$75.00/米输出
200K 代币

Mistral's cutting-edge language model for coding released end of July 2025. Codestral specializes in low-latency, high-frequency tasks such as fill-in-the-middle (FIM), code correction and test generation. [Blog Post](https://mistral.ai/news/codestral-25-08)

经过 |八月 2025 |256K 上下文 |$0.3000/米输入 |$0.9000/米输出
256K 代币

Qwen3-Coder-30B-A3B-Instruct is a 30.5B parameter Mixture-of-Experts (教育部) model with 128 experts (8 active per forward pass), designed for advanced code generation, repository-scale understanding, and agentic tool use. Built on the...

经过 |Jul 2025 |262K 上下文 |$0.0700/米输入 |$0.2800/米输出
262K 代币

Qwen3-30B-A3B-Instruct-2507 is a 30.5B-parameter mixture-of-experts language model from Qwen, with 3.3B active parameters per inference. It operates in non-thinking mode and is designed for high-quality instruction following, multilingual understanding, and...

经过 |Jul 2025 |262K 上下文 |$0.0482/米输入 |$0.1931/米输出
262K 代币

GLM-4.5 is our latest flagship foundation model, purpose-built for agent-based applications. It leverages a Mixture-of-Experts (教育部) architecture and supports a context length of up to 128k tokens. GLM-4.5 delivers significantly...

经过 |Jul 2025 |131K 上下文 |$0.6000/米输入 |$2.20/米输出
131K 代币

GLM-4.5-Air is the lightweight variant of our latest flagship model family, also purpose-built for agent-centric applications. Like GLM-4.5, it adopts the Mixture-of-Experts (教育部) architecture but with a more compact parameter...

经过 |Jul 2025 |131K 上下文 |$0.1300/米输入 |$0.8500/米输出
131K 代币

Qwen3-235B-A22B-Thinking-2507 is a high-performance, open-weight Mixture-of-Experts (教育部) language model optimized for complex reasoning tasks. It activates 22B of its 235B parameters per forward pass and natively supports up to 262,144...

经过 |Jul 2025 |262K 上下文 |$0.2300/米输入 |$2.30/米输出
262K 代币

Qwen3-Coder-480B-A35B-Instruct is a Mixture-of-Experts (教育部) code generation model developed by the Qwen team. It is optimized for agentic coding tasks such as function calling, 工具使用, and long-context reasoning over...

经过 |Jul 2025 |262K 上下文 |$0.3000/米输入 |$1.00/米输出
262K 代币

UI-TARS-1.5 is a multimodal vision-language agent optimized for GUI-based environments, 包括桌面界面, web browsers, mobile systems, and games. 由字节跳动打造, it builds upon the UI-TARS framework with reinforcement...

经过 |Jul 2025 |128K 上下文 |$0.1000/米输入 |$0.2000/米输出
128K 代币

双子座 2.5 Flash-Lite is a lightweight reasoning model in the Gemini 2.5 family, optimized for ultra-low latency and cost efficiency. It offers improved throughput, 更快的代币生成, and better performance...

经过 |Jul 2025 |1M 上下文 |$0.0500/米输入 |$0.2000/米输出
1M代币

双子座 2.5 Flash-Lite is a lightweight reasoning model in the Gemini 2.5 family, optimized for ultra-low latency and cost efficiency. It offers improved throughput, 更快的代币生成, and better performance...

经过 |Jul 2025 |1M 上下文 |$0.1000/米输入 |$0.4000/米输出
1M代币

Qwen3-235B-A22B-Instruct-2507 is a multilingual, instruction-tuned mixture-of-experts language model based on the Qwen3-235B architecture, with 22B active parameters per forward pass. It is optimized for general-purpose text generation, including instruction following,...

经过 |Jul 2025 |262K 上下文 |$0.0900/米输入 |$0.5500/米输出
262K 代币

Kimi K2 Instruct is a large-scale Mixture-of-Experts (教育部) language model developed by Moonshot AI, featuring 1 trillion total parameters with 32 billion active per forward pass. It is optimized for...

经过 |Jul 2025 |131K 上下文 |$0.5700/米输入 |$2.30/米输出
131K 代币

Venice Uncensored Dolphin Mistral 24B Venice Edition is a fine-tuned variant of Mistral-Small-24B-Instruct-2501, developed by dphn.ai in collaboration with Venice.ai. This model is designed as an “uncensored” instruct-tuned LLM, preserving...

经过 |Jul 2025 |128K 上下文 |$0.2000/米输入 |$0.9000/米输出
128K 代币

Hunyuan-A13B is a 13B active parameter Mixture-of-Experts (教育部) language model developed by Tencent, with a total parameter count of 80B and support for reasoning via Chain-of-Thought. It offers competitive benchmark...

经过 |Jul 2025 |131K 上下文 |$0.1400/米输入 |$0.5700/米输出
131K 代币

Morph's high-accuracy apply model for complex code edits. ~4,500 tokens/sec with 98% accuracy for precise code transformations. The model requires the prompt to be in the following format: {instruction} {initial_code}...

经过 |Jul 2025 |262K 上下文 |$0.9000/米输入 |$1.90/米输出
262K 代币

Morph's fastest apply model for code edits. ~10,500 tokens/sec with 96% accuracy for rapid code transformations. The model requires the prompt to be in the following format: {instruction} {initial_code} {edit_snippet}...

经过 |Jul 2025 |82K 上下文 |$0.8000/米输入 |$1.20/米输出
82K 代币

ERNIE-4.5-VL-424B-A47B is a multimodal Mixture-of-Experts (教育部) model from Baidu’s ERNIE 4.5 系列, featuring 424B total parameters with 47B active per token. It is trained jointly on text and image data...

经过 |六月 2025 |123K 上下文 |$0.4200/米输入 |$1.25/米输出
123K 代币

Mistral-Small-3.2-24B-Instruct-2506 is an updated 24B parameter model from Mistral optimized for instruction following, repetition reduction, and improved function calling. 相比于 3.1 release, version 3.2 significantly improves accuracy on...

经过 |六月 2025 |256K 上下文 |$0.0938/米输入 |$0.2500/米输出
256K 代币

MiniMax-M1 is a large-scale, open-weight reasoning model designed for extended context and high-efficiency inference. It leverages a hybrid Mixture-of-Experts (教育部) architecture paired with a custom "lightning attention" mechanism, allowing it...

经过 |六月 2025 |1M 上下文 |$0.5500/米输入 |$2.20/米输出
1M代币

双子座 2.5 Flash is Google's state-of-the-art workhorse model, specifically designed for advanced reasoning, coding, mathematics, and scientific tasks. It includes built-in "thinking" capabilities, enabling it to provide responses with greater...

经过 |六月 2025 |1M 上下文 |$0.3000/米输入 |$2.50/米输出
1M代币

双子座 2.5 Flash is Google's state-of-the-art workhorse model, specifically designed for advanced reasoning, coding, mathematics, and scientific tasks. It includes built-in "thinking" capabilities, enabling it to provide responses with greater...

经过 |六月 2025 |1M 上下文 |$0.1500/米输入 |$1.25/米输出
1M代币