AI 모델

464 모델 Free & Paid 업데이트: 12 hours trước

Qwen3-Coder-480B-A35B-Instruct is a Mixture-of-Experts (MoE) code generation model developed by the Qwen team. It is optimized for agentic coding tasks such as function calling, tool use, and long-context reasoning over...

by |Jul 2025 |262K context |$0.3000/M input |$1.00/M output
262K tokens ⓘ

UI-TARS-1.5 is a multimodal vision-language agent optimized for GUI-based environments, including desktop interfaces, web browsers, mobile systems, and games. Built by ByteDance, it builds upon the UI-TARS framework with reinforcement...

by |Jul 2025 |128K context |$0.1000/M input |$0.2000/M output
128K tokens ⓘ

쌍둥이자리 2.5 Flash-Lite is a lightweight reasoning model in the Gemini 2.5 가족, optimized for ultra-low latency and cost efficiency. It offers improved throughput, faster token generation, and better performance...

by |Jul 2025 |1M context |$0.0500/M input |$0.2000/M output
1M tokens ⓘ

쌍둥이자리 2.5 Flash-Lite is a lightweight reasoning model in the Gemini 2.5 가족, optimized for ultra-low latency and cost efficiency. It offers improved throughput, faster token generation, and better performance...

by |Jul 2025 |1M context |$0.1000/M input |$0.4000/M output
1M tokens ⓘ

Qwen3-235B-A22B-Instruct-2507 is a multilingual, instruction-tuned mixture-of-experts language model based on the Qwen3-235B architecture, with 22B active parameters per forward pass. It is optimized for general-purpose text generation, including instruction following,...

by |Jul 2025 |262K context |$0.0875/M input |$0.3500/M output
262K tokens ⓘ

Kimi K2 Instruct is a large-scale Mixture-of-Experts (MoE) language model developed by Moonshot AI, featuring 1 trillion total parameters with 32 billion active per forward pass. It is optimized for...

by |Jul 2025 |131K context |$0.5700/M input |$2.30/M output
131K tokens ⓘ

Venice Uncensored Dolphin Mistral 24B Venice Edition is a fine-tuned variant of Mistral-Small-24B-Instruct-2501, developed by dphn.ai in collaboration with Venice.ai. This model is designed as an “uncensored” instruct-tuned LLM, preserving...

by |Jul 2025 |128K context |$0.2000/M input |$0.9000/M output
128K tokens ⓘ

Hunyuan-A13B is a 13B active parameter Mixture-of-Experts (MoE) language model developed by Tencent, with a total parameter count of 80B and support for reasoning via Chain-of-Thought. It offers competitive benchmark...

by |Jul 2025 |131K context |$0.1400/M input |$0.5700/M output
131K tokens ⓘ

Morph's high-accuracy apply model for complex code edits. ~4,500 tokens/sec with 98% accuracy for precise code transformations. The model requires the prompt to be in the following format: {instruction} {initial_code}...

by |Jul 2025 |262K context |$0.9000/M input |$1.90/M output
262K tokens ⓘ

Morph's fastest apply model for code edits. ~10,500 tokens/sec with 96% accuracy for rapid code transformations. The model requires the prompt to be in the following format: {instruction} {initial_code} {edit_snippet}...

by |Jul 2025 |82K context |$0.8000/M input |$1.20/M output
82K tokens ⓘ

ERNIE-4.5-VL-424B-A47B is a multimodal Mixture-of-Experts (MoE) model from Baidu’s ERNIE 4.5 series, featuring 424B total parameters with 47B active per token. It is trained jointly on text and image data...

by |6 월 2025 |123K context |$0.4200/M input |$1.25/M output
123K tokens ⓘ

Mistral-Small-3.2-24B-Instruct-2506 is an updated 24B parameter model from Mistral optimized for instruction following, repetition reduction, and improved function calling. Compared to the 3.1 release, version 3.2 significantly improves accuracy on...

by |6 월 2025 |256K context |$0.0938/M input |$0.2500/M output
256K tokens ⓘ

MiniMax-M1 is a large-scale, open-weight reasoning model designed for extended context and high-efficiency inference. It leverages a hybrid Mixture-of-Experts (MoE) architecture paired with a custom "lightning attention" mechanism, allowing it...

by |6 월 2025 |1M context |$0.5500/M input |$2.20/M output
1M tokens ⓘ

쌍둥이자리 2.5 Flash is Google's state-of-the-art workhorse model, specifically designed for advanced reasoning, coding, mathematics, and scientific tasks. It includes built-in "thinking" capabilities, enabling it to provide responses with greater...

by |6 월 2025 |1M context |$0.1500/M input |$1.25/M output
1M tokens ⓘ

쌍둥이자리 2.5 Flash is Google's state-of-the-art workhorse model, specifically designed for advanced reasoning, coding, mathematics, and scientific tasks. It includes built-in "thinking" capabilities, enabling it to provide responses with greater...

by |6 월 2025 |1M context |$0.3000/M input |$2.50/M output
1M tokens ⓘ

쌍둥이자리 2.5 Pro is Google’s state-of-the-art AI model designed for advanced reasoning, coding, mathematics, and scientific tasks. It employs “thinking” capabilities, enabling it to reason through responses with enhanced accuracy...

by |6 월 2025 |1M context |$0.6250/M input |$5.00/M output
1M tokens ⓘ

쌍둥이자리 2.5 Pro is Google’s state-of-the-art AI model designed for advanced reasoning, coding, mathematics, and scientific tasks. It employs “thinking” capabilities, enabling it to reason through responses with enhanced accuracy...

by |6 월 2025 |1M context |$1.25/M input |$10.00/M output
1M tokens ⓘ

The o-series of models are trained with reinforcement learning to think before they answer and perform complex reasoning. The o3-pro model uses more compute to think harder and provide consistently...

by |6 월 2025 |200K context |$20.00/M input |$80.00/M output
200K tokens ⓘ

쌍둥이자리 2.5 Pro is Google’s state-of-the-art AI model designed for advanced reasoning, coding, mathematics, and scientific tasks. It employs “thinking” capabilities, enabling it to reason through responses with enhanced accuracy...

by |6 월 2025 |1M context |$1.25/M input |$10.00/M output
1M tokens ⓘ

May 28th update to the [original DeepSeek R1](/deepseek/deepseek-r1) Performance on par with [OpenAI o1](/openai/o1), but open-sourced and with fully open reasoning tokens. It's 671B parameters in size, with 37B active...

by |5월 2025 |164K context |$0.5000/M input |$2.15/M output
164K tokens ⓘ

Claude Sonnet 4 significantly enhances the capabilities of its predecessor, Sonnet 3.7, excelling in both coding and reasoning tasks with improved precision and controllability. Achieving state-of-the-art performance on SWE-bench (72.7%),...

by |5월 2025 |200K context |$3.00/M input |$15.00/M output
200K tokens ⓘ

Mistral Medium 3 is a high-performance enterprise-grade language model designed to deliver frontier-level capabilities at significantly reduced operational cost. It balances state-of-the-art reasoning and multimodal performance with 8× lower cost...

by |5월 2025 |131K context |$0.4000/M input |$2.00/M output
131K tokens ⓘ

Llama Guard 4 is a Llama 4 Scout-derived multimodal pretrained model, fine-tuned for content safety classification. Similar to previous versions, it can be used to classify content in both LLM...

by |4월 2025 |164K context |$0.1800/M input |$0.1800/M output
164K tokens ⓘ

Qwen3, the latest generation in the Qwen large language model series, features both dense and mixture-of-experts (MoE) architectures to excel in reasoning, multilingual support, and advanced agent tasks. Its unique...

by |4월 2025 |131K context |$0.1200/M input |$0.5000/M output
131K tokens ⓘ

Qwen3-8B is a dense 8.2B parameter causal language model from the Qwen3 series, designed for both reasoning-heavy tasks and efficient dialogue. It supports seamless switching between "thinking" mode for math,...

by |4월 2025 |131K context |$0.1170/M input |$0.4550/M output
131K tokens ⓘ

Qwen3-14B is a dense 14.8B parameter causal language model from the Qwen3 series, designed for both complex reasoning and efficient dialogue. It supports seamless switching between a "thinking" mode for...

by |4월 2025 |131K context |$0.1200/M input |$0.2400/M output
131K tokens ⓘ

Qwen3-32B is a dense 32.8B parameter causal language model from the Qwen3 series, optimized for both complex reasoning and efficient dialogue. It supports seamless switching between a "thinking" mode for...

by |4월 2025 |131K context |$0.0800/M input |$0.2800/M output
131K tokens ⓘ

Qwen3-235B-A22B is a 235B parameter mixture-of-experts (MoE) model developed by Qwen, activating 22B parameters per forward pass. It supports seamless switching between a "thinking" mode for complex reasoning, math, and...

by |4월 2025 |131K context |$0.4550/M input |$1.82/M output
131K tokens ⓘ

OpenAI o4-mini-high is the same model as [o4-mini](/openai/o4-mini) with reasoning_effort set to high. OpenAI o4-mini is a compact reasoning model in the o-series, optimized for fast, cost-efficient performance while retaining...

by |4월 2025 |200K context |$1.10/M input |$4.40/M output
200K tokens ⓘ

o3 is a well-rounded and powerful model across domains. It sets a new standard for math, science, coding, and visual reasoning tasks. It also excels at technical writing and instruction-following....

by |4월 2025 |200K context |$2.00/M input |$8.00/M output
200K tokens ⓘ

o3 is a well-rounded and powerful model across domains. It sets a new standard for math, science, coding, and visual reasoning tasks. It also excels at technical writing and instruction-following....

by |4월 2025 |200K context |$1.00/M input |$4.00/M output
200K tokens ⓘ

OpenAI o4-mini is a compact reasoning model in the o-series, optimized for fast, cost-efficient performance while retaining strong multimodal and agentic capabilities. It supports tool use and demonstrates competitive reasoning...

by |4월 2025 |200K context |$1.10/M input |$4.40/M output
200K tokens ⓘ

OpenAI o4-mini is a compact reasoning model in the o-series, optimized for fast, cost-efficient performance while retaining strong multimodal and agentic capabilities. It supports tool use and demonstrates competitive reasoning...

by |4월 2025 |200K context |$0.5500/M input |$2.20/M output
200K tokens ⓘ

GPT-4.1 is a flagship large language model optimized for advanced instruction following, real-world software engineering, and long-context reasoning. It supports a 1 million token context window and outperforms GPT-4o and...

by |4월 2025 |1M context |$2.00/M input |$8.00/M output
1M tokens ⓘ

GPT-4.1 is a flagship large language model optimized for advanced instruction following, real-world software engineering, and long-context reasoning. It supports a 1 million token context window and outperforms GPT-4o and...

by |4월 2025 |1M context |$1.00/M input |$4.00/M output
1M tokens ⓘ

GPT-4.1 Mini is a mid-sized model delivering performance competitive with GPT-4o at substantially lower latency and cost. It retains a 1 million token context window and scores 45.1% on hard...

by |4월 2025 |1M context |$0.2000/M input |$0.8000/M output
1M tokens ⓘ

GPT-4.1 Mini is a mid-sized model delivering performance competitive with GPT-4o at substantially lower latency and cost. It retains a 1 million token context window and scores 45.1% on hard...

by |4월 2025 |1M context |$0.4000/M input |$1.60/M output
1M tokens ⓘ

For tasks that demand low latency, GPT‑4.1 nano is the fastest and cheapest model in the GPT-4.1 series. It delivers exceptional performance at a small size with its 1 million...

by |4월 2025 |1M context |$0.1000/M input |$0.4000/M output
1M tokens ⓘ

For tasks that demand low latency, GPT‑4.1 nano is the fastest and cheapest model in the GPT-4.1 series. It delivers exceptional performance at a small size with its 1 million...

by |4월 2025 |1M context |$0.0500/M input |$0.2000/M output
1M tokens ⓘ

Llama 4 Maverick 17B Instruct (128E) is a high-capacity multimodal language model from Meta, built on a mixture-of-experts (MoE) architecture with 128 experts and 17 billion active parameters per forward...

by |4월 2025 |1M context |$0.1875/M input |$0.6525/M output
1M tokens ⓘ

Llama 4 Scout 17B Instruct (16E) is a mixture-of-experts (MoE) language model developed by Meta, activating 17 billion parameters out of a total of 109B. It supports native multimodal input...

by |4월 2025 |1.3M context |$0.1000/M input |$0.3000/M output
1.3M tokens ⓘ

DeepSeek V3, a 685B-parameter, mixture-of-experts model, is the latest iteration of the flagship chat model family from the DeepSeek team. It succeeds the [DeepSeek V3](/deepseek/deepseek-chat-v3) model and performs really well...

by |3월 2025 |164K context |$0.2900/M input |$1.14/M output
164K tokens ⓘ

The o1 series of models are trained with reinforcement learning to think before they answer and perform complex reasoning. The o1-pro model uses more compute to think harder and provide...

by |3월 2025 |200K context |$150.00/M input |$600.00/M output
200K tokens ⓘ

Mistral Small 3.1 24B Instruct is an upgraded variant of Mistral Small 3 (2501), featuring 24 billion parameters with advanced multimodal capabilities. It provides state-of-the-art performance in text-based reasoning and...

by |3월 2025 |128K context |$0.3510/M input |$0.5550/M output
128K tokens ⓘ

Gemma 3 introduces multimodality, supporting vision-language input and text outputs. It handles context windows up to 128k tokens, understands over 140 languages, and offers improved math, reasoning, and chat capabilities,...

by |3월 2025 |131K context |$0.0500/M input |$0.1000/M output
131K tokens ⓘ

Gemma 3 introduces multimodality, supporting vision-language input and text outputs. It handles context windows up to 128k tokens, understands over 140 languages, and offers improved math, reasoning, and chat capabilities,...

by |3월 2025 |131K context |$0.0500/M input |$0.1500/M output
131K tokens ⓘ

Command A is an open-weights 111B parameter model with a 256k context window focused on delivering great performance across agentic, multilingual, and coding use cases. Compared to other leading proprietary...

by |3월 2025 |256K context |$2.50/M input |$10.00/M output
256K tokens ⓘ

Reka Flash 3 is a general-purpose, instruction-tuned large language model with 21 billion parameters, developed by Reka. It excels at general chat, coding tasks, instruction-following, and function calling. Featuring a...

by |3월 2025 |66K context |$0.1000/M input |$0.2000/M output
66K tokens ⓘ

Gemma 3 introduces multimodality, supporting vision-language input and text outputs. It handles context windows up to 128k tokens, understands over 140 languages, and offers improved math, reasoning, and chat capabilities,...

by |3월 2025 |131K context |$0.0800/M input |$0.4500/M output
131K tokens ⓘ

Skyfall 36B v2 is an enhanced iteration of Mistral Small 2501, specifically fine-tuned for improved creativity, nuanced writing, role-playing, and coherent storytelling.

by |3월 2025 |33K context |$0.5500/M input |$0.8000/M output