Skip to content
  • Models
  • Rankings
  • Ori
Sign Up
Sign Up
OpenRouterOpenRouter
© 2026 OpenRouter, Inc

Product

  • Chat
  • Rankings
  • Benchmarks
  • Apps
  • Discover
  • Models
  • Collections
  • Providers
  • Pricing
  • Business
  • Enterprise
  • Labs

Company

  • About
  • Blog
  • Careers
    Hiring
  • Privacy
  • Terms of Service
  • Trust Center
  • Support
  • Works With OR
  • Data
  • Brand

Developer

  • Documentation
  • API Reference
  • Developer Platform
  • Status

Connect

  • Discord
  • GitHub
  • LinkedIn
  • X
  • YouTube

Models

Models

CompareDiscover Models
Favicon for anthropic
Favicon for openai
  • Favicon for prism-ml
    PrismML: Ternary Bonsai 2 27BTernary Bonsai 2 27B
    8.23M tokens

    Bonsai 2 27B is a 27B-parameter reasoning model from PrismML derived from Qwen3.8-27B. It supports coding, mathematics, tool calling, and image understanding with a 262K-token context window. Ternary compression shrinks the language-model weights to roughly 8.5 GB while retaining 98.2% of the base model's average score across PrismML's 14 thinking-mode benchmarks, enabling efficient inference on consumer hardware. The model thinks by default and defaults to xhigh reasoning effort.

    by prism-mlSep 18, 2026262K context$0.075/M input tokens$0.50/M output tokens
  • Favicon for typesafe
    TypeSafe: Jev LatestJev Latest

    This model always redirects to the latest model in the Jev family.

    by typesafeSep 18, 202632K context$0.042/M input tokens$0/M output tokens
  • Favicon for typesafe
    TypeSafe: Jev 1.13Jev 1.13
    82.7B tokens

    Jev is a structured decision model from TypeSafe, and the first of its System One models. System One models make fast, structured decisions for software, returning a typed choice rather than free-form text. It is suited for routing, classification, and other decision points inside an application where a fast, predictable answer matters more than generated prose. Learn more in TypeSafe's docs: https://docs.typesafe.ai/concepts/system-one

    by typesafeSep 18, 202632K context$0.042/M input tokens$0/M output tokens
  • Favicon for unbiased
    ParetoPareto
    4.18B tokens

    Pareto is a multimodal composite model built for research, coding, and agentic workflows, while delivering frontier-level performance across a broad range of general-purpose tasks.

    by unbiasedSep 17, 2026262K context$2.50/M input tokens$7.50/M output tokens
  • Favicon for deepseek
    DeepSeek: DeepSeek Pro LatestDeepSeek Pro Latest
    56% off
    53.2B tokens

    This model always redirects to the latest model in the DeepSeek Pro family.

    by deepseekSep 14, 20261.05M context$0.5782/M input tokens$1.734/M output tokens
  • Favicon for deepseek
    DeepSeek: DeepSeek Flash LatestDeepSeek Flash Latest
    55% off
    255B tokens

    This model always redirects to the latest model in the DeepSeek Flash family.

    by deepseekSep 14, 20261.05M context$0.135/M input tokens$0.54/M output tokens
  • Favicon for inference-net
    Inference.net: Schematron V2 TurboSchematron V2 Turbo
    206M tokens

    Schematron V2 Turbo is a 3B-parameter HTML-to-JSON extraction model from Inference.net. It prioritizes throughput for high-volume extraction workloads. Extraction instructions must be supplied through a JSON schema in response_format rather than through system or user prompts.

    by inference-netSep 12, 2026128K context$0.03/M input tokens$0.15/M output tokens
  • Favicon for inference-net
    Inference.net: Schematron V2 SmallSchematron V2 Small
    76.4M tokens

    Schematron V2 Small is a 3B-parameter HTML-to-JSON extraction model from Inference.net. It prioritizes extraction quality for complex schemas and long pages. Extraction instructions must be supplied through a JSON schema in response_format rather than through system or user prompts.

    by inference-netSep 12, 2026128K context$0.05/M input tokens$0.23/M output tokens
  • Favicon for meta
    Meta: Muse Voice Transcribe 1.0Muse Voice Transcribe 1.0
    18+
    8.61M characters

    Muse Voice Transcribe 1.0 is a synchronous speech-to-text model from Meta. It is suited for push-to-talk, endpointing, and speaker-aware transcription, with keyword biasing for domain terms and language biasing through language-name hints. It accepts mono 16-bit PCM WAV audio at 16 kHz or 24 kHz for recordings up to 10 minutes. It does not provide word-level timestamps or confidence scores, and other audio formats must be converted to WAV before upload.

    by metaSep 11, 2026$0.00005/second
  • Favicon for openai
    OpenAI: GPT Astra LatestGPT Astra Latest
    22.7B tokens

    This model always redirects to the latest model in the GPT Astra family.

    by openaiSep 11, 20261.05M context$10/M input tokens$50/M output tokens
  • Favicon for openai
    OpenAI: GPT Sol LatestGPT Sol Latest
    50% off
    49.7B tokens

    This model always redirects to the latest model in the GPT Sol family.

    by openaiSep 11, 20261.05M context$2/M input tokens$10/M output tokens
  • Favicon for openai
    OpenAI: GPT Terra LatestGPT Terra Latest
    4.09B tokens

    This model always redirects to the latest model in the GPT Terra family.

    by openaiSep 11, 20261.05M context$2/M input tokens$12/M output tokens
  • Favicon for openai
    OpenAI: GPT Luna LatestGPT Luna Latest
    47.2B tokens

    This model always redirects to the latest model in the GPT Luna family.

    by openaiSep 11, 20261.05M context$0.20/M input tokens$1.20/M output tokens
  • Favicon for sakana
    Sakana: Fugu Ultra v2Fugu Ultra v2
    4.32B tokens

    Fugu Ultra v2 is the higher-performance model in Sakana AI's Fugu family. Rather than a single monolithic model, Fugu is a learned multi-agent orchestration system: a language model trained to route tasks across a fixed pool of open and specialized models and to recursively call instances of itself. Fugu Ultra v2 prioritizes answer quality on complex multi-step reasoning, autonomous research, and full-stack software development, and does not rely on individual proprietary frontier models in its pool. It supports configurable reasoning effort (high, xhigh, max), function calling, structured outputs, image and PDF input, and built-in web search. Prompts above 272K tokens are billed at a higher rate. Orchestration tokens consumed by the system are billed as standard input/output tokens.

    by sakanaSep 11, 20261M context$5/M input tokens$30/M output tokens
  • Favicon for sakana
    Sakana: Fugu MaxFugu Max
    6.51B tokens

    Fugu Max is the cost-performance model in Sakana AI's Fugu family. Rather than a single monolithic model, Fugu is a learned multi-agent orchestration system: a language model trained to route tasks across a fixed pool of open-weights and specialized models, including NVIDIA's Nemotron family, and to recursively call instances of itself. Fugu Max dynamically selects efficient combinations of expert agents to improve quality and cost together, and is priced flat regardless of context length. It supports configurable reasoning effort (high, xhigh, max), function calling, structured outputs, image and PDF input, and built-in web search and web fetch. Orchestration tokens consumed by the system are billed as standard input/output tokens.

    by sakanaSep 11, 20261M context$2/M input tokens$6/M output tokens
  • Favicon for black-forest-labs
    Black Forest Labs: FLUX Video EditFLUX Video Edit
    2 hours

    FLUX Video Edit [fast] takes a source video and an edit prompt and returns a precisely edited video. Add, remove, or replace objects and characters, rebuild the setting, edit on-screen text, change colors, materials and visual effects, change or translate dialogue with lip sync, or change the events of a clip. Output preserves the source duration, aspect ratio, and audio; inputs above 720p are downscaled to 720p. Source clips up to 15 seconds and 50 MiB.

    by black-forest-labsSep 10, 2026$0.03/second
  • Favicon for inclusionai
    inclusionAI: Ling 3.0 Flash VL (free)Ling 3.0 Flash VL (free)Free variant
    632B tokens
    Programming (#49)

    Ling 3.0 Flash VL builds on Ling 3.0 Flash (124B total / 5.5B active MoE from InclusionAI), further strengthening its language capabilities while adding native visual perception and advanced visual agent capabilities. Hybrid instant/reasoning model with tool calling.

    by inclusionaiSep 10, 2026262K context$0/M input tokens$0/M output tokens
  • Favicon for inclusionai
    inclusionAI: Ling 3.0 Flash VLLing 3.0 Flash VL
    2.08B tokens

    Ling 3.0 Flash VL builds on Ling 3.0 Flash (124B total / 5.5B active MoE from InclusionAI), further strengthening its language capabilities while adding native visual perception and advanced visual agent capabilities. Hybrid instant/reasoning model with tool calling.

    by inclusionaiSep 10, 2026131K context$0.06/M input tokens$0.18/M output tokens
  • Favicon for deepseek
    DeepSeek: DeepSeek V4.1 FlashDeepSeek V4.1 Flash
    55% off
    14.4T tokens
    Academia (#5)
    Finance (#2)
    Health (#2)
    Legal (#5)
    Marketing (#6)

    DeepSeek V4.1 Flash is a sparse mixture-of-experts model from DeepSeek, and the first built on the company's Causal Encoder-Decoder (CED) architecture. It activates 8B parameters on input and 16B on output from a 552B-parameter backbone, an asymmetric split that keeps per-token compute low relative to the model's total size. Image understanding is native to the architecture, with visual and text embeddings trained jointly from the start of pre-training rather than added afterward as in the earlier experimental V4 Flash Vision Exp. It is suited for coding, terminal, and computer-use agents, along with long-horizon tasks that must run to completion across many steps and long-context analysis. Compressed KV caching cuts cache memory to roughly a quarter of the previous Flash generation, significantly reducing costs on agentic workloads. DeepSeek positions it as the cost-efficient tier of the V4.1 family and reports that it exceeds V4 Pro on performance, speed, and task completion time.

    by deepseekSep 10, 20261.05M context$0.135/M input tokens$0.54/M output tokens
  • Favicon for openai
    OpenAI: GPT Image 2.5 SunburstGPT Image 2.5 Sunburst
    1.91B tokens

    GPT Image 2.5 Sunburst is an image generation and editing model from OpenAI, positioned as the precision-oriented tier of the GPT Image 2.5 series. It is suited to detailed creative work where editing accuracy matters more than generation speed, via the dedicated Images API.

    by openaiSep 9, 2026400K context$8/M input tokens$30/M output tokens