HyperInferSign in

Models

Sign in
You’re browsing the public catalog. Sign in to call any model with your API key or in Chat.Sign in to use →

DeepSeek: DeepSeek V4 Flash

MMLU-Pro86.4%HighGPQA Diamond87.4%HighSWE-bench Verified78.6%HighLiveCodeBench88.4%High

DeepSeek's efficiency-optimized Mixture-of-Experts model in the V4 family, built for fast, high-throughput inference with a 1M-token context window.

by deepseek|Apr 24, 2026|1.05M context|$0.15/M input tokens|$0.30/M output tokensHuggingFace

DeepSeek: DeepSeek V4 Pro

MMLU-Pro87.1%HighGPQA Diamond89.1%HighSWE-bench Verified79.4%HighLiveCodeBench89.8%High

The flagship of DeepSeek's V4 family — a large Mixture-of-Experts model with a 1M-token context window, aimed at demanding reasoning, coding, and long-horizon agent workloads.

by deepseek|Apr 24, 2026|1.05M context|$1.83/M input tokens|$3.66/M output tokensHuggingFace

DeepSeek: DeepSeek V3.2

MMLU-Pro85%GPQA Diamond82.4%SWE-bench Verified73.1%LiveCodeBench83.3%AIME 202593.1%

DeepSeek's unified chat-and-reasoning model using DeepSeek Sparse Attention (DSA) for efficient long-context inference; replaced the separate V3 and R1 lines at a single price point.

by deepseek|Dec 1, 2025|164K context|$0.59/M input tokens|$1.77/M output tokensHuggingFace

Alibaba: Qwen3.5 397B A17B

MMLU-Pro87.8%SWE-bench Verified76.4%LiveCodeBench83.6%MMMU85%IFEval92.6%

Alibaba's Qwen3.5 flagship: a 397B-parameter (17B active) native vision-language MoE that accepts text, image, and video input, with strong reasoning, coding, and agentic performance across 200+ languages.

by qwen|Feb 16, 2026|262K context|$0.79/M input tokens|$4.73/M output tokensHuggingFace

Alibaba: Qwen3 Coder 480B A35B

SWE-bench Pro (Scale AI)38.7%

A 480B-parameter (35B active) Mixture-of-Experts code model optimized for agentic coding — function calling, tool use, and repository-scale long-context reasoning — from the Qwen team.

by qwen|Jul 23, 2025|1.05M context|$1.05/M input tokens|$5.12/M output tokensHuggingFace

Z.ai: GLM 5.2

GPQA Diamond91.2%AIME 202699.2%SWE-bench Pro62.1%HLE40.5%Terminal-Bench 2.181%

Z.ai's flagship large-scale reasoning model, tuned for long-horizon agent workflows, project-level software engineering, and complex multi-step automation over a 1M-token context.

by z-ai|Jun 16, 2026|1.05M context|$1.47/M input tokens|$4.62/M output tokensHuggingFace

Z.ai: GLM 5.1

GPQA Diamond86.2%AIME 202695.3%SWE-bench Pro58.4%

Z.ai's prior-generation flagship, delivering strong coding and long-horizon agentic performance over a ~200K-token context.

by z-ai|Apr 7, 2026|203K context|$1.62/M input tokens|$5.09/M output tokensHuggingFace

Google: Gemma 4 31B

MMLU-Pro85.2%GPQA Diamond84.3%LiveCodeBench80%AIME 202689.2%MMMU-Pro76.9%

Google DeepMind's Gemma 4 31B — a 30.7B-parameter dense multimodal model taking text and image input, with a 256K-token context window and multilingual coverage across 140+ languages. Apache 2.0 license.

by google|Apr 2, 2026|262K context|$0.41/M input tokens|$1.02/M output tokensHuggingFace

Google: Gemma 4 26B A4B

MMLU-Pro82.6%GPQA Diamond82.3%LiveCodeBench77.1%AIME 202688.3%MMMU-Pro73.8%

Google DeepMind's Gemma 4 26B A4B — an instruction-tuned Mixture-of-Experts model with 25.2B total parameters and only 3.8B active per token, delivering near-31B quality at a fraction of the compute. Apache 2.0 license.

by google|Apr 3, 2026|262K context|$0.16/M input tokens|$0.63/M output tokensHuggingFace

Meta: Llama 4 Maverick

MMLU-Pro80.5%GPQA Diamond69.8%LiveCodeBench43.4%MMMU73.4%MMMU-Pro59.6%

Meta's Llama 4 mixture-of-experts flagship — 17B active parameters (400B total, 128 experts) with native early-fusion multimodality and a 1M-token context window, instruction-tuned for assistant behavior and image reasoning.

by meta-llama|Apr 5, 2025|1.05M context|$0.37/M input tokens|$1.21/M output tokensHuggingFace

Meta: Llama 3.3 70B Instruct

MMLU-Pro68.9%GPQA Diamond50.5%HumanEval88.4%IFEval92.1%MMLU86%

Meta's 70B dense, instruction-tuned multilingual model optimized for dialogue across eight languages; text in, text out. Knowledge cutoff Dec 2023.

by meta-llama|Dec 6, 2024|131K context|$1.10/M input tokens|$1.10/M output tokensHuggingFace

Meta: Llama 4 Scout

MMLU-Pro74.3%GPQA Diamond57.2%LiveCodeBench32.8%MMMU69.4%MMMU-Pro52.2%

Llama 4 Scout is a 17B-active (109B total, 16 experts) MoE model from Meta with native multimodal (text+image) input and an exceptionally long context window, tuned for retrieval-heavy and long-document workloads.

by meta-llama|Apr 5, 2025|1.31M context|$0.27/M input tokens|$0.74/M output tokensHuggingFace

Mistral AI: Mistral Small 3.2 24B

MMLU-Pro69.06%GPQA Diamond46.13%MMMU62.5%MMLU80.5%HumanEval Plus p@592.9%

An updated 24B instruction-tuned model from Mistral optimized for precise instruction following, reduced repetition, and robust function calling, with text + image input and structured output. Apache 2.0 license.

by mistralai|Jun 20, 2025|131K context|$0.11/M input tokens|$0.32/M output tokensHuggingFace

Mistral AI: Mistral Large 3

Mistral's most capable open-weight model to date — a granular mixture-of-experts (41B active / 675B total) general-purpose multimodal model with a natively fused vision encoder, released under Apache 2.0.

by mistralai|Dec 1, 2025|262K context|$0.53/M input tokens|$1.58/M output tokensHuggingFace

Moonshot AI: Kimi K2.6

GPQA Diamond90.5%SWE-bench Verified80.2%LiveCodeBench89.6%AIME 202696.4%MMMU-Pro79.4%

Moonshot AI's flagship open-weights model — a 1T-parameter native-multimodal mixture-of-experts (32B active) built for long-horizon coding, UI generation, and multi-agent orchestration.

by moonshotai|Apr 20, 2026|262K context|$1.26/M input tokens|$4.83/M output tokensHuggingFace

Moonshot AI: Kimi K2.7 Code

Kimi Code Bench v262%Program Bench53.6%MLS Bench Lite35.1%Kimi Claw 24/7 Bench46.9%MCP Atlas76%

The newest member of Moonshot AI's Kimi K2 family — a coding-focused agentic model that always reasons in a thinking mode, built to complete end-to-end programming tasks over long contexts.

by moonshotai|Jun 12, 2026|262K context|$2.00/M input tokens|$8.40/M output tokensHuggingFace