# Small language models under 100B

Qwen3.8-27B leads this source-checked under-100B index at 58.2 overall. 32 of 47 models have a public overall rank. Use the size, weight-access, and license filters to narrow your candidates; these scores measure published capability, not performance on your local machine.

Choose a small language model by total parameter count first, then compare the evidence for the task you need. This index covers source-checked models below 100B total parameters, with weight access and license terms listed separately. The size bands are navigation aids: a 70B model belongs in the index, but it is a very different deployment from a 3B model.

By [Glevd](/authors/glevd)

Last updated: October 5, 2026

Parameter sources checked: 2026-10-05

Canonical page: https://benchlm.ai/best/small-language-models

## Compare small LLM benchmark scores

47 documented model variants have fewer than 100 billion total parameters; 32 have a public overall rank. Ranked rows follow the public BenchAlign v5.8 overall order. An unranked row remains visible without a score.

Open weights means downloadable parameters. MIT and Apache 2.0 describe license terms; they do not by themselves establish that an AI system is fully open source. Every listed model needs a documented total below 100B.

| Rank | Model | Weight access | Size band | Architecture | Total params | Active params | Score | Evidence | Coding | Knowledge | Agentic | Context | 4-bit weights (estimate) | License | Parameter source |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 1 | [Qwen3.8-27B](/models/qwen3-8-27b) | Open weights | small | dense | 27.781B | 27.781B | 58.2 | Supported | 48.1 | 49.7 | 61.4 | 262,144 native; up to 1M extended | 13.891 GB | [Apache-2.0](https://huggingface.co/Qwen/Qwen3.8-27B) | [Model card](https://huggingface.co/Qwen/Qwen3.8-27B) · [Parameter source](https://huggingface.co/api/models/Qwen/Qwen3.8-27B) |
| 2 | [Ternary Bonsai 2 27B](/models/ternary-bonsai-2-27b) | Open weights | small | dense | 27.36B | 27.36B | 53.1 | Estimated | Not measured | Not measured | Not measured | 262K | 13.68 GB | [Apache-2.0](https://huggingface.co/prism-ml/Ternary-Bonsai-2-27B-gguf) | [Model card](https://huggingface.co/prism-ml/Ternary-Bonsai-2-27B-gguf) · [Parameter source](https://huggingface.co/prism-ml/Ternary-Bonsai-2-27B-gguf) |
| 3 | [Qwen3.5-27B](/models/qwen3-5-27b) | Open weights | small | dense | 27.781B | 27.781B | 48.2 | Estimated | 34.0 | 44.6 | 35.3 | 262,144 native; up to 1M extended | 13.891 GB | [Apache-2.0](https://huggingface.co/Qwen/Qwen3.5-27B) | [Model card](https://huggingface.co/Qwen/Qwen3.5-27B) · [Parameter source](https://huggingface.co/api/models/Qwen/Qwen3.5-27B) |
| 4 | [Qwen3.6-27B](/models/qwen3-6-27b) | Open weights | small | dense | 27.781B | 27.781B | 47.9 | Estimated | 36.7 | 45.2 | 27.8 | 262,144 native; up to 1M extended | 13.891 GB | [Apache-2.0](https://huggingface.co/Qwen/Qwen3.6-27B) | [Model card](https://huggingface.co/Qwen/Qwen3.6-27B) · [Parameter source](https://huggingface.co/api/models/Qwen/Qwen3.6-27B) |
| 5 | [Qwen3.5-35B-A3B](/models/qwen3-5-35b-a3b) | Open weights | small | MoE | 35.952B | 3B | 46.4 | Supported | 27.3 | 41.7 | 31.4 | 262,144 native; up to 1M extended | 17.976 GB | [Apache-2.0](https://huggingface.co/Qwen/Qwen3.5-35B-A3B) | [Model card](https://huggingface.co/Qwen/Qwen3.5-35B-A3B) · [Parameter source](https://huggingface.co/api/models/Qwen/Qwen3.5-35B-A3B) |
| 6 | [Gemma 4 26B A4B](/models/gemma-4-26b-a4b) | Open weights | small | MoE | 25.806B | 3.8B | 45.9 | Supported | 32.6 | 36.5 | Not measured | 256K | 12.903 GB | [Apache-2.0](https://huggingface.co/google/gemma-4-26B-A4B-it) | [Model card](https://huggingface.co/google/gemma-4-26B-A4B-it) · [Parameter source](https://huggingface.co/api/models/google/gemma-4-26B-A4B-it) |
| 7 | [Qwen3.6-35B-A3B](/models/qwen3-6-35b-a3b) | Open weights | small | MoE | 35.952B | 3B | 43.4 | Estimated | 34.7 | 41.9 | 28.7 | 262,144 native; up to 1M extended | 17.976 GB | [Apache-2.0](https://huggingface.co/Qwen/Qwen3.6-35B-A3B) | [Model card](https://huggingface.co/Qwen/Qwen3.6-35B-A3B) · [Parameter source](https://huggingface.co/api/models/Qwen/Qwen3.6-35B-A3B) |
| 8 | [Muse Glimmer 30B](/models/muse-glimmer-30b) | Open weights | small | dense | 29.777B | 29.777B | 41.8 | Estimated | 32.4 | 42.8 | 24.3 | 131,072+ | 14.889 GB | [Apache-2.0](https://huggingface.co/meta-models/Muse-Glimmer-30B) | [Model card](https://huggingface.co/meta-models/Muse-Glimmer-30B) · [Parameter source](https://huggingface.co/api/models/meta-models/Muse-Glimmer-30B) |
| 9 | [Gemma 4 31B](/models/gemma-4-31b) | Open weights | small | dense | 31.273B | 31.273B | 40.4 | Estimated | 33.2 | 38.6 | 18.9 | 256K | 15.637 GB | [Apache-2.0](https://huggingface.co/google/gemma-4-31B-it) | [Model card](https://huggingface.co/google/gemma-4-31B-it) · [Parameter source](https://huggingface.co/api/models/google/gemma-4-31B-it) |
| 10 | [GLM-4.7-Flash](/models/glm-4-7-flash) | Open weights | small | MoE | 31.221B | 3B | 39.4 | Estimated | Not measured | Not measured | Not measured | 202,752 | 15.611 GB | [MIT](https://huggingface.co/zai-org/GLM-4.7-Flash) | [Model card](https://huggingface.co/zai-org/GLM-4.7-Flash) · [Parameter source](https://huggingface.co/api/models/zai-org/GLM-4.7-Flash) |
| 11 | [GPT-OSS 20B](/models/gpt-oss-20b) | Open weights | small | MoE | 21B | 3.6B | 33.6 | Supported | 20.1 | 33.0 | 14.4 | 131,072 | 10.5 GB | [Apache-2.0](https://huggingface.co/openai/gpt-oss-20b) | [Model card](https://huggingface.co/openai/gpt-oss-20b) · [Parameter source](https://huggingface.co/openai/gpt-oss-20b) |
| 12 | [Granite 4.2 8B](/models/granite-4-2-8b) | Open weights | tiny | dense | 8.792B | 8.792B | 31.8 | Estimated | 18.1 | 28.8 | 14.0 | 128K native; up to 512K extended | 4.396 GB | [Apache-2.0](https://huggingface.co/ibm-granite/granite-4.2-8b) | [Model card](https://huggingface.co/ibm-granite/granite-4.2-8b) · [Parameter source](https://huggingface.co/api/models/ibm-granite/granite-4.2-8b) |
| 13 | [Gemma 4 12B](/models/gemma-4-12b) | Open weights | small | dense | 11.96B | 11.96B | 31.8 | Estimated | 23.2 | 32.5 | Not measured | 256K | 5.98 GB | [Apache-2.0](https://huggingface.co/google/gemma-4-12B-it) | [Model card](https://huggingface.co/google/gemma-4-12B-it) · [Parameter source](https://huggingface.co/api/models/google/gemma-4-12B-it) |
| 14 | [ZAYA1-8B](/models/zaya1-8b) | Open weights | tiny | MoE | 8.84B | 0.76B | 31.5 | Estimated | Not measured | 33.9 | Not measured | 131K | 4.42 GB | [Apache-2.0](https://huggingface.co/Zyphra/ZAYA1-8B) | [Model card](https://huggingface.co/Zyphra/ZAYA1-8B) · [Parameter source](https://huggingface.co/api/models/Zyphra/ZAYA1-8B) |
| 15 | [DeepSeek R1 Distill Qwen 32B](/models/deepseek-r1-distill-qwen-32b) | Open weights | small | dense | 32.764B | 32.764B | 31.2 | Estimated | Not measured | 28.8 | Not measured | 128K | 16.382 GB | [MIT](https://huggingface.co/deepseek-ai/DeepSeek-R1-Distill-Qwen-32B) | [Model card](https://huggingface.co/deepseek-ai/DeepSeek-R1-Distill-Qwen-32B) · [Parameter source](https://huggingface.co/api/models/deepseek-ai/DeepSeek-R1-Distill-Qwen-32B) |
| 16 | [Gemma 4 E4B](/models/gemma-4-e4b) | Open weights | tiny | dense | 7.996B | Not documented | 31.2 | Estimated | 14.2 | 26.6 | Not measured | 128K | 3.998 GB | [Apache-2.0](https://huggingface.co/google/gemma-4-E4B-it) | [Model card](https://huggingface.co/google/gemma-4-E4B-it) · [Parameter source](https://huggingface.co/api/models/google/gemma-4-E4B-it) |
| 17 | [Nemotron 3 Nano Omni 30B A3B](/models/nemotron-3-nano-omni-30b-a3b) | Open weights | small | MoE | 33.016B | 3B | 30.9 | Estimated | 16.6 | 31.2 | Not measured | 256K | 16.508 GB | [NVIDIA Open Model Agreement](https://huggingface.co/nvidia/Nemotron-3-Nano-Omni-30B-A3B-Reasoning-BF16) | [Model card](https://huggingface.co/nvidia/Nemotron-3-Nano-Omni-30B-A3B-Reasoning-BF16) · [Parameter source](https://huggingface.co/api/models/nvidia/Nemotron-3-Nano-Omni-30B-A3B-Reasoning-BF16) |
| 18 | [LFM2.5-2.6B](/models/lfm2-5-2-6b) | Open weights | tiny | hybrid | 2.697B | 2.697B | 30.6 | Estimated | 13.3 | 25.3 | 15.5 | 131,072 | 1.349 GB | [LFM Open License v1.0](https://huggingface.co/LiquidAI/LFM2.5-2.6B) | [Model card](https://huggingface.co/LiquidAI/LFM2.5-2.6B) · [Parameter source](https://huggingface.co/api/models/LiquidAI/LFM2.5-2.6B) |
| 19 | [Gemma 4 E2B](/models/gemma-4-e2b) | Open weights | tiny | dense | 5.123B | Not documented | 30.1 | Estimated | 12.6 | 24.8 | Not measured | 128K | 2.562 GB | [Apache-2.0](https://huggingface.co/google/gemma-4-E2B-it) | [Model card](https://huggingface.co/google/gemma-4-E2B-it) · [Parameter source](https://huggingface.co/api/models/google/gemma-4-E2B-it) |
| 20 | [LFM2.5-8B-A1B](/models/lfm2-5-8b-a1b) | Open weights | tiny | MoE | 8.468B | 1.5B | 29.4 | Estimated | Not measured | 25.1 | Not measured | 128,000 | 4.234 GB | [LFM Open License v1.0](https://huggingface.co/LiquidAI/LFM2.5-8B-A1B) | [Model card](https://huggingface.co/LiquidAI/LFM2.5-8B-A1B) · [Parameter source](https://huggingface.co/api/models/LiquidAI/LFM2.5-8B-A1B) |
| 21 | [Qwen2.5-72B](/models/qwen2-5-72b) | Open weights | medium | dense | 72.706B | 72.706B | 29.2 | Estimated | Not measured | Not measured | Not measured | 32,768 default; up to 131,072 extended | 36.353 GB | [Qwen License](https://huggingface.co/Qwen/Qwen2.5-72B-Instruct) | [Model card](https://huggingface.co/Qwen/Qwen2.5-72B-Instruct) · [Parameter source](https://huggingface.co/api/models/Qwen/Qwen2.5-72B-Instruct) |
| 22 | [Sarvam 30B](/models/sarvam-30b) | Open weights | small | MoE | 32.153B | Not documented | 28.8 | Estimated | Not measured | 26.3 | Not measured | 131,072 | 16.077 GB | [Apache-2.0](https://huggingface.co/sarvamai/sarvam-30b) | [Model card](https://huggingface.co/sarvamai/sarvam-30b) · [Parameter source](https://huggingface.co/api/models/sarvamai/sarvam-30b) |
| 23 | [Exaone 4.0 32B](/models/exaone-4-0-32b) | Open weights | small | dense | 32.003B | 32.003B | 28.6 | Estimated | Not measured | 27.0 | Not measured | 131,072 | 16.002 GB | [EXAONE Model License](https://huggingface.co/LGAI-EXAONE/EXAONE-4.0-32B) | [Model card](https://huggingface.co/LGAI-EXAONE/EXAONE-4.0-32B) · [Parameter source](https://huggingface.co/api/models/LGAI-EXAONE/EXAONE-4.0-32B) |
| 24 | [Exaone 4.0 1.2B](/models/exaone-4-0-1-2b) | Open weights | tiny | dense | 1.279B | 1.279B | 27.5 | Estimated | Not measured | 22.8 | Not measured | 65,536 | 0.64 GB | [EXAONE Model License](https://huggingface.co/LGAI-EXAONE/EXAONE-4.0-1.2B) | [Model card](https://huggingface.co/LGAI-EXAONE/EXAONE-4.0-1.2B) · [Parameter source](https://huggingface.co/api/models/LGAI-EXAONE/EXAONE-4.0-1.2B) |
| 25 | [Qwen2.5 Coder 32B Instruct](/models/qwen2-5-coder-32b-instruct) | Open weights | small | dense | 32.764B | 32.764B | 26.8 | Estimated | Not measured | 24.4 | Not measured | 32,768 default; up to 131,072 extended | 16.382 GB | [Apache-2.0](https://huggingface.co/Qwen/Qwen2.5-Coder-32B-Instruct) | [Model card](https://huggingface.co/Qwen/Qwen2.5-Coder-32B-Instruct) · [Parameter source](https://huggingface.co/api/models/Qwen/Qwen2.5-Coder-32B-Instruct) |
| 26 | [Phi-4](/models/phi-4) | Open weights | small | dense | 14.66B | 14.66B | 25.7 | Supported | Not measured | 23.1 | Not measured | 16K | 7.33 GB | [MIT](https://huggingface.co/microsoft/phi-4) | [Model card](https://huggingface.co/microsoft/phi-4) · [Parameter source](https://huggingface.co/api/models/microsoft/phi-4) |
| 27 | [Ministral 3 14B](/models/ministral-3-14b) | Open weights | small | dense | 13.945B | 13.945B | 25.1 | Estimated | 12.0 | Not measured | 6.4 | 256K | 6.973 GB | [Apache-2.0](https://huggingface.co/mistralai/Ministral-3-14B-Instruct-2512) | [Model card](https://huggingface.co/mistralai/Ministral-3-14B-Instruct-2512) · [Parameter source](https://huggingface.co/api/models/mistralai/Ministral-3-14B-Instruct-2512) |
| 28 | [Llama 3 70B](/models/llama-3-70b) | Open weights | medium | dense | 70.554B | 70.554B | 20.9 | Estimated | Not measured | Not measured | Not measured | 8K | 35.277 GB | [Llama 3 Community License](https://huggingface.co/meta-llama/Meta-Llama-3-70B-Instruct) | [Model card](https://huggingface.co/meta-llama/Meta-Llama-3-70B-Instruct) · [Parameter source](https://huggingface.co/api/models/meta-llama/Meta-Llama-3-70B-Instruct) |
| 29 | [Ministral 3 8B](/models/ministral-3-8b) | Open weights | tiny | dense | 8.918B | 8.918B | 20.3 | Estimated | 10.6 | Not measured | 5.7 | 256K | 4.459 GB | [Apache-2.0](https://huggingface.co/mistralai/Ministral-3-8B-Instruct-2512) | [Model card](https://huggingface.co/mistralai/Ministral-3-8B-Instruct-2512) · [Parameter source](https://huggingface.co/api/models/mistralai/Ministral-3-8B-Instruct-2512) |
| 30 | [Ministral 3 3B](/models/ministral-3-3b) | Open weights | tiny | dense | 3.849B | 3.849B | 19.1 | Estimated | 9.0 | Not measured | 5.0 | 256K | 1.925 GB | [Apache-2.0](https://huggingface.co/mistralai/Ministral-3-3B-Instruct-2512) | [Model card](https://huggingface.co/mistralai/Ministral-3-3B-Instruct-2512) · [Parameter source](https://huggingface.co/api/models/mistralai/Ministral-3-3B-Instruct-2512) |
| 31 | [Mistral 7B v0.3](/models/mistral-7b-v0-3) | Open weights | tiny | dense | 7.248B | 7.248B | 7.3 | Estimated | Not measured | Not measured | Not measured | 32K | 3.624 GB | [Apache-2.0](https://huggingface.co/mistralai/Mistral-7B-Instruct-v0.3) | [Model card](https://huggingface.co/mistralai/Mistral-7B-Instruct-v0.3) · [Parameter source](https://huggingface.co/api/models/mistralai/Mistral-7B-Instruct-v0.3) |
| 32 | [MiniCPM5-1B](/models/minicpm5-1b) | Open weights | tiny | dense | 1.081B | 1.081B | 4.7 | Estimated | Not measured | 7.6 | Not measured | 131,072 | 0.541 GB | [Apache-2.0](https://huggingface.co/openbmb/MiniCPM5-1B) | [Model card](https://huggingface.co/openbmb/MiniCPM5-1B) · [Parameter source](https://huggingface.co/api/models/openbmb/MiniCPM5-1B) |
| Unranked | [LFM2.5-230M](/models/lfm2-5-230m) | Open weights | tiny | hybrid | 0.23B | 0.23B | Not ranked | Not ranked | Not measured | Not measured | Not measured | 32,768 | 0.115 GB | [LFM Open License v1.0](https://huggingface.co/LiquidAI/LFM2.5-230M) | [Model card](https://huggingface.co/LiquidAI/LFM2.5-230M) · [Parameter source](https://huggingface.co/api/models/LiquidAI/LFM2.5-230M) |
| Unranked | [LFM2.5-350M](/models/lfm2-5-350m) | Open weights | tiny | hybrid | 0.354B | 0.354B | Not ranked | Not ranked | Not measured | Not measured | Not measured | 32,768 | 0.177 GB | [LFM Open License v1.0](https://huggingface.co/LiquidAI/LFM2.5-350M) | [Model card](https://huggingface.co/LiquidAI/LFM2.5-350M) · [Parameter source](https://huggingface.co/api/models/LiquidAI/LFM2.5-350M) |
| Unranked | [MiniCPM5-2B](/models/minicpm5-2b) | Open weights | tiny | dense | 2.517B | 2.517B | Not ranked | Not ranked | Not measured | Not measured | Not measured | 131,072 | 1.259 GB | [Apache-2.0](https://huggingface.co/openbmb/MiniCPM5-2B) | [Model card](https://huggingface.co/openbmb/MiniCPM5-2B) · [Parameter source](https://huggingface.co/api/models/openbmb/MiniCPM5-2B) |
| Unranked | [Granite 4.2 3B](/models/granite-4-2-3b) | Open weights | tiny | dense | 3.66B | 3.66B | Not ranked | Not ranked | Not measured | Not measured | Not measured | 128K native; up to 512K extended | 1.83 GB | [Apache-2.0](https://huggingface.co/ibm-granite/granite-4.2-3b) | [Model card](https://huggingface.co/ibm-granite/granite-4.2-3b) · [Parameter source](https://huggingface.co/api/models/ibm-granite/granite-4.2-3b) |
| Unranked | [Ministral 3 3B (Reasoning)](/models/ministral-3-3b-reasoning) | Open weights | tiny | dense | 4.252B | 4.252B | Not ranked | Not ranked | Not measured | Not measured | Not measured | 256K | 2.126 GB | [Apache-2.0](https://huggingface.co/mistralai/Ministral-3-3B-Reasoning-2512) | [Model card](https://huggingface.co/mistralai/Ministral-3-3B-Reasoning-2512) · [Parameter source](https://huggingface.co/api/models/mistralai/Ministral-3-3B-Reasoning-2512) |
| Unranked | [Ministral 3 8B (Reasoning)](/models/ministral-3-8b-reasoning) | Open weights | tiny | dense | 8.918B | 8.918B | Not ranked | Not ranked | Not measured | Not measured | Not measured | 256K | 4.459 GB | [Apache-2.0](https://huggingface.co/mistralai/Ministral-3-8B-Reasoning-2512) | [Model card](https://huggingface.co/mistralai/Ministral-3-8B-Reasoning-2512) · [Parameter source](https://huggingface.co/api/models/mistralai/Ministral-3-8B-Reasoning-2512) |
| Unranked | [Mellum2-12B-A2.5B-Instruct](/models/mellum2-12b-a2-5b-instruct) | Open weights | small | MoE | 12.15B | 2.5B | Not ranked | Not ranked | Not measured | Not measured | Not measured | 131,072 | 6.075 GB | [Apache-2.0](https://huggingface.co/JetBrains/Mellum2-12B-A2.5B-Instruct) | [Model card](https://huggingface.co/JetBrains/Mellum2-12B-A2.5B-Instruct) · [Parameter source](https://huggingface.co/api/models/JetBrains/Mellum2-12B-A2.5B-Instruct) |
| Unranked | [Mellum2-12B-A2.5B-Thinking](/models/mellum2-12b-a2-5b-thinking) | Open weights | small | MoE | 12.15B | 2.5B | Not ranked | Not ranked | Not measured | Not measured | Not measured | 131,072 | 6.075 GB | [Apache-2.0](https://huggingface.co/JetBrains/Mellum2-12B-A2.5B-Thinking) | [Model card](https://huggingface.co/JetBrains/Mellum2-12B-A2.5B-Thinking) · [Parameter source](https://huggingface.co/api/models/JetBrains/Mellum2-12B-A2.5B-Thinking) |
| Unranked | [Ministral 3 14B (Reasoning)](/models/ministral-3-14b-reasoning) | Open weights | small | dense | 13.945B | 13.945B | Not ranked | Not ranked | Not measured | Not measured | Not measured | 256K | 6.973 GB | [Apache-2.0](https://huggingface.co/mistralai/Ministral-3-14B-Reasoning-2512) | [Model card](https://huggingface.co/mistralai/Ministral-3-14B-Reasoning-2512) · [Parameter source](https://huggingface.co/api/models/mistralai/Ministral-3-14B-Reasoning-2512) |
| Unranked | [Pokee-Isaac 28B](/models/pokee-isaac-28b) | Closed weights | small | dense | 28B | 28B | Not ranked | Not ranked | Not measured | Not measured | Not measured | Up to 10M (provider) | Not available (closed weights) | [Proprietary](https://pokee.ai/legal/terms-of-service) | [Technical report](https://console.pokee.ai/pokee-isaac-28b-v0-technical-report.pdf) · [Parameter source](https://console.pokee.ai/pokee-isaac-28b-v0-technical-report.pdf) |
| Unranked | [Granite 4.2 30B](/models/granite-4-2-30b) | Open weights | small | dense | 29.277B | 29.277B | Not ranked | Not ranked | Not measured | Not measured | Not measured | 128K native; up to 512K extended | 14.639 GB | [Apache-2.0](https://huggingface.co/ibm-granite/granite-4.2-30b) | [Model card](https://huggingface.co/ibm-granite/granite-4.2-30b) · [Parameter source](https://huggingface.co/api/models/ibm-granite/granite-4.2-30b) |
| Unranked | [Nemotron 3 Nano 30B](/models/nemotron-3-nano-30b) | Open weights | small | MoE | 31.578B | 3.5B | Not ranked | Not ranked | Not measured | Not measured | Not measured | 256K default; up to 1M | 15.789 GB | [NVIDIA Nemotron Open Model License](https://huggingface.co/nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16) | [Model card](https://huggingface.co/nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16) · [Parameter source](https://huggingface.co/api/models/nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16) |
| Unranked | [Mistral 8x7B](/models/mistral-8x7b) | Open weights | medium | MoE | 46.703B | Not documented | Not ranked | Not ranked | Not measured | Not measured | Not measured | 32K | 23.352 GB | [Apache-2.0](https://huggingface.co/mistralai/Mixtral-8x7B-Instruct-v0.1) | [Model card](https://huggingface.co/mistralai/Mixtral-8x7B-Instruct-v0.1) · [Parameter source](https://huggingface.co/api/models/mistralai/Mixtral-8x7B-Instruct-v0.1) |
| Unranked | [LongCat-Flash-Lite-Sparse](/models/longcat-flash-lite-sparse) | Open weights | medium | MoE | 69.127B | 3B | Not ranked | Not ranked | Not measured | Not measured | Not measured | 1M | 34.564 GB | [MIT](https://huggingface.co/meituan-longcat/LongCat-Flash-Lite-Sparse) | [Model card](https://huggingface.co/meituan-longcat/LongCat-Flash-Lite-Sparse) · [Parameter source](https://huggingface.co/api/models/meituan-longcat/LongCat-Flash-Lite-Sparse) |
| Unranked | [ZAYA1-74B-Preview](/models/zaya1-74b-preview) | Open weights | medium | MoE | 74.794B | 4B | Not ranked | Not ranked | Not measured | Not measured | Not measured | 256K | 37.397 GB | [Apache-2.0](https://huggingface.co/Zyphra/ZAYA1-74B-Preview) | [Model card](https://huggingface.co/Zyphra/ZAYA1-74B-Preview) · [Parameter source](https://huggingface.co/api/models/Zyphra/ZAYA1-74B-Preview) |

Total parameters determine inclusion. Active parameters describe how much of a mixture-of-experts model runs for each token. The 4-bit column estimates weight storage; runtime memory, KV cache, and quantization overhead require more memory. Public benchmark scores do not measure a local quantized run.

### Parameter notes

- [Qwen3.8-27B](/models/qwen3-8-27b): 27B language model; the 27.781B checkpoint total also includes vision and auxiliary parameters.
- [Ternary Bonsai 2 27B](/models/ternary-bonsai-2-27b): Published 27.36B total includes the vision tower. Native ternary packing differs from the hypothetical 4-bit weight estimate; use the publisher file sizes for deployment.
- [Qwen3.5-27B](/models/qwen3-5-27b): 27B language model; the 27.781B checkpoint total also includes vision and auxiliary parameters.
- [Qwen3.6-27B](/models/qwen3-6-27b): 27B language model; the 27.781B checkpoint total also includes vision and auxiliary parameters.
- [Qwen3.5-35B-A3B](/models/qwen3-5-35b-a3b): 35B language model with 3B active; 35.952B checkpoint total includes vision and auxiliary parameters.
- [Gemma 4 26B A4B](/models/gemma-4-26b-a4b): 25.2B language model with 3.8B active; 25.806B checkpoint total also includes vision components.
- [Qwen3.6-35B-A3B](/models/qwen3-6-35b-a3b): 35B language model with 3B active; 35.952B checkpoint total includes vision and auxiliary parameters.
- [Muse Glimmer 30B](/models/muse-glimmer-30b): 29.777B total checkpoint parameters; the card describes roughly 29.6B including the perception encoder.
- [Gemma 4 31B](/models/gemma-4-31b): 30.7B language model plus vision components; 31.273B total stored checkpoint parameters.
- [GLM-4.7-Flash](/models/glm-4-7-flash): 31.221B total checkpoint parameters; the model card describes a 30B-A3B MoE.
- [GPT-OSS 20B](/models/gpt-oss-20b): OpenAI reports 21B total and 3.6B active. Native MXFP4 packing means tensor-element counts do not directly equal logical parameters.
- [Granite 4.2 8B](/models/granite-4-2-8b): Total checkpoint parameters from the publisher repository metadata, rounded to three decimals in billions.
- [Gemma 4 12B](/models/gemma-4-12b): 11.960B total checkpoint parameters. The unified architecture handles text, images, and audio without separate encoders.
- [ZAYA1-8B](/models/zaya1-8b): 8.840B stored checkpoint parameters; the model card describes 8.4B total and 0.760B active. The index uses the larger stored total.
- [DeepSeek R1 Distill Qwen 32B](/models/deepseek-r1-distill-qwen-32b): Total checkpoint parameters from the publisher repository metadata, rounded to three decimals in billions.
- [Gemma 4 E4B](/models/gemma-4-e4b): 7.996B total with embeddings and multimodal components. E4B names 4.5B effective parameters; that is not a published active count.
- [Nemotron 3 Nano Omni 30B A3B](/models/nemotron-3-nano-omni-30b-a3b): 33.016B stored checkpoint parameters, including vision and speech encoders. The card reports approximately 3B active for the language backbone.
- [LFM2.5-2.6B](/models/lfm2-5-2-6b): Total checkpoint parameters from the publisher repository metadata, rounded to three decimals in billions.
- [Gemma 4 E2B](/models/gemma-4-e2b): 5.123B total with embeddings and multimodal components. E2B names 2.3B effective parameters; that is not a published active count.
- [LFM2.5-8B-A1B](/models/lfm2-5-8b-a1b): 8.468B checkpoint total; the model card reports 1.5B active despite the A1B name.
- [Qwen2.5-72B](/models/qwen2-5-72b): Total checkpoint parameters from the publisher repository metadata, rounded to three decimals in billions.
- [Sarvam 30B](/models/sarvam-30b): 32.153B total checkpoint parameters. The published 2.4B active count excludes embeddings, so a full active count is left blank.
- [Exaone 4.0 32B](/models/exaone-4-0-32b): Total checkpoint parameters from the publisher repository metadata, rounded to three decimals in billions.
- [Exaone 4.0 1.2B](/models/exaone-4-0-1-2b): Total checkpoint parameters from the publisher repository metadata, rounded to three decimals in billions.
- [Qwen2.5 Coder 32B Instruct](/models/qwen2-5-coder-32b-instruct): Total checkpoint parameters from the publisher repository metadata, rounded to three decimals in billions.
- [Phi-4](/models/phi-4): Total checkpoint parameters from the publisher repository metadata, rounded to three decimals in billions.
- [Ministral 3 14B](/models/ministral-3-14b): Total checkpoint parameters from the publisher repository metadata, rounded to three decimals in billions.
- [Llama 3 70B](/models/llama-3-70b): Total checkpoint parameters from the publisher repository metadata, rounded to three decimals in billions.
- [Ministral 3 8B](/models/ministral-3-8b): Total checkpoint parameters from the publisher repository metadata, rounded to three decimals in billions.
- [Ministral 3 3B](/models/ministral-3-3b): 3.849B total checkpoint parameters, including the vision encoder. The official release uses FP8 for language-model weights.
- [Mistral 7B v0.3](/models/mistral-7b-v0-3): Total checkpoint parameters from the publisher repository metadata, rounded to three decimals in billions.
- [MiniCPM5-1B](/models/minicpm5-1b): 1.081B total, including embeddings; the card separately reports 0.680B non-embedding parameters.
- [LFM2.5-230M](/models/lfm2-5-230m): Total checkpoint parameters from the publisher repository metadata, rounded to three decimals in billions.
- [LFM2.5-350M](/models/lfm2-5-350m): Total checkpoint parameters from the publisher repository metadata, rounded to three decimals in billions.
- [MiniCPM5-2B](/models/minicpm5-2b): 2.517B total, including embeddings; the card separately reports 1.982B non-embedding parameters.
- [Granite 4.2 3B](/models/granite-4-2-3b): Total checkpoint parameters from the publisher repository metadata, rounded to three decimals in billions.
- [Ministral 3 3B (Reasoning)](/models/ministral-3-3b-reasoning): Repository metadata reports 4.252B total. The card separately rounds the language model to 3.4B and the vision encoder to 0.4B.
- [Ministral 3 8B (Reasoning)](/models/ministral-3-8b-reasoning): Total checkpoint parameters from the publisher repository metadata, rounded to three decimals in billions.
- [Mellum2-12B-A2.5B-Instruct](/models/mellum2-12b-a2-5b-instruct): Total checkpoint parameters from the publisher repository metadata, rounded to three decimals in billions.
- [Mellum2-12B-A2.5B-Thinking](/models/mellum2-12b-a2-5b-thinking): Total checkpoint parameters from the publisher repository metadata, rounded to three decimals in billions.
- [Ministral 3 14B (Reasoning)](/models/ministral-3-14b-reasoning): Total checkpoint parameters from the publisher repository metadata, rounded to three decimals in billions.
- [Pokee-Isaac 28B](/models/pokee-isaac-28b): The original v0 technical report lists 28B dense in Table 10. This is a provider-reported rounded size, not an exact checkpoint count. The weights are proprietary; local or VPC access does not make them publicly downloadable.
- [Granite 4.2 30B](/models/granite-4-2-30b): Total checkpoint parameters from the publisher repository metadata, rounded to three decimals in billions.
- [Nemotron 3 Nano 30B](/models/nemotron-3-nano-30b): 31.578B total checkpoint parameters; the card reports 3.5B active in a Mamba2-Transformer hybrid MoE.
- [Mistral 8x7B](/models/mistral-8x7b): 46.703B total. The eight experts share attention and embeddings, so 8 times 7B would overstate its size.
- [LongCat-Flash-Lite-Sparse](/models/longcat-flash-lite-sparse): 69.127B stored checkpoint parameters; the card reports approximately 3B active per token. N-gram embeddings remain part of the total.
- [ZAYA1-74B-Preview](/models/zaya1-74b-preview): 74.794B stored checkpoint parameters and 4B active. This is a reasoning-base preview without chat tuning or RL post-training.

## Start with the strongest ranked model in your size band

These picks follow the current overall ranking within each band. Check a task score before treating the general leader as your coding or agent choice.

| Size band | Current overall leader | Where this pick can lose |
|---|---|---|
| Tiny: below 10B | [Granite 4.2 8B](/models/granite-4-2-8b) — 31.8 overall, 8.792B total | A low parameter count does not establish reliable results on your task. |
| Small: 10–40B | [Qwen3.8-27B](/models/qwen3-8-27b) — 58.2 overall, 27.781B total | The complete runtime needs more memory than the theoretical weight size. |
| Larger: above 40B, below 100B | [Qwen2.5-72B](/models/qwen2-5-72b) — 29.2 overall, 72.706B total | These are larger deployments; inclusion does not establish laptop fit. |

## A size limit gives you a shortlist

Small language model, or SLM, has no single parameter cutoff. We use three bands here: tiny below 10B, small from 10B through 40B, and larger models above 40B but below 100B. They let you compare models within a size budget without implying that everything on this page fits a laptop.

Start in the smallest band your application can use, and compare its answers with a larger candidate on the same prompts. For a short extraction task, the smaller checkpoint can be enough even when its general score is lower. For a coding agent that must recover from errors across many steps, a broader capability gap can matter more.

The starting points above are the highest overall ranked checkpoints in each band of this source-checked index. They are benchmark candidates to evaluate, not winners from a local hardware test. A tiny-band leader can still lose when your task needs knowledge, document understanding, or reliable tool use that its score does not establish.

## Total parameters decide the 100B cutoff

A dense model uses its main network for each token. A mixture-of-experts model, or MoE, selects part of its expert network for a token. That produces two useful numbers: total parameters describe the full checkpoint, while active parameters describe the portion used for a token.

We apply the strict cutoff to total parameters. A 120B MoE with 12B active parameters is outside this index. Its active count does not turn it into a 12B download or establish a 12B memory requirement. The runtime still has to store and manage the other experts, whether they reside on a GPU, in system RAM, or elsewhere.

Read the parameter note beside the source link when a name uses an effective count, a rounded size, or a language-only count. Embeddings and multimodal components can make the full checkpoint larger than the number in its name. The documented count and its scope are more useful than the marketing label.

## Weight access and license terms answer different questions

Use the weight-access filter to separate downloadable weights from proprietary weights. A provider can offer a closed-weight model through an API, a private deployment, or a VPC without making its parameters publicly downloadable. Closed-weight models enter this index only when a publisher documents a model size below the 100B limit.

The license filter separates MIT or Apache 2.0 from community or custom terms. This describes the recorded license, not the openness of the entire AI system. Downloadable weights and a permissive license do not, by themselves, establish access to the training code and data information needed to reproduce the system. Check the publisher terms for the use you intend.

Pokee-Isaac 28B uses the original v0 report's rounded 28B dense size. Its size is provider-reported rather than counted from a public checkpoint. The index does not estimate downloadable weight storage for closed-weight models, and their private-deployment availability does not establish public access.

See the [Open Source AI Definition](https://opensource.org/ai/open-source-ai-definition) for requirements beyond weight access.

## Four-bit weight size is only a starting estimate

At exactly four bits per parameter, weight storage is total parameters multiplied by 0.5GB per billion parameters, using decimal GB. That gives 4GB for 8B, 13.5GB for 27B, and 35GB for 70B. The index computes this arithmetic for each downloadable checkpoint; these are theoretical weights-only figures.

A real quantized artifact also contains metadata and tensors that may use other precision. Runtime buffers, KV cache, context length, parallel requests, and any extra encoders add memory. Some checkpoint formats use mixed precision. The estimate therefore is not a measured file size, a minimum RAM specification, or a guarantee that a model fits a GPU.

Reserve memory for the operating system and your other applications on a shared-memory machine. On a discrete GPU, check VRAM separately from system RAM. CPU offload can make a configuration possible, but it changes the speed you need to measure. Increasing context or concurrency can break a configuration that worked for one short prompt.

For a configuration estimate, use the [VRAM calculator](/tools/llm-vram-calculator), then verify the result in your runtime.

## Benchmark scores compare capability, not local speed

The overall and task scores come from the same current public ranking used by the main leaderboards. We filter that ranking by the source-checked parameter registry and retain its order. We do not rescore a checkpoint for being smaller, borrow a sibling model's results, or replace missing scores with zero.

Supported means the public rank has enough independent evidence to stand behind it. Estimated means the position is visible with greater uncertainty. Not ranked means this checkpoint has no published overall position in the current ranking. A missing task score remains marked Not measured, even when the checkpoint has an overall score.

Published evaluations can use a different reasoning budget, precision, harness, or runtime from your local deployment. A score does not measure quantized answer quality, tokens per second, time to first token, or performance at the context length you plan to serve. Model pages carry the individual benchmark results and their sources so you can inspect what a score rests on.

This is a maintained index of documented checkpoints in our catalog, not every language model ever released. Models without a checked total count stay outside the index. Open weights also do not imply identical license terms: read the official model card and license before commercial use, redistribution, or fine-tuning.

## Your own prompts decide which model earns the memory

Choose two candidates within your memory budget and one larger reference if you can run it. Keep the prompt, output limit, context, and scoring criteria the same. Record the exact checkpoint, quantization, runtime version, hardware, and reasoning settings so a later change can be compared with the first run.

Use examples from the work you actually need done. For extraction, check exact fields and unsupported values. For coding, run the tests and inspect the patch. For retrieval, include questions the supplied documents cannot answer and check whether the model invents an answer. Count successful tasks as well as retries; a fast failed answer does not finish the job.

Measure peak memory and latency with the context and concurrency you intend to use. Move to a larger checkpoint when it solves failures that matter enough to justify the extra memory or waiting time. Keep the smaller one when it meets the same acceptance criteria on your workload.

## Questions

### Are all models under 100B small language models?

No universal cutoff defines a small language model. This page uses under 100B total parameters as an explicit inclusion limit and separates tiny, small, and larger checkpoints. A 70B model can qualify for the index while requiring much more memory than the compact models usually described as SLMs.

### Can I run an under-100B model on a laptop?

The cutoff alone cannot establish laptop fit. Available memory, checkpoint format, quantization, context, and runtime support decide whether a configuration works. The four-bit column estimates only theoretical weight storage. Check the exact artifact, leave room for the operating system and KV cache, and measure the complete configuration before relying on it.

### Does a lower active parameter count mean less RAM?

An active parameter count describes the part of an MoE used for a token. It does not describe the entire checkpoint. Inactive experts still need storage and memory management. Total parameters are the safer starting point for weight-size estimates, and the runtime determines how experts are loaded or offloaded.

### Are smaller language models always faster?

Parameter count alone does not prove a speed advantage on your machine. Hardware bandwidth, quantization kernels, attention, context length, reasoning output, and CPU offload can change latency. Compare the same workload on the same runtime and hardware, and record time to a usable answer as well as token generation speed.

## Related guides and tools

- [Best local LLMs](/best/local-llm): Compare local deployment needs.
- [Best Ollama models](/best/ollama-models): Check runner-specific model choices.
- [Open-weight leaderboard](/best/open-source): Compare the wider benchmark ranking.
- [LLM VRAM calculator](/tools/llm-vram-calculator): Estimate memory for a model and quantization.
