Skip to main content
BenchLM
Data
Guide and benchmark index

Small language models under 100B

Qwen3.8-27B leads this source-checked under-100B index at 58.2 overall. 32 of 47 models have a public overall rank. Use the size, weight-access, and license filters to narrow your candidates; these scores measure published capability, not performance on your local machine.

Benchmark data as of

Choose a small language model by total parameter count first, then compare the evidence for the task you need. This index covers source-checked models below 100B total parameters, with weight access and license terms listed separately. The size bands are navigation aids: a 70B model belongs in the index, but it is a very different deployment from a 3B model.

Guide by Glevd · Parameter sources checked 2026-10-05.

Compare small LLM benchmark scores

Overall, Coding, Agentic, and Knowledge scores use BenchAlign v5.8. Missing results stay visible. Four-bit figures are theoretical weights-only estimates in decimal GB, before KV cache and runtime overhead. Context is the catalog specification, not a measured local capacity. How the score is built.

Open weights means downloadable parameters. MIT and Apache 2.0 describe license terms; they do not by themselves establish that an AI system is fully open source. Every listed model needs a documented total below 100B.

47 of 47 models shown. Overall rank is within this index.

Source-checked models below 100B total parameters. Scores compare published capability; memory figures estimate only four-bit weight storage.
Overall rankModel and sourceTotal / activeOverall scoreCodingAgenticKnowledge4-bit weightsContext
1Qwen3.8-27BAlibaba · Open weightsApache-2.0Official model cardParameter source
Parameter note

27B language model; the 27.781B checkpoint total also includes vision and auxiliary parameters.

27.781BDense58.2supported48.161.449.713.891 GBWeights only262,144 native; up to 1M extended
2Ternary Bonsai 2 27BPrism ML · Open weightsApache-2.0Official model cardParameter source
Parameter note

Published 27.36B total includes the vision tower. Native ternary packing differs from the hypothetical 4-bit weight estimate; use the publisher file sizes for deployment.

27.36BDense53.1estimatedNot measuredNot measuredNot measured13.68 GBWeights only262K
3Qwen3.5-27BAlibaba · Open weightsApache-2.0Official model cardParameter source
Parameter note

27B language model; the 27.781B checkpoint total also includes vision and auxiliary parameters.

27.781BDense48.2estimated34.035.344.613.891 GBWeights only262,144 native; up to 1M extended
4Qwen3.6-27BAlibaba · Open weightsApache-2.0Official model cardParameter source
Parameter note

27B language model; the 27.781B checkpoint total also includes vision and auxiliary parameters.

27.781BDense47.9estimated36.727.845.213.891 GBWeights only262,144 native; up to 1M extended
5Qwen3.5-35B-A3BAlibaba · Open weightsApache-2.0Official model cardParameter source
Parameter note

35B language model with 3B active; 35.952B checkpoint total includes vision and auxiliary parameters.

35.952B3B active · MoE46.4supported27.331.441.717.976 GBWeights only262,144 native; up to 1M extended
6Gemma 4 26B A4BGoogle · Open weightsApache-2.0Official model cardParameter source
Parameter note

25.2B language model with 3.8B active; 25.806B checkpoint total also includes vision components.

25.806B3.8B active · MoE45.9supported32.6Not measured36.512.903 GBWeights only256K
7Qwen3.6-35B-A3BAlibaba · Open weightsApache-2.0Official model cardParameter source
Parameter note

35B language model with 3B active; 35.952B checkpoint total includes vision and auxiliary parameters.

35.952B3B active · MoE43.4estimated34.728.741.917.976 GBWeights only262,144 native; up to 1M extended
8Muse Glimmer 30BMeta · Open weightsApache-2.0Official model cardParameter source
Parameter note

29.777B total checkpoint parameters; the card describes roughly 29.6B including the perception encoder.

29.777BDense41.8estimated32.424.342.814.889 GBWeights only131,072+
9Gemma 4 31BGoogle · Open weightsApache-2.0Official model cardParameter source
Parameter note

30.7B language model plus vision components; 31.273B total stored checkpoint parameters.

31.273BDense40.4estimated33.218.938.615.637 GBWeights only256K
10GLM-4.7-FlashZ.AI · Open weightsMITOfficial model cardParameter source
Parameter note

31.221B total checkpoint parameters; the model card describes a 30B-A3B MoE.

31.221B3B active · MoE39.4estimatedNot measuredNot measuredNot measured15.611 GBWeights only202,752
11GPT-OSS 20BOpenAI · Open weightsApache-2.0Official model cardParameter source
Parameter note

OpenAI reports 21B total and 3.6B active. Native MXFP4 packing means tensor-element counts do not directly equal logical parameters.

21B3.6B active · MoE33.6supported20.114.433.010.5 GBWeights only131,072
12Granite 4.2 8BIBM · Open weightsApache-2.0Official model cardParameter source
Parameter note

Total checkpoint parameters from the publisher repository metadata, rounded to three decimals in billions.

8.792BDense31.8estimated18.114.028.84.396 GBWeights only128K native; up to 512K extended
13Gemma 4 12BGoogle · Open weightsApache-2.0Official model cardParameter source
Parameter note

11.960B total checkpoint parameters. The unified architecture handles text, images, and audio without separate encoders.

11.96BDense31.8estimated23.2Not measured32.55.98 GBWeights only256K
14ZAYA1-8BZyphra · Open weightsApache-2.0Official model cardParameter source
Parameter note

8.840B stored checkpoint parameters; the model card describes 8.4B total and 0.760B active. The index uses the larger stored total.

8.84B0.76B active · MoE31.5estimatedNot measuredNot measured33.94.42 GBWeights only131K
15DeepSeek R1 Distill Qwen 32BDeepSeek · Open weightsMITOfficial model cardParameter source
Parameter note

Total checkpoint parameters from the publisher repository metadata, rounded to three decimals in billions.

32.764BDense31.2estimatedNot measuredNot measured28.816.382 GBWeights only128K
16Gemma 4 E4BGoogle · Open weightsApache-2.0Official model cardParameter source
Parameter note

7.996B total with embeddings and multimodal components. E4B names 4.5B effective parameters; that is not a published active count.

7.996BDense31.2estimated14.2Not measured26.63.998 GBWeights only128K
17Nemotron 3 Nano Omni 30B A3BNVIDIA · Open weightsNVIDIA Open Model AgreementOfficial model cardParameter source
Parameter note

33.016B stored checkpoint parameters, including vision and speech encoders. The card reports approximately 3B active for the language backbone.

33.016B3B active · MoE30.9estimated16.6Not measured31.216.508 GBWeights only256K
18LFM2.5-2.6BLiquidAI · Open weightsLFM Open License v1.0Official model cardParameter source
Parameter note

Total checkpoint parameters from the publisher repository metadata, rounded to three decimals in billions.

2.697BHybrid30.6estimated13.315.525.31.349 GBWeights only131,072
19Gemma 4 E2BGoogle · Open weightsApache-2.0Official model cardParameter source
Parameter note

5.123B total with embeddings and multimodal components. E2B names 2.3B effective parameters; that is not a published active count.

5.123BDense30.1estimated12.6Not measured24.82.562 GBWeights only128K
20LFM2.5-8B-A1BLiquidAI · Open weightsLFM Open License v1.0Official model cardParameter source
Parameter note

8.468B checkpoint total; the model card reports 1.5B active despite the A1B name.

8.468B1.5B active · MoE29.4estimatedNot measuredNot measured25.14.234 GBWeights only128,000
21Qwen2.5-72BAlibaba · Open weightsQwen LicenseOfficial model cardParameter source
Parameter note

Total checkpoint parameters from the publisher repository metadata, rounded to three decimals in billions.

72.706BDense29.2estimatedNot measuredNot measuredNot measured36.353 GBWeights only32,768 default; up to 131,072 extended
22Sarvam 30BSarvam · Open weightsApache-2.0Official model cardParameter source
Parameter note

32.153B total checkpoint parameters. The published 2.4B active count excludes embeddings, so a full active count is left blank.

32.153BActive count not documented · MoE28.8estimatedNot measuredNot measured26.316.077 GBWeights only131,072
23Exaone 4.0 32BLG AI Research · Open weightsEXAONE Model LicenseOfficial model cardParameter source
Parameter note

Total checkpoint parameters from the publisher repository metadata, rounded to three decimals in billions.

32.003BDense28.6estimatedNot measuredNot measured27.016.002 GBWeights only131,072
24Exaone 4.0 1.2BLG AI Research · Open weightsEXAONE Model LicenseOfficial model cardParameter source
Parameter note

Total checkpoint parameters from the publisher repository metadata, rounded to three decimals in billions.

1.279BDense27.5estimatedNot measuredNot measured22.80.64 GBWeights only65,536
25Qwen2.5 Coder 32B InstructAlibaba · Open weightsApache-2.0Official model cardParameter source
Parameter note

Total checkpoint parameters from the publisher repository metadata, rounded to three decimals in billions.

32.764BDense26.8estimatedNot measuredNot measured24.416.382 GBWeights only32,768 default; up to 131,072 extended
26Phi-4Microsoft · Open weightsMITOfficial model cardParameter source
Parameter note

Total checkpoint parameters from the publisher repository metadata, rounded to three decimals in billions.

14.66BDense25.7supportedNot measuredNot measured23.17.33 GBWeights only16K
27Ministral 3 14BMistral · Open weightsApache-2.0Official model cardParameter source
Parameter note

Total checkpoint parameters from the publisher repository metadata, rounded to three decimals in billions.

13.945BDense25.1estimated12.06.4Not measured6.973 GBWeights only256K
28Llama 3 70BMeta · Open weightsLlama 3 Community LicenseOfficial model cardParameter source
Parameter note

Total checkpoint parameters from the publisher repository metadata, rounded to three decimals in billions.

70.554BDense20.9estimatedNot measuredNot measuredNot measured35.277 GBWeights only8K
29Ministral 3 8BMistral · Open weightsApache-2.0Official model cardParameter source
Parameter note

Total checkpoint parameters from the publisher repository metadata, rounded to three decimals in billions.

8.918BDense20.3estimated10.65.7Not measured4.459 GBWeights only256K
30Ministral 3 3BMistral · Open weightsApache-2.0Official model cardParameter source
Parameter note

3.849B total checkpoint parameters, including the vision encoder. The official release uses FP8 for language-model weights.

3.849BDense19.1estimated9.05.0Not measured1.925 GBWeights only256K
31Mistral 7B v0.3Mistral · Open weightsApache-2.0Official model cardParameter source
Parameter note

Total checkpoint parameters from the publisher repository metadata, rounded to three decimals in billions.

7.248BDense7.3estimatedNot measuredNot measuredNot measured3.624 GBWeights only32K
32MiniCPM5-1BOpenBMB · Open weightsApache-2.0Official model cardParameter source
Parameter note

1.081B total, including embeddings; the card separately reports 0.680B non-embedding parameters.

1.081BDense4.7estimatedNot measuredNot measured7.60.541 GBWeights only131,072
Not rankedGranite 4.2 30BIBM · Open weightsApache-2.0Official model cardParameter source
Parameter note

Total checkpoint parameters from the publisher repository metadata, rounded to three decimals in billions.

29.277BDenseNot rankedNot measuredNot measuredNot measured14.639 GBWeights only128K native; up to 512K extended
Not rankedGranite 4.2 3BIBM · Open weightsApache-2.0Official model cardParameter source
Parameter note

Total checkpoint parameters from the publisher repository metadata, rounded to three decimals in billions.

3.66BDenseNot rankedNot measuredNot measuredNot measured1.83 GBWeights only128K native; up to 512K extended
Not rankedLFM2.5-230MLiquidAI · Open weightsLFM Open License v1.0Official model cardParameter source
Parameter note

Total checkpoint parameters from the publisher repository metadata, rounded to three decimals in billions.

0.23BHybridNot rankedNot measuredNot measuredNot measured0.115 GBWeights only32,768
Not rankedLFM2.5-350MLiquidAI · Open weightsLFM Open License v1.0Official model cardParameter source
Parameter note

Total checkpoint parameters from the publisher repository metadata, rounded to three decimals in billions.

0.354BHybridNot rankedNot measuredNot measuredNot measured0.177 GBWeights only32,768
Not rankedLongCat-Flash-Lite-SparseMeituan · Open weightsMITOfficial model cardParameter source
Parameter note

69.127B stored checkpoint parameters; the card reports approximately 3B active per token. N-gram embeddings remain part of the total.

69.127B3B active · MoENot rankedNot measuredNot measuredNot measured34.564 GBWeights only1M
Not rankedMellum2-12B-A2.5B-InstructJetBrains · Open weightsApache-2.0Official model cardParameter source
Parameter note

Total checkpoint parameters from the publisher repository metadata, rounded to three decimals in billions.

12.15B2.5B active · MoENot rankedNot measuredNot measuredNot measured6.075 GBWeights only131,072
Not rankedMellum2-12B-A2.5B-ThinkingJetBrains · Open weightsApache-2.0Official model cardParameter source
Parameter note

Total checkpoint parameters from the publisher repository metadata, rounded to three decimals in billions.

12.15B2.5B active · MoENot rankedNot measuredNot measuredNot measured6.075 GBWeights only131,072
Not rankedMiniCPM5-2BOpenBMB · Open weightsApache-2.0Official model cardParameter source
Parameter note

2.517B total, including embeddings; the card separately reports 1.982B non-embedding parameters.

2.517BDenseNot rankedNot measuredNot measuredNot measured1.259 GBWeights only131,072
Not rankedMinistral 3 14B (Reasoning)Mistral · Open weightsApache-2.0Official model cardParameter source
Parameter note

Total checkpoint parameters from the publisher repository metadata, rounded to three decimals in billions.

13.945BDenseNot rankedNot measuredNot measuredNot measured6.973 GBWeights only256K
Not rankedMinistral 3 3B (Reasoning)Mistral · Open weightsApache-2.0Official model cardParameter source
Parameter note

Repository metadata reports 4.252B total. The card separately rounds the language model to 3.4B and the vision encoder to 0.4B.

4.252BDenseNot rankedNot measuredNot measuredNot measured2.126 GBWeights only256K
Not rankedMinistral 3 8B (Reasoning)Mistral · Open weightsApache-2.0Official model cardParameter source
Parameter note

Total checkpoint parameters from the publisher repository metadata, rounded to three decimals in billions.

8.918BDenseNot rankedNot measuredNot measuredNot measured4.459 GBWeights only256K
Not rankedMistral 8x7BMistral · Open weightsApache-2.0Official model cardParameter source
Parameter note

46.703B total. The eight experts share attention and embeddings, so 8 times 7B would overstate its size.

46.703BActive count not documented · MoENot rankedNot measuredNot measuredNot measured23.352 GBWeights only32K
Not rankedNemotron 3 Nano 30BNVIDIA · Open weightsNVIDIA Nemotron Open Model LicenseOfficial model cardParameter source
Parameter note

31.578B total checkpoint parameters; the card reports 3.5B active in a Mamba2-Transformer hybrid MoE.

31.578B3.5B active · MoENot rankedNot measuredNot measuredNot measured15.789 GBWeights only256K default; up to 1M
Not rankedPokee-Isaac 28BPokee AI · Closed weightsProprietaryTechnical reportParameter source
Parameter note

The original v0 technical report lists 28B dense in Table 10. This is a provider-reported rounded size, not an exact checkpoint count. The weights are proprietary; local or VPC access does not make them publicly downloadable.

28BDenseNot rankedNot measuredNot measuredNot measuredNot availableClosed weightsUp to 10M (provider)
Not rankedZAYA1-74B-PreviewZyphra · Open weightsApache-2.0Official model cardParameter source
Parameter note

74.794B stored checkpoint parameters and 4B active. This is a reasoning-base preview without chat tuning or RL post-training.

74.794B4B active · MoENot rankedNot measuredNot measuredNot measured37.397 GBWeights only256K

Start with the strongest ranked model in your size band

These picks follow the current overall ranking within each band. Check a task score before treating the general leader as your coding or agent choice.

Size bandCurrent overall leaderWhere this pick can lose
Tiny: below 10BGranite 4.2 8B31.8 overall · 8.792B totalA low parameter count does not establish reliable results on your task.
Small: 10–40BQwen3.8-27B58.2 overall · 27.781B totalThe complete runtime needs more memory than the theoretical weight size.
Larger: above 40B, below 100BQwen2.5-72B29.2 overall · 72.706B totalThese are larger deployments; inclusion does not establish laptop fit.

A size limit gives you a shortlist

Small language model, or SLM, has no single parameter cutoff. We use three bands here: tiny below 10B, small from 10B through 40B, and larger models above 40B but below 100B. They let you compare models within a size budget without implying that everything on this page fits a laptop.

Start in the smallest band your application can use, and compare its answers with a larger candidate on the same prompts. For a short extraction task, the smaller checkpoint can be enough even when its general score is lower. For a coding agent that must recover from errors across many steps, a broader capability gap can matter more.

The starting points above are the highest overall ranked checkpoints in each band of this source-checked index. They are benchmark candidates to evaluate, not winners from a local hardware test. A tiny-band leader can still lose when your task needs knowledge, document understanding, or reliable tool use that its score does not establish.

Total parameters decide the 100B cutoff

A dense model uses its main network for each token. A mixture-of-experts model, or MoE, selects part of its expert network for a token. That produces two useful numbers: total parameters describe the full checkpoint, while active parameters describe the portion used for a token.

We apply the strict cutoff to total parameters. A 120B MoE with 12B active parameters is outside this index. Its active count does not turn it into a 12B download or establish a 12B memory requirement. The runtime still has to store and manage the other experts, whether they reside on a GPU, in system RAM, or elsewhere.

Read the parameter note beside the source link when a name uses an effective count, a rounded size, or a language-only count. Embeddings and multimodal components can make the full checkpoint larger than the number in its name. The documented count and its scope are more useful than the marketing label.

Weight access and license terms answer different questions

Use the weight-access filter to separate downloadable weights from proprietary weights. A provider can offer a closed-weight model through an API, a private deployment, or a VPC without making its parameters publicly downloadable. Closed-weight models enter this index only when a publisher documents a model size below the 100B limit.

The license filter separates MIT or Apache 2.0 from community or custom terms. This describes the recorded license, not the openness of the entire AI system. Downloadable weights and a permissive license do not, by themselves, establish access to the training code and data information needed to reproduce the system. Check the publisher terms for the use you intend.

Pokee-Isaac 28B uses the original v0 report's rounded 28B dense size. Its size is provider-reported rather than counted from a public checkpoint. The index does not estimate downloadable weight storage for closed-weight models, and their private-deployment availability does not establish public access.

See the Open Source AI Definition for requirements beyond weight access.

Four-bit weight size is only a starting estimate

At exactly four bits per parameter, weight storage is total parameters multiplied by 0.5GB per billion parameters, using decimal GB. That gives 4GB for 8B, 13.5GB for 27B, and 35GB for 70B. The index computes this arithmetic for each downloadable checkpoint; these are theoretical weights-only figures.

A real quantized artifact also contains metadata and tensors that may use other precision. Runtime buffers, KV cache, context length, parallel requests, and any extra encoders add memory. Some checkpoint formats use mixed precision. The estimate therefore is not a measured file size, a minimum RAM specification, or a guarantee that a model fits a GPU.

Reserve memory for the operating system and your other applications on a shared-memory machine. On a discrete GPU, check VRAM separately from system RAM. CPU offload can make a configuration possible, but it changes the speed you need to measure. Increasing context or concurrency can break a configuration that worked for one short prompt.

For a configuration estimate, use the VRAM calculator, then verify the result in your runtime.

Benchmark scores compare capability, not local speed

The overall and task scores come from the same current public ranking used by the main leaderboards. We filter that ranking by the source-checked parameter registry and retain its order. We do not rescore a checkpoint for being smaller, borrow a sibling model's results, or replace missing scores with zero.

Supported means the public rank has enough independent evidence to stand behind it. Estimated means the position is visible with greater uncertainty. Not ranked means this checkpoint has no published overall position in the current ranking. A missing task score remains marked Not measured, even when the checkpoint has an overall score.

Published evaluations can use a different reasoning budget, precision, harness, or runtime from your local deployment. A score does not measure quantized answer quality, tokens per second, time to first token, or performance at the context length you plan to serve. Model pages carry the individual benchmark results and their sources so you can inspect what a score rests on.

This is a maintained index of documented checkpoints in our catalog, not every language model ever released. Models without a checked total count stay outside the index. Open weights also do not imply identical license terms: read the official model card and license before commercial use, redistribution, or fine-tuning.

Your own prompts decide which model earns the memory

Choose two candidates within your memory budget and one larger reference if you can run it. Keep the prompt, output limit, context, and scoring criteria the same. Record the exact checkpoint, quantization, runtime version, hardware, and reasoning settings so a later change can be compared with the first run.

Use examples from the work you actually need done. For extraction, check exact fields and unsupported values. For coding, run the tests and inspect the patch. For retrieval, include questions the supplied documents cannot answer and check whether the model invents an answer. Count successful tasks as well as retries; a fast failed answer does not finish the job.

Measure peak memory and latency with the context and concurrency you intend to use. Move to a larger checkpoint when it solves failures that matter enough to justify the extra memory or waiting time. Keep the smaller one when it meets the same acceptance criteria on your workload.

Questions

Are all models under 100B small language models?

No universal cutoff defines a small language model. This page uses under 100B total parameters as an explicit inclusion limit and separates tiny, small, and larger checkpoints. A 70B model can qualify for the index while requiring much more memory than the compact models usually described as SLMs.

Can I run an under-100B model on a laptop?

The cutoff alone cannot establish laptop fit. Available memory, checkpoint format, quantization, context, and runtime support decide whether a configuration works. The four-bit column estimates only theoretical weight storage. Check the exact artifact, leave room for the operating system and KV cache, and measure the complete configuration before relying on it.

Does a lower active parameter count mean less RAM?

An active parameter count describes the part of an MoE used for a token. It does not describe the entire checkpoint. Inactive experts still need storage and memory management. Total parameters are the safer starting point for weight-size estimates, and the runtime determines how experts are loaded or offloaded.

Are smaller language models always faster?

Parameter count alone does not prove a speed advantage on your machine. Hardware bandwidth, quantization kernels, attention, context length, reasoning output, and CPU offload can change latency. Compare the same workload on the same runtime and hardware, and record time to a usable answer as well as token generation speed.