Small language models under 100B
Qwen3.8-27B leads this source-checked under-100B index at 58.2 overall. 32 of 47 models have a public overall rank. Use the size, weight-access, and license filters to narrow your candidates; these scores measure published capability, not performance on your local machine.
Choose a small language model by total parameter count first, then compare the evidence for the task you need. This index covers source-checked models below 100B total parameters, with weight access and license terms listed separately. The size bands are navigation aids: a 70B model belongs in the index, but it is a very different deployment from a 3B model.
Guide by Glevd · Parameter sources checked 2026-10-05.
Compare small LLM benchmark scores
Overall, Coding, Agentic, and Knowledge scores use BenchAlign v5.8. Missing results stay visible. Four-bit figures are theoretical weights-only estimates in decimal GB, before KV cache and runtime overhead. Context is the catalog specification, not a measured local capacity. How the score is built.
Open weights means downloadable parameters. MIT and Apache 2.0 describe license terms; they do not by themselves establish that an AI system is fully open source. Every listed model needs a documented total below 100B.
47 of 47 models shown. Overall rank is within this index.
| Overall rank | Model and source | Total / active | Overall score | Coding | Agentic | Knowledge | 4-bit weights | Context |
|---|---|---|---|---|---|---|---|---|
| 1 | Qwen3.8-27BAlibaba · Open weightsApache-2.0Official model cardParameter sourceParameter note27B language model; the 27.781B checkpoint total also includes vision and auxiliary parameters. | 27.781BDense | 58.2supported | 48.1 | 61.4 | 49.7 | 13.891 GBWeights only | 262,144 native; up to 1M extended |
| 2 | Ternary Bonsai 2 27BPrism ML · Open weightsApache-2.0Official model cardParameter sourceParameter notePublished 27.36B total includes the vision tower. Native ternary packing differs from the hypothetical 4-bit weight estimate; use the publisher file sizes for deployment. | 27.36BDense | 53.1estimated | Not measured | Not measured | Not measured | 13.68 GBWeights only | 262K |
| 3 | Qwen3.5-27BAlibaba · Open weightsApache-2.0Official model cardParameter sourceParameter note27B language model; the 27.781B checkpoint total also includes vision and auxiliary parameters. | 27.781BDense | 48.2estimated | 34.0 | 35.3 | 44.6 | 13.891 GBWeights only | 262,144 native; up to 1M extended |
| 4 | Qwen3.6-27BAlibaba · Open weightsApache-2.0Official model cardParameter sourceParameter note27B language model; the 27.781B checkpoint total also includes vision and auxiliary parameters. | 27.781BDense | 47.9estimated | 36.7 | 27.8 | 45.2 | 13.891 GBWeights only | 262,144 native; up to 1M extended |
| 5 | Qwen3.5-35B-A3BAlibaba · Open weightsApache-2.0Official model cardParameter sourceParameter note35B language model with 3B active; 35.952B checkpoint total includes vision and auxiliary parameters. | 35.952B3B active · MoE | 46.4supported | 27.3 | 31.4 | 41.7 | 17.976 GBWeights only | 262,144 native; up to 1M extended |
| 6 | Gemma 4 26B A4BGoogle · Open weightsApache-2.0Official model cardParameter sourceParameter note25.2B language model with 3.8B active; 25.806B checkpoint total also includes vision components. | 25.806B3.8B active · MoE | 45.9supported | 32.6 | Not measured | 36.5 | 12.903 GBWeights only | 256K |
| 7 | Qwen3.6-35B-A3BAlibaba · Open weightsApache-2.0Official model cardParameter sourceParameter note35B language model with 3B active; 35.952B checkpoint total includes vision and auxiliary parameters. | 35.952B3B active · MoE | 43.4estimated | 34.7 | 28.7 | 41.9 | 17.976 GBWeights only | 262,144 native; up to 1M extended |
| 8 | Muse Glimmer 30BMeta · Open weightsApache-2.0Official model cardParameter sourceParameter note29.777B total checkpoint parameters; the card describes roughly 29.6B including the perception encoder. | 29.777BDense | 41.8estimated | 32.4 | 24.3 | 42.8 | 14.889 GBWeights only | 131,072+ |
| 9 | Gemma 4 31BGoogle · Open weightsApache-2.0Official model cardParameter sourceParameter note30.7B language model plus vision components; 31.273B total stored checkpoint parameters. | 31.273BDense | 40.4estimated | 33.2 | 18.9 | 38.6 | 15.637 GBWeights only | 256K |
| 10 | GLM-4.7-FlashZ.AI · Open weightsMITOfficial model cardParameter sourceParameter note31.221B total checkpoint parameters; the model card describes a 30B-A3B MoE. | 31.221B3B active · MoE | 39.4estimated | Not measured | Not measured | Not measured | 15.611 GBWeights only | 202,752 |
| 11 | GPT-OSS 20BOpenAI · Open weightsApache-2.0Official model cardParameter sourceParameter noteOpenAI reports 21B total and 3.6B active. Native MXFP4 packing means tensor-element counts do not directly equal logical parameters. | 21B3.6B active · MoE | 33.6supported | 20.1 | 14.4 | 33.0 | 10.5 GBWeights only | 131,072 |
| 12 | Granite 4.2 8BIBM · Open weightsApache-2.0Official model cardParameter sourceParameter noteTotal checkpoint parameters from the publisher repository metadata, rounded to three decimals in billions. | 8.792BDense | 31.8estimated | 18.1 | 14.0 | 28.8 | 4.396 GBWeights only | 128K native; up to 512K extended |
| 13 | Gemma 4 12BGoogle · Open weightsApache-2.0Official model cardParameter sourceParameter note11.960B total checkpoint parameters. The unified architecture handles text, images, and audio without separate encoders. | 11.96BDense | 31.8estimated | 23.2 | Not measured | 32.5 | 5.98 GBWeights only | 256K |
| 14 | ZAYA1-8BZyphra · Open weightsApache-2.0Official model cardParameter sourceParameter note8.840B stored checkpoint parameters; the model card describes 8.4B total and 0.760B active. The index uses the larger stored total. | 8.84B0.76B active · MoE | 31.5estimated | Not measured | Not measured | 33.9 | 4.42 GBWeights only | 131K |
| 15 | DeepSeek R1 Distill Qwen 32BDeepSeek · Open weightsMITOfficial model cardParameter sourceParameter noteTotal checkpoint parameters from the publisher repository metadata, rounded to three decimals in billions. | 32.764BDense | 31.2estimated | Not measured | Not measured | 28.8 | 16.382 GBWeights only | 128K |
| 16 | Gemma 4 E4BGoogle · Open weightsApache-2.0Official model cardParameter sourceParameter note7.996B total with embeddings and multimodal components. E4B names 4.5B effective parameters; that is not a published active count. | 7.996BDense | 31.2estimated | 14.2 | Not measured | 26.6 | 3.998 GBWeights only | 128K |
| 17 | Nemotron 3 Nano Omni 30B A3BNVIDIA · Open weightsNVIDIA Open Model AgreementOfficial model cardParameter sourceParameter note33.016B stored checkpoint parameters, including vision and speech encoders. The card reports approximately 3B active for the language backbone. | 33.016B3B active · MoE | 30.9estimated | 16.6 | Not measured | 31.2 | 16.508 GBWeights only | 256K |
| 18 | LFM2.5-2.6BLiquidAI · Open weightsLFM Open License v1.0Official model cardParameter sourceParameter noteTotal checkpoint parameters from the publisher repository metadata, rounded to three decimals in billions. | 2.697BHybrid | 30.6estimated | 13.3 | 15.5 | 25.3 | 1.349 GBWeights only | 131,072 |
| 19 | Gemma 4 E2BGoogle · Open weightsApache-2.0Official model cardParameter sourceParameter note5.123B total with embeddings and multimodal components. E2B names 2.3B effective parameters; that is not a published active count. | 5.123BDense | 30.1estimated | 12.6 | Not measured | 24.8 | 2.562 GBWeights only | 128K |
| 20 | LFM2.5-8B-A1BLiquidAI · Open weightsLFM Open License v1.0Official model cardParameter sourceParameter note8.468B checkpoint total; the model card reports 1.5B active despite the A1B name. | 8.468B1.5B active · MoE | 29.4estimated | Not measured | Not measured | 25.1 | 4.234 GBWeights only | 128,000 |
| 21 | Qwen2.5-72BAlibaba · Open weightsQwen LicenseOfficial model cardParameter sourceParameter noteTotal checkpoint parameters from the publisher repository metadata, rounded to three decimals in billions. | 72.706BDense | 29.2estimated | Not measured | Not measured | Not measured | 36.353 GBWeights only | 32,768 default; up to 131,072 extended |
| 22 | Sarvam 30BSarvam · Open weightsApache-2.0Official model cardParameter sourceParameter note32.153B total checkpoint parameters. The published 2.4B active count excludes embeddings, so a full active count is left blank. | 32.153BActive count not documented · MoE | 28.8estimated | Not measured | Not measured | 26.3 | 16.077 GBWeights only | 131,072 |
| 23 | Exaone 4.0 32BLG AI Research · Open weightsEXAONE Model LicenseOfficial model cardParameter sourceParameter noteTotal checkpoint parameters from the publisher repository metadata, rounded to three decimals in billions. | 32.003BDense | 28.6estimated | Not measured | Not measured | 27.0 | 16.002 GBWeights only | 131,072 |
| 24 | Exaone 4.0 1.2BLG AI Research · Open weightsEXAONE Model LicenseOfficial model cardParameter sourceParameter noteTotal checkpoint parameters from the publisher repository metadata, rounded to three decimals in billions. | 1.279BDense | 27.5estimated | Not measured | Not measured | 22.8 | 0.64 GBWeights only | 65,536 |
| 25 | Qwen2.5 Coder 32B InstructAlibaba · Open weightsApache-2.0Official model cardParameter sourceParameter noteTotal checkpoint parameters from the publisher repository metadata, rounded to three decimals in billions. | 32.764BDense | 26.8estimated | Not measured | Not measured | 24.4 | 16.382 GBWeights only | 32,768 default; up to 131,072 extended |
| 26 | Phi-4Microsoft · Open weightsMITOfficial model cardParameter sourceParameter noteTotal checkpoint parameters from the publisher repository metadata, rounded to three decimals in billions. | 14.66BDense | 25.7supported | Not measured | Not measured | 23.1 | 7.33 GBWeights only | 16K |
| 27 | Ministral 3 14BMistral · Open weightsApache-2.0Official model cardParameter sourceParameter noteTotal checkpoint parameters from the publisher repository metadata, rounded to three decimals in billions. | 13.945BDense | 25.1estimated | 12.0 | 6.4 | Not measured | 6.973 GBWeights only | 256K |
| 28 | Llama 3 70BMeta · Open weightsLlama 3 Community LicenseOfficial model cardParameter sourceParameter noteTotal checkpoint parameters from the publisher repository metadata, rounded to three decimals in billions. | 70.554BDense | 20.9estimated | Not measured | Not measured | Not measured | 35.277 GBWeights only | 8K |
| 29 | Ministral 3 8BMistral · Open weightsApache-2.0Official model cardParameter sourceParameter noteTotal checkpoint parameters from the publisher repository metadata, rounded to three decimals in billions. | 8.918BDense | 20.3estimated | 10.6 | 5.7 | Not measured | 4.459 GBWeights only | 256K |
| 30 | Ministral 3 3BMistral · Open weightsApache-2.0Official model cardParameter sourceParameter note3.849B total checkpoint parameters, including the vision encoder. The official release uses FP8 for language-model weights. | 3.849BDense | 19.1estimated | 9.0 | 5.0 | Not measured | 1.925 GBWeights only | 256K |
| 31 | Mistral 7B v0.3Mistral · Open weightsApache-2.0Official model cardParameter sourceParameter noteTotal checkpoint parameters from the publisher repository metadata, rounded to three decimals in billions. | 7.248BDense | 7.3estimated | Not measured | Not measured | Not measured | 3.624 GBWeights only | 32K |
| 32 | MiniCPM5-1BOpenBMB · Open weightsApache-2.0Official model cardParameter sourceParameter note1.081B total, including embeddings; the card separately reports 0.680B non-embedding parameters. | 1.081BDense | 4.7estimated | Not measured | Not measured | 7.6 | 0.541 GBWeights only | 131,072 |
| Not ranked | Granite 4.2 30BIBM · Open weightsApache-2.0Official model cardParameter sourceParameter noteTotal checkpoint parameters from the publisher repository metadata, rounded to three decimals in billions. | 29.277BDense | Not ranked | Not measured | Not measured | Not measured | 14.639 GBWeights only | 128K native; up to 512K extended |
| Not ranked | Granite 4.2 3BIBM · Open weightsApache-2.0Official model cardParameter sourceParameter noteTotal checkpoint parameters from the publisher repository metadata, rounded to three decimals in billions. | 3.66BDense | Not ranked | Not measured | Not measured | Not measured | 1.83 GBWeights only | 128K native; up to 512K extended |
| Not ranked | LFM2.5-230MLiquidAI · Open weightsLFM Open License v1.0Official model cardParameter sourceParameter noteTotal checkpoint parameters from the publisher repository metadata, rounded to three decimals in billions. | 0.23BHybrid | Not ranked | Not measured | Not measured | Not measured | 0.115 GBWeights only | 32,768 |
| Not ranked | LFM2.5-350MLiquidAI · Open weightsLFM Open License v1.0Official model cardParameter sourceParameter noteTotal checkpoint parameters from the publisher repository metadata, rounded to three decimals in billions. | 0.354BHybrid | Not ranked | Not measured | Not measured | Not measured | 0.177 GBWeights only | 32,768 |
| Not ranked | LongCat-Flash-Lite-SparseMeituan · Open weightsMITOfficial model cardParameter sourceParameter note69.127B stored checkpoint parameters; the card reports approximately 3B active per token. N-gram embeddings remain part of the total. | 69.127B3B active · MoE | Not ranked | Not measured | Not measured | Not measured | 34.564 GBWeights only | 1M |
| Not ranked | Mellum2-12B-A2.5B-InstructJetBrains · Open weightsApache-2.0Official model cardParameter sourceParameter noteTotal checkpoint parameters from the publisher repository metadata, rounded to three decimals in billions. | 12.15B2.5B active · MoE | Not ranked | Not measured | Not measured | Not measured | 6.075 GBWeights only | 131,072 |
| Not ranked | Mellum2-12B-A2.5B-ThinkingJetBrains · Open weightsApache-2.0Official model cardParameter sourceParameter noteTotal checkpoint parameters from the publisher repository metadata, rounded to three decimals in billions. | 12.15B2.5B active · MoE | Not ranked | Not measured | Not measured | Not measured | 6.075 GBWeights only | 131,072 |
| Not ranked | MiniCPM5-2BOpenBMB · Open weightsApache-2.0Official model cardParameter sourceParameter note2.517B total, including embeddings; the card separately reports 1.982B non-embedding parameters. | 2.517BDense | Not ranked | Not measured | Not measured | Not measured | 1.259 GBWeights only | 131,072 |
| Not ranked | Ministral 3 14B (Reasoning)Mistral · Open weightsApache-2.0Official model cardParameter sourceParameter noteTotal checkpoint parameters from the publisher repository metadata, rounded to three decimals in billions. | 13.945BDense | Not ranked | Not measured | Not measured | Not measured | 6.973 GBWeights only | 256K |
| Not ranked | Ministral 3 3B (Reasoning)Mistral · Open weightsApache-2.0Official model cardParameter sourceParameter noteRepository metadata reports 4.252B total. The card separately rounds the language model to 3.4B and the vision encoder to 0.4B. | 4.252BDense | Not ranked | Not measured | Not measured | Not measured | 2.126 GBWeights only | 256K |
| Not ranked | Ministral 3 8B (Reasoning)Mistral · Open weightsApache-2.0Official model cardParameter sourceParameter noteTotal checkpoint parameters from the publisher repository metadata, rounded to three decimals in billions. | 8.918BDense | Not ranked | Not measured | Not measured | Not measured | 4.459 GBWeights only | 256K |
| Not ranked | Mistral 8x7BMistral · Open weightsApache-2.0Official model cardParameter sourceParameter note46.703B total. The eight experts share attention and embeddings, so 8 times 7B would overstate its size. | 46.703BActive count not documented · MoE | Not ranked | Not measured | Not measured | Not measured | 23.352 GBWeights only | 32K |
| Not ranked | Nemotron 3 Nano 30BNVIDIA · Open weightsNVIDIA Nemotron Open Model LicenseOfficial model cardParameter sourceParameter note31.578B total checkpoint parameters; the card reports 3.5B active in a Mamba2-Transformer hybrid MoE. | 31.578B3.5B active · MoE | Not ranked | Not measured | Not measured | Not measured | 15.789 GBWeights only | 256K default; up to 1M |
| Not ranked | Pokee-Isaac 28BPokee AI · Closed weightsProprietaryTechnical reportParameter sourceParameter noteThe original v0 technical report lists 28B dense in Table 10. This is a provider-reported rounded size, not an exact checkpoint count. The weights are proprietary; local or VPC access does not make them publicly downloadable. | 28BDense | Not ranked | Not measured | Not measured | Not measured | Not availableClosed weights | Up to 10M (provider) |
| Not ranked | ZAYA1-74B-PreviewZyphra · Open weightsApache-2.0Official model cardParameter sourceParameter note74.794B stored checkpoint parameters and 4B active. This is a reasoning-base preview without chat tuning or RL post-training. | 74.794B4B active · MoE | Not ranked | Not measured | Not measured | Not measured | 37.397 GBWeights only | 256K |
Start with the strongest ranked model in your size band
These picks follow the current overall ranking within each band. Check a task score before treating the general leader as your coding or agent choice.
| Size band | Current overall leader | Where this pick can lose |
|---|---|---|
| Tiny: below 10B | Granite 4.2 8B31.8 overall · 8.792B total | A low parameter count does not establish reliable results on your task. |
| Small: 10–40B | Qwen3.8-27B58.2 overall · 27.781B total | The complete runtime needs more memory than the theoretical weight size. |
| Larger: above 40B, below 100B | Qwen2.5-72B29.2 overall · 72.706B total | These are larger deployments; inclusion does not establish laptop fit. |
A size limit gives you a shortlist
Small language model, or SLM, has no single parameter cutoff. We use three bands here: tiny below 10B, small from 10B through 40B, and larger models above 40B but below 100B. They let you compare models within a size budget without implying that everything on this page fits a laptop.
Start in the smallest band your application can use, and compare its answers with a larger candidate on the same prompts. For a short extraction task, the smaller checkpoint can be enough even when its general score is lower. For a coding agent that must recover from errors across many steps, a broader capability gap can matter more.
The starting points above are the highest overall ranked checkpoints in each band of this source-checked index. They are benchmark candidates to evaluate, not winners from a local hardware test. A tiny-band leader can still lose when your task needs knowledge, document understanding, or reliable tool use that its score does not establish.
Total parameters decide the 100B cutoff
A dense model uses its main network for each token. A mixture-of-experts model, or MoE, selects part of its expert network for a token. That produces two useful numbers: total parameters describe the full checkpoint, while active parameters describe the portion used for a token.
We apply the strict cutoff to total parameters. A 120B MoE with 12B active parameters is outside this index. Its active count does not turn it into a 12B download or establish a 12B memory requirement. The runtime still has to store and manage the other experts, whether they reside on a GPU, in system RAM, or elsewhere.
Read the parameter note beside the source link when a name uses an effective count, a rounded size, or a language-only count. Embeddings and multimodal components can make the full checkpoint larger than the number in its name. The documented count and its scope are more useful than the marketing label.
Weight access and license terms answer different questions
Use the weight-access filter to separate downloadable weights from proprietary weights. A provider can offer a closed-weight model through an API, a private deployment, or a VPC without making its parameters publicly downloadable. Closed-weight models enter this index only when a publisher documents a model size below the 100B limit.
The license filter separates MIT or Apache 2.0 from community or custom terms. This describes the recorded license, not the openness of the entire AI system. Downloadable weights and a permissive license do not, by themselves, establish access to the training code and data information needed to reproduce the system. Check the publisher terms for the use you intend.
Pokee-Isaac 28B uses the original v0 report's rounded 28B dense size. Its size is provider-reported rather than counted from a public checkpoint. The index does not estimate downloadable weight storage for closed-weight models, and their private-deployment availability does not establish public access.
See the Open Source AI Definition for requirements beyond weight access.
Four-bit weight size is only a starting estimate
At exactly four bits per parameter, weight storage is total parameters multiplied by 0.5GB per billion parameters, using decimal GB. That gives 4GB for 8B, 13.5GB for 27B, and 35GB for 70B. The index computes this arithmetic for each downloadable checkpoint; these are theoretical weights-only figures.
A real quantized artifact also contains metadata and tensors that may use other precision. Runtime buffers, KV cache, context length, parallel requests, and any extra encoders add memory. Some checkpoint formats use mixed precision. The estimate therefore is not a measured file size, a minimum RAM specification, or a guarantee that a model fits a GPU.
Reserve memory for the operating system and your other applications on a shared-memory machine. On a discrete GPU, check VRAM separately from system RAM. CPU offload can make a configuration possible, but it changes the speed you need to measure. Increasing context or concurrency can break a configuration that worked for one short prompt.
For a configuration estimate, use the VRAM calculator, then verify the result in your runtime.
Benchmark scores compare capability, not local speed
The overall and task scores come from the same current public ranking used by the main leaderboards. We filter that ranking by the source-checked parameter registry and retain its order. We do not rescore a checkpoint for being smaller, borrow a sibling model's results, or replace missing scores with zero.
Supported means the public rank has enough independent evidence to stand behind it. Estimated means the position is visible with greater uncertainty. Not ranked means this checkpoint has no published overall position in the current ranking. A missing task score remains marked Not measured, even when the checkpoint has an overall score.
Published evaluations can use a different reasoning budget, precision, harness, or runtime from your local deployment. A score does not measure quantized answer quality, tokens per second, time to first token, or performance at the context length you plan to serve. Model pages carry the individual benchmark results and their sources so you can inspect what a score rests on.
This is a maintained index of documented checkpoints in our catalog, not every language model ever released. Models without a checked total count stay outside the index. Open weights also do not imply identical license terms: read the official model card and license before commercial use, redistribution, or fine-tuning.
Your own prompts decide which model earns the memory
Choose two candidates within your memory budget and one larger reference if you can run it. Keep the prompt, output limit, context, and scoring criteria the same. Record the exact checkpoint, quantization, runtime version, hardware, and reasoning settings so a later change can be compared with the first run.
Use examples from the work you actually need done. For extraction, check exact fields and unsupported values. For coding, run the tests and inspect the patch. For retrieval, include questions the supplied documents cannot answer and check whether the model invents an answer. Count successful tasks as well as retries; a fast failed answer does not finish the job.
Measure peak memory and latency with the context and concurrency you intend to use. Move to a larger checkpoint when it solves failures that matter enough to justify the extra memory or waiting time. Keep the smaller one when it meets the same acceptance criteria on your workload.
Questions
Are all models under 100B small language models?
No universal cutoff defines a small language model. This page uses under 100B total parameters as an explicit inclusion limit and separates tiny, small, and larger checkpoints. A 70B model can qualify for the index while requiring much more memory than the compact models usually described as SLMs.
Can I run an under-100B model on a laptop?
The cutoff alone cannot establish laptop fit. Available memory, checkpoint format, quantization, context, and runtime support decide whether a configuration works. The four-bit column estimates only theoretical weight storage. Check the exact artifact, leave room for the operating system and KV cache, and measure the complete configuration before relying on it.
Does a lower active parameter count mean less RAM?
An active parameter count describes the part of an MoE used for a token. It does not describe the entire checkpoint. Inactive experts still need storage and memory management. Total parameters are the safer starting point for weight-size estimates, and the runtime determines how experts are loaded or offloaded.
Are smaller language models always faster?
Parameter count alone does not prove a speed advantage on your machine. Hardware bandwidth, quantization kernels, attention, context length, reasoning output, and CPU offload can change latency. Compare the same workload on the same runtime and hardware, and record time to a usable answer as well as token generation speed.