Small language model evaluation
Research indexCompact
Language Model
Index
A fixed evaluation of compact language models across six benchmarks, with chance-normalized scores and 95% intervals
Explore the leaderboardRanked results
Core leaderboard
| Rank | Release date | Vocab. size | Track | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 1–2 | LFM2 350M LiquidAI/LFM2-350M | 2025-07-10 | 354M | 65,536 | instruction | 43.242.0–44.2 | 32.030.6–33.3 | 39.235.0–43.2 | 55.252.6–57.7 | 19.515.7–23.3 | 31.327.9–34.9 | 58.157.2–58.4 |
| 1–2 | Qwen2.5 0.5B Qwen/Qwen2.5-0.5B | 2024-09-15 | 494M | 151,936 | base | 43.041.9–44.0 | 36.234.8–37.5 | 38.834.7–42.8 | 44.641.9–47.2 | 9.35.8–13.1 | 26.222.9–29.7 | 68.767.9–68.9 |
| 3 | Gemma 3 270M google/gemma-3-270m | 2025-08-05 | 270M | 262,144 | base | 38.937.8–39.9 | 21.920.7–23.3 | 37.032.9–41.2 | 43.040.4–45.6 | 4.00.6–7.3 | 28.224.8–31.7 | 64.363.5–64.5 |
| 4 | SmolLM2 135M HuggingFaceTB/SmolLM2-135M | 2024-10-31 | 135M | 49,152 | base | 37.836.7–38.8 | 24.323.0–25.6 | 36.332.1–40.6 | 44.942.3–47.5 | 6.32.6–9.7 | 19.416.2–22.8 | 62.761.9–63.1 |
| 5 | GPT-X2 125M AxiomicLabs/GPT-X2-125M | 2026-03-25 | 125M | 32,768 | base | 35.534.3–36.5 | 20.719.4–21.9 | 34.229.7–38.4 | 35.532.8–38.2 | 3.4-0.1–7.1 | 18.314.9–21.6 | 65.865.0–66.1 |
| 6 | MobileLLM-R1 140M Base facebook/MobileLLM-R1-140M-base | 2025-09-10 | 140M | 128,256 | base | 31.230.1–32.2 | 12.110.9–13.3 | 26.622.2–30.8 | 33.230.5–35.7 | -1.1-4.6–2.3 | 19.316.0–22.8 | 61.660.7–61.8 |
| 7 | Baguettotron PleIAs/Baguettotron | 2025-11-10 | 321M | 65,536 | instruction | 29.628.4–30.7 | 13.912.6–15.1 | 24.219.6–28.7 | 34.131.5–36.7 | 7.23.8–10.6 | 13.310.0–16.6 | 56.956.0–57.2 |
| 8–9 | GPT-2 124M openai-community/gpt2 | 2022-03-02 | 124M | 50,257 | base | 27.926.8–28.9 | 8.27.0–9.4 | 25.020.7–29.4 | 19.316.7–21.9 | -3.4-6.6–-0.2 | 12.39.2–15.5 | 67.566.8–67.8 |
| 8–9 | Supra 50M Base SupraLabs/Supra-50M-Base | 2026-05-21 | 51.8M | 32,000 | base | 27.326.2–28.3 | 8.87.6–10.0 | 23.418.9–28.0 | 27.524.7–30.2 | -0.1-3.5–3.2 | 12.59.2–15.7 | 59.057.9–59.1 |
| 10 | Veyra2 Apricot 50M Base veyra-ai/Veyra2-Apricot-50M-Base | 2026-06-14 | 49.3M | 8,192 | base | 26.124.9–27.1 | 8.47.2–9.5 | 23.419.0–27.9 | 23.720.9–26.4 | -2.0-5.2–1.1 | 10.87.7–14.0 | 58.757.7–58.9 |
| 11–13 | Falcon H1 Tiny R 90M tiiuae/Falcon-H1-Tiny-R-90M | 2026-01-12 | 90.0M | 32,768 | instruction | 20.819.6–21.8 | 7.76.6–8.9 | 21.517.0–25.9 | 19.016.3–21.5 | -0.1-3.4–3.2 | 10.47.4–13.6 | 42.541.6–43.0 |
| 11–13 | SLM 10M liodon-ai/slm-10m | 2026-06-13 | 9.97M | 8,192 | base | 20.319.1–21.2 | 3.12.0–4.2 | 14.19.7–18.5 | 14.311.6–16.8 | -1.9-5.1–1.3 | 7.54.5–10.6 | 54.453.5–54.7 |
| 11–14 | Pythia 160M EleutherAI/pythia-160m | 2023-02-08 | 160M | 50,304 | base | 20.219.1–21.2 | 7.36.2–8.5 | 16.912.5–21.2 | 15.813.2–18.4 | -2.6-5.7–0.6 | 11.78.4–14.8 | 45.744.8–46.2 |
| 13–14 | GPT-S2 5M AxiomicLabs/GPT-S2-5M | 2026-06-19 | 5.38M | 4,096 | base | 19.318.2–20.3 | 3.42.3–4.5 | 12.98.3–17.4 | 11.38.6–13.9 | -3.8-6.9–-0.7 | 6.03.2–9.2 | 54.653.9–55.0 |
| 15–16 | nanowhale 100M Base HuggingFaceTB/nanowhale-100m-base | 2026-04-24 | 110M | 129,280 | base | 17.916.8–18.9 | 2.81.6–4.0 | 14.710.2–19.3 | 16.413.7–19.0 | -3.4-6.5–-0.2 | 9.76.6–12.9 | 41.841.0–42.3 |
| 15–16 | Pythia 31M EleutherAI/pythia-31m | 2026-02-24 | 31.0M | 50,304 | base | 16.915.8–18.0 | 2.91.8–4.1 | 13.69.1–18.1 | 12.19.7–14.5 | -4.4-7.6–-1.1 | 8.15.1–11.1 | 43.342.5–43.9 |
| 17 | Atom 3.4M UniversalComputingResearch/Atom3.4m | 2026-06-19 | 3.41M | 4,096 | base | 14.313.3–15.3 | 3.52.4–4.5 | 11.46.9–15.9 | 10.88.4–13.2 | -4.2-7.4–-1.1 | 5.92.9–9.2 | 36.435.6–37.0 |
| 18 | GPT-S 1.4M AxiomicLabs/GPT-S-1.4M | 2026-06-01 | 1.43M | 4,096 | base | 12.311.2–13.3 | 2.51.3–3.6 | 10.35.7–14.7 | 9.06.5–11.5 | -4.0-7.1–-0.7 | 4.31.5–7.4 | 32.031.2–32.6 |
ARC-E* = ARC-Easy · ARC-C* = ARC-Challenge · CQA* = CommonsenseQA
Chart metric
Select a benchmark to update both the scaling chart and score comparison below.
Model scaling
Parameters and Average score
Model score comparison
The top 10 matching models are shown by default. This comparison uses the shared metric selected above.
Scoring method
Chance-normalized composite score
100 × (accuracy − chance) / (1 − chance)
For each benchmark, the observed result is normalized against its chance level: chance maps to 0 and a perfect result maps to 100.
HellaSwag, PIQA, and CommonsenseQA form a commonsense domain worth 50% of the composite, with each benchmark weighted equally within that domain. Science and world knowledge contribute 25% after combining ARC-Easy (75%) and ARC-Challenge (25%). BLiMP contributes the remaining 25% for grammatical competence and is macro-averaged across its 12 linguistic categories.
CommonsenseQA scores the likelihood of each full answer text with length normalization, rather than scoring the A–E answer label. This prevents a model’s letter preference from being mistaken for commonsense capability.
Each benchmark interval is a 95% percentile bootstrap. Model comparisons use paired bootstrap resampling on matched evaluation items. Results below chance remain negative, and rank ranges reflect comparisons that do not distinguish neighbouring models.
Evaluation protocol
Curated evaluation
CompactLMIndex is a curated benchmark, not an open intake list. We are establishing a Pareto-frontier baseline across model sizes and highlighting training or inference strategies that put a model ahead of its peers at a comparable parameter count.
- Complete core suite required for a rank
- Zero-shot, non-chat log-likelihood prompts
- Paired bootstrap comparisons on matched questions
Every published result is independently checked through our internal verification process. Final inclusion and presentation decisions remain with Universal Computing Research.
If you have interesting results for a model you think is suitable for this benchmark, you can open a discussion on the Space for consideration.
Credits & scope
Efficiency at smaller scales
CompactLMIndex looks at smaller language models. At this scale, architecture, training, and inference choices can make a real difference. Every model follows the same evaluation protocol, so results are compared fairly rather than by size alone.
Inspired by the Open LLM Leaderboard, Axiomic Labs Open SLM Leaderboard, and OpenAI Parameter Golf.