Small language model evaluation

Research index

Compact
Language Model
Index

A fixed evaluation of compact language models across six benchmarks, with chance-normalized scores and 95% intervals

Explore the leaderboard
Evaluation breadth
6core benchmarks
Linguistic depth
12BLiMP categories
Current sample
18models measured

Ranked results

Core leaderboard

RankTrack
1–2LFM2 350M LiquidAI/LFM2-350M354Minstruction43.242.0–44.232.030.6–33.339.235.0–43.255.252.6–57.719.515.7–23.331.327.9–34.958.157.2–58.4
1–2Qwen2.5 0.5B Qwen/Qwen2.5-0.5B494Mbase43.041.9–44.036.234.8–37.538.834.7–42.844.641.9–47.29.35.8–13.126.222.9–29.768.767.9–68.9
3Gemma 3 270M google/gemma-3-270m270Mbase38.937.8–39.921.920.7–23.337.032.9–41.243.040.4–45.64.00.6–7.328.224.8–31.764.363.5–64.5
4SmolLM2 135M HuggingFaceTB/SmolLM2-135M135Mbase37.836.7–38.824.323.0–25.636.332.1–40.644.942.3–47.56.32.6–9.719.416.2–22.862.761.9–63.1
5GPT-X2 125M AxiomicLabs/GPT-X2-125M125Mbase35.534.3–36.520.719.4–21.934.229.7–38.435.532.8–38.23.4-0.1–7.118.314.9–21.665.865.0–66.1
6MobileLLM-R1 140M Base facebook/MobileLLM-R1-140M-base140Mbase31.230.1–32.212.110.9–13.326.622.2–30.833.230.5–35.7-1.1-4.6–2.319.316.0–22.861.660.7–61.8
7Baguettotron PleIAs/Baguettotron321Minstruction29.628.4–30.713.912.6–15.124.219.6–28.734.131.5–36.77.23.8–10.613.310.0–16.656.956.0–57.2
8–9GPT-2 124M openai-community/gpt2124Mbase27.926.8–28.98.27.0–9.425.020.7–29.419.316.7–21.9-3.4-6.6–-0.212.39.2–15.567.566.8–67.8
8–9Supra 50M Base SupraLabs/Supra-50M-Base51.8Mbase27.326.2–28.38.87.6–10.023.418.9–28.027.524.7–30.2-0.1-3.5–3.212.59.2–15.759.057.9–59.1
10Veyra2 Apricot 50M Base veyra-ai/Veyra2-Apricot-50M-Base49.3Mbase26.124.9–27.18.47.2–9.523.419.0–27.923.720.9–26.4-2.0-5.2–1.110.87.7–14.058.757.7–58.9
11–13Falcon H1 Tiny R 90M tiiuae/Falcon-H1-Tiny-R-90M90.0Minstruction20.819.6–21.87.76.6–8.921.517.0–25.919.016.3–21.5-0.1-3.4–3.210.47.4–13.642.541.6–43.0
11–13SLM 10M liodon-ai/slm-10m9.97Mbase20.319.1–21.23.12.0–4.214.19.7–18.514.311.6–16.8-1.9-5.1–1.37.54.5–10.654.453.5–54.7
11–14Pythia 160M EleutherAI/pythia-160m160Mbase20.219.1–21.27.36.2–8.516.912.5–21.215.813.2–18.4-2.6-5.7–0.611.78.4–14.845.744.8–46.2
13–14GPT-S2 5M AxiomicLabs/GPT-S2-5M5.38Mbase19.318.2–20.33.42.3–4.512.98.3–17.411.38.6–13.9-3.8-6.9–-0.76.03.2–9.254.653.9–55.0
15–16nanowhale 100M Base HuggingFaceTB/nanowhale-100m-base110Mbase17.916.8–18.92.81.6–4.014.710.2–19.316.413.7–19.0-3.4-6.5–-0.29.76.6–12.941.841.0–42.3
15–16Pythia 31M EleutherAI/pythia-31m31.0Mbase16.915.8–18.02.91.8–4.113.69.1–18.112.19.7–14.5-4.4-7.6–-1.18.15.1–11.143.342.5–43.9
17Atom 3.4M UniversalComputingResearch/Atom3.4m3.41Mbase14.313.3–15.33.52.4–4.511.46.9–15.910.88.4–13.2-4.2-7.4–-1.15.92.9–9.236.435.6–37.0
18GPT-S 1.4M AxiomicLabs/GPT-S-1.4M1.43Mbase12.311.2–13.32.51.3–3.610.35.7–14.79.06.5–11.5-4.0-7.1–-0.74.31.5–7.432.031.2–32.6

ARC-E* = ARC-Easy · ARC-C* = ARC-Challenge · CQA* = CommonsenseQA

Chart metric

Select a benchmark to update both the scaling chart and score comparison below.

Model scaling

Parameters and Average score

Base Instruction-tuned Pareto trend
Chance-normalized score-5.08.021.034.047.01.00M66.2M233M502MParameters (square-root scale)

Model score comparison

The top 10 matching models are shown by default. This comparison uses the shared metric selected above.

1 LFM2 350M
43.2
2 Qwen2.5 0.5B
43.0
3 Gemma 3 270M
38.9
4 SmolLM2 135M
37.8
5 GPT-X2 125M
35.5
6 MobileLLM-R1 140M Base
31.2
7 Baguettotron
29.6
8 GPT-2 124M
27.9
9 Supra 50M Base
27.3
10 Veyra2 Apricot 50M Base
26.1

Scoring method

Chance-normalized composite score

100 × (accuracy − chance) / (1 − chance)

For each benchmark, the observed result is normalized against its chance level: chance maps to 0 and a perfect result maps to 100.

HellaSwag, PIQA, and CommonsenseQA form a commonsense domain worth 50% of the composite, with each benchmark weighted equally within that domain. Science and world knowledge contribute 25% after combining ARC-Easy (75%) and ARC-Challenge (25%). BLiMP contributes the remaining 25% for grammatical competence and is macro-averaged across its 12 linguistic categories.

CommonsenseQA scores the likelihood of each full answer text with length normalization, rather than scoring the A–E answer label. This prevents a model’s letter preference from being mistaken for commonsense capability.

Each benchmark interval is a 95% percentile bootstrap. Model comparisons use paired bootstrap resampling on matched evaluation items. Results below chance remain negative, and rank ranges reflect comparisons that do not distinguish neighbouring models.

Evaluation protocol

Curated evaluation

CompactLMIndex is a curated benchmark, not an open intake list. We are establishing a Pareto-frontier baseline across model sizes and highlighting training or inference strategies that put a model ahead of its peers at a comparable parameter count.

  • Complete core suite required for a rank
  • Zero-shot, non-chat log-likelihood prompts
  • Paired bootstrap comparisons on matched questions

Every published result is independently checked through our internal verification process. Final inclusion and presentation decisions remain with Universal Computing Research.

If you have interesting results for a model you think is suitable for this benchmark, you can open a discussion on the Space for consideration.

Credits & scope

Efficiency at smaller scales

CompactLMIndex looks at smaller language models. At this scale, architecture, training, and inference choices can make a real difference. Every model follows the same evaluation protocol, so results are compared fairly rather than by size alone.

Inspired by the Open LLM Leaderboard, Axiomic Labs Open SLM Leaderboard, and OpenAI Parameter Golf.