All 49 Benchmarked AI Models — Complete Results
IWV Digital Solutions LLC | 75-model benchmark sweep | refreshed 12 August 2026
Reference data — full candidate sweep. This page tabulates the complete NotesXML model-selection sweep: all 76 benchmarked candidate models on Ryzen 7 7800X3D + RTX 5070 Ti 16 GB, Linux, GPU-accelerated inference, identical sampling parameters per model, sorted by HMS (descending). ★ marks the models in the current 10-model shipping catalog (v2026.08.13.02) — all 10 appear in this sweep. The 13 August 2026 revision (v2026.08.13.02) replaced LFM2 VL 3B with LFM2.5 VL 3B; the 12 August revision had added GPT-OSS 20B and retired Llama 3.2 3B Instruct and Phi-4 Mini 3.8B; retired models remain listed here as evaluated candidates. All 10 shipping models sit on the sweep’s Pareto frontier. Sorting is by HMS rather than raw quality, so fast small models are not buried beneath large slow ones.
Refreshed 13 August 2026 — catalog v2026.08.13.02, tool-calling suite updated, LFM2.5 VL 3B added. The scatter plot and the full results table on this page are regenerated from the current master benchmark catalog: 76 evaluated models, of which the 10 in the shipping NotesXML catalog are highlighted. LFM2.5 VL 1.6B (Liquid AI) enters as the recommended default on both platforms and a free-tier model — 584.8 TPS at 87% quality, vision-capable in 1.3 GB — and Gemma 4 31B joins the catalog; Llama 3.2 1B Instruct and Gemma 4 E2B were retired from the catalog and remain in the table below as evaluated candidates. The current catalog is on the AI Models page.
Performance Scatter Plot
Each point represents one of the 76 benchmarked models (refreshed 13 August 2026). The X-axis shows the token generation interval (1/TPS in milliseconds, log scale — lower is faster). The Y-axis shows the average quality score across the seven scored benchmarks (Code excluded). Models in the current 10-model catalog are highlighted in green (all 10 appear in this sweep).
Complete Benchmark Table
All 76 evaluated models, sorted by HMS descending. ★ = in the v2026.08.13.02 shipping catalog (10 models) (v2026.08.13.02). A dash means the benchmark does not apply to that model.
| Model | Params (B) | Vision | TPS | ms/tok | Model Bench (text) | HEE | PLE | PhD | SDB | LongCtx | Tool | Quality | HMS |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| ★ Gemma 4 26B-A4B | 26 | Yes | 51.4 | 19.5 | 100% | 100% | 99% | 100% | 80% | 85% | 94% | 100% | 1.00 |
| ★ Gemma 4 12B | 12 | Yes | 70.7 | 14.1 | 100% | 98% | 96% | 99% | 54% | 81% | 93% | 99% | 0.99 |
| ★ Ministral 3 14B | 14 | Yes | 74.6 | 13.4 | 100% | 98% | 95% | 96% | 47% | 87% | 78% | 98% | 0.99 |
| ★ Ministral 3 8B | 8 | Yes | 116 | 8.6 | 100% | 96% | 95% | 96% | 72% | 86% | 83% | 97% | 0.99 |
| ★ Gemma 4 E4B | 4.5 | Yes | 118.8 | 8.4 | 100% | 96% | 92% | 99% | 46% | 77% | 97% | 97% | 0.99 |
| Phi-4 14B | 14 | — | 77.1 | 13.0 | 100% | 96% | 95% | 98% | 56% | 56% | 91% | 97% | 0.99 |
| Phi-4 Reasoning Vision 15B | 15 | Yes | 73.1 | 13.7 | 100% | 96% | 95% | 97% | 53% | 15% | 87% | 97% | 0.98 |
| LFM 24B-A2B (MOE) | 24 | — | 81.1 | 12.3 | 100% | 96% | 87% | 98% | 32% | 14% | 0% | 95% | 0.98 |
| Mistral Small 3.2 24B | 24 | Yes | 48 | 20.8 | 97% | 96% | 96% | 100% | 37% | 75% | 94% | 97% | 0.97 |
| Devstral Small 2 24B Instruct | 24 | — | 48.4 | 20.7 | 100% | 94% | 93% | 100% | — | 69% | 94% | 97% | 0.97 |
| ★ GPT-OSS 20B | 21 | — | 141.5 | 7.1 | 89% | 100% | 93% | 96% | 69% | 82% | 4% | 94% | 0.97 |
| ★ Ministral 3 3B Instruct | 3 | Yes | 225.1 | 4.4 | 100% | 88% | 89% | 94% | 41% | 88% | 84% | 94% | 0.97 |
| Gemma 3 12B | 12 | Yes | 66.4 | 15.1 | 100% | 92% | 91% | 93% | — | 82% | 93% | 94% | 0.97 |
| Granite 4.1 3B | 3 | — | 203.8 | 4.9 | 89% | 94% | 89% | 96% | 40% | 81% | 0% | 92% | 0.96 |
| Gemma 4 E2B | 2.3 | Yes | 206.8 | 4.8 | 96% | 91% | 87% | 93% | 23% | 80% | 71% | 92% | 0.96 |
| Granite 4.1 3B Instruct | 3 | — | 211 | 4.7 | 89% | 94% | 89% | 95% | 40% | 81% | 60% | 92% | 0.96 |
| Pixtral 12B | 12 | Yes | 84.8 | 11.8 | 94% | 82% | 85% | 96% | 26% | 89% | 86% | 91% | 0.96 |
| Mistral NeMo 12B | 12 | — | 83.6 | 12.0 | 100% | 88% | 83% | 94% | 29% | 80% | 91% | 91% | 0.95 |
| Granite 4.1 8B Instruct | 8 | — | 127 | 7.9 | 86% | 94% | 89% | 95% | 29% | 84% | 92% | 91% | 0.95 |
| OLMo-2-13B | 13 | — | 73.2 | 13.7 | 96% | 84% | 85% | 97% | — | 15% | 84% | 90% | 0.95 |
| Llama 3.2 8B Instruct | 8 | — | 132 | 7.6 | 100% | 86% | 85% | 89% | 20% | 81% | 84% | 90% | 0.95 |
| ★ LFM2.5 VL 3B | 3 | Yes | 279.3 | 3.6 | 92% | 88% | 83% | 89% | 41% | 30% | 0% | 90% | 0.95 |
| LFM2 VL 3B | 3 | Yes | 270.3 | 3.7 | 94% | 84% | 79% | 93% | 46% | 31% | 0% | 90% | 0.95 |
| Granite 4.0 H-Tiny (MOE) | 7 | — | 235 | 4.3 | 89% | 90% | 86% | 90% | 38% | 79% | 84% | 89% | 0.94 |
| Granite 3.3 8B Instruct | 8 | — | 106.6 | 9.4 | 83% | 90% | 84% | 90% | 27% | 0% | 76% | 87% | 0.93 |
| OLMo-2-1124 7B Instruct | 7 | — | 137 | 7.3 | 100% | 80% | 76% | 91% | — | 16% | 65% | 87% | 0.93 |
| ★ LFM2.5 VL 1.6B | 1.6 | Yes | 584.8 | 1.7 | 92% | 78% | 74% | 90% | 28% | 79% | 4% | 87% | 0.93 |
| Gemma 3n E4B | 4 | — | 84.3 | 11.9 | 71% | 96% | 84% | 96% | — | 45% | 75% | 87% | 0.93 |
| Phi-4 Mini 3.8B | 3.8 | — | 229 | 4.4 | 92% | 86% | 81% | 86% | 48% | 84% | 82% | 86% | 0.93 |
| Devstral Small 2507 24B | 24 | — | 46.7 | 21.4 | 71% | 98% | 94% | 99% | — | 51% | 91% | 91% | 0.92 |
| Mistral 7B Instruct v0.3 | 7 | — | 120 | 8.3 | 89% | 84% | 81% | 88% | — | 0% | 76% | 86% | 0.92 |
| OLMo-2-Specialized 7B | 7 | — | 127 | 7.9 | 76% | 84% | 77% | 90% | — | 16% | 65% | 82% | 0.90 |
| SmolLM3 3B (3B) | 3 | — | 248.2 | 4.0 | 81% | 84% | 75% | 83% | 8% | 19% | 55% | 81% | 0.89 |
| Gemma 3 4B | 4 | Yes | 111.9 | 8.9 | 65% | 90% | 80% | 86% | — | 82% | 31% | 80% | 0.89 |
| Granite 3.0 2B | 2 | — | 242.2 | 4.1 | 81% | 76% | 76% | 87% | 25% | 72% | 77% | 80% | 0.89 |
| Gemma 3n E2B | 2 | — | 112.1 | 8.9 | 66% | 86% | 83% | 84% | — | 44% | 67% | 80% | 0.89 |
| Llama 3.2 3B Instruct | 3.2 | — | 257 | 3.9 | 94% | 80% | 71% | 73% | 32% | 77% | 64% | 80% | 0.89 |
| LFM 2.5 1.2B Thinking | 1.2 | — | 644 | 1.6 | 67% | 86% | 81% | 84% | 41% | 14% | 0% | 79% | 0.89 |
| Granite 3.3 2B | 2.5 | — | 231 | 4.3 | 78% | 82% | 74% | 82% | 40% | 0% | 81% | 79% | 0.88 |
| Granite 4.0 H-1B (Hybrid) | 1 | — | 265 | 3.8 | 100% | 78% | 68% | 69% | 39% | 18% | 71% | 79% | 0.88 |
| Granite 4.1 4B | 4 | Yes | 182.5 | 5.5 | 83% | 94% | 87% | 92% | 32% | 88% | 81% | 76% | 0.87 |
| SmolVLM (2B) | 2 | Yes | 364 | 2.7 | 79% | 74% | 72% | 65% | 29% | 29% | 30% | 72% | 0.84 |
| Index 1.9B Chat | 1.9 | — | 243.4 | 4.1 | 75% | 72% | 66% | 71% | 29% | 0% | 57% | 71% | 0.83 |
| Llama 3.2 1B Instruct | 1.2 | — | 559 | 1.8 | 88% | 60% | 63% | 61% | 22% | 80% | 55% | 68% | 0.81 |
| SmolLM2 1.7B Instruct | 1.7 | — | 312 | 3.2 | 72% | 64% | 65% | 63% | 28% | 32% | 82% | 66% | 0.80 |
| LFM 2.5 1.2B Instruct | 1.2 | — | 517 | 1.9 | 88% | 56% | 52% | 68% | 25% | 18% | 57% | 66% | 0.79 |
| LFM 2.5 VL 450M | 0.45 | Yes | 724.7 | 1.4 | 79% | 48% | 48% | 48% | 24% | 25% | 55% | 65% | 0.79 |
| Granite 3.0 8B | 8 | — | 105.8 | 9.5 | 85% | 45% | 17% | 87% | 34% | 0% | 86% | 58% | 0.74 |
| Granite Vision 3.2 2B | 2 | Yes | 253.8 | 3.9 | 74% | 80% | 71% | 0% | 43% | 0% | 77% | 56% | 0.72 |
| MythoMax L2 13B | 13 | — | 69.2 | 14.5 | 79% | 70% | 68% | 0% | — | 0% | 0% | 54% | 0.70 |
| LFM 2.5 350M | 0.35 | — | 675 | 1.5 | 79% | 32% | 35% | 54% | 13% | 16% | 8% | 50% | 0.67 |
| Gemma 3 1B | 1 | — | 152.4 | 6.6 | 76% | 44% | 40% | 36% | 10% | 16% | 37% | 49% | 0.66 |
| BioMistral 7B | 7 | — | 95.7 | 10.4 | 58% | 74% | 63% | 0% | — | 0% | 0% | 49% | 0.66 |
| Granite 3.0 1B-A400M | 1 | — | 444 | 2.3 | 100% | 24% | 25% | 36% | 19% | 0% | 40% | 46% | 0.63 |
| Qwen3.5 9B | 9 | Yes | 106.6 | 9.4 | 100% | 10% | 28% | 17% | 30% | 49% | 97% | 39% | 0.56 |
| Qwen3.5 35B-A3B (MoE) | 35 | Yes | 37.3 | 26.8 | 100% | 10% | 16% | 18% | 30% | 47% | 91% | 44% | 0.55 |
| Granite 4.0 H-Small | 32 | — | 17.3 | 57.8 | 89% | 96% | 93% | 99% | 52% | 83% | 94% | 94% | 0.51 |
| SmolLM2 360M Instruct | 0.36 | — | 486.9 | 2.1 | 42% | 32% | 38% | 27% | 25% | 16% | 30% | 35% | 0.51 |
| Qwen3.5 4B | 4 | Yes | 158.1 | 6.3 | 61% | 10% | 16% | 79% | 30% | 29% | 95% | 33% | 0.50 |
| Qwen3.5 0.8B | 0.8 | Yes | 364.2 | 2.7 | 46% | 10% | 16% | 16% | 30% | 29% | 86% | 33% | 0.49 |
| LFM2.5 2.6B | 2.6 | — | 301 | 3.3 | 90% | 6% | 13% | 6% | 30% | 14% | 8% | 29% | 0.45 |
| LFM 2.5 230M | 0.23 | — | 724 | 1.4 | 67% | 10% | 16% | 19% | 30% | 18% | 0% | 28% | 0.44 |
| H2O-Danube 3 500M | 0.5 | — | 634.4 | 1.6 | 75% | 6% | 16% | 14% | 30% | 0% | 0% | 28% | 0.43 |
| SmolLM2 135M Instruct | 0.135 | — | 960.1 | 1.0 | 42% | 16% | 32% | 20% | 28% | 15% | 7% | 28% | 0.43 |
| MiniCPM5 1B | 1 | — | 583.6 | 1.7 | 69% | 10% | 12% | 17% | 30% | 29% | 29% | 27% | 0.43 |
| LFM2.5 8B-A1B (MOE) | 8 | — | 55.9 | 17.9 | 69% | 6% | 15% | 14% | 24% | 28% | 0% | 26% | 0.41 |
| Gemma 3 27B | 27 | — | 12.7 | 78.7 | 61% | 98% | 93% | 97% | — | 0% | 94% | 87% | 0.39 |
| Codestral 22B | 22 | — | 12 | 83.3 | 94% | 80% | 68% | 83% | — | 0% | 72% | 81% | 0.37 |
| Qwen3.5 27B | 27 | Yes | 14.4 | 69.4 | 100% | 10% | 16% | 18% | 30% | 49% | 97% | 38% | 0.33 |
| MiniCPM-V 4.6 | 0.8 | Yes | 461 | 2.2 | 69% | 4% | 9% | 18% | 30% | 27% | 29% | 20% | 0.33 |
| Granite 4.1 30 | 30 | — | 8.9 | 112.4 | 83% | 96% | 94% | 100% | 45% | 90% | 86% | 93% | 0.30 |
| ★ Gemma 4 31B | 31 | Yes | 8.5 | 117.6 | 100% | 100% | 100% | 100% | 51% | 92% | 91% | 100% | 0.29 |
| MobileLLM 1.5B (R1.5) | 1.5 | — | 166 | 6.0 | 21% | 8% | 17% | 13% | — | 0% | 0% | 15% | 0.26 |
| Qwen3.5 2B | 2 | Yes | 284.7 | 3.5 | 13% | 10% | 16% | 18% | 30% | 29% | 44% | 11% | 0.20 |
| Granite 4.0 1B (Dense) | 1 | — | 140.3 | 7.1 | 35% | 0% | 0% | 0% | 1% | 10% | 0% | 9% | 0.16 |
| Granite 4.0 350M (Dense) | 0.35 | — | 1502 | 0.7 | 24% | 0% | 0% | 0% | 0% | 2% | 0% | 6% | 0.11 |
TPS = tokens per second on the benchmark hardware. ms/tok = 1000/TPS (token generation interval). Model Bench = Model Benchmark score (in-app task suite: six text tasks + one vision OCR, graded A–F). HEE, PLE, PhD, SDB, LongCtx, Tool = normalized scores (0–100%). Avg Score = unweighted mean of the seven normalized benchmark scores (Code excluded). HMS = 2 × min(TPS/200, 1) × AvgScore / (min(TPS/200, 1) + AvgScore).
Benchmark Descriptions
Key Takeaways
Gemma 4 26B-A4B (MoE) leads on average quality at 94.0%, with a 100% PhD score, 100% on HEE, and the highest SDB-100 score (80%) of any model in the sweep. At 51.4 TPS it anchors the Desktop Pro tier alongside GPT-OSS 20B and Gemma 4 31B; Gemma 4 31B holds the raw quality ceiling (100%) but pays for it in speed on 16 GB VRAM hardware.
The 2–4B parameter class is remarkably competitive. Gemma 4 E2B, Phi-4 Mini 3.8B, Granite 4.1 3B Instruct, and Ministral 3 3B all deliver quality scores competitive with many 8B–14B models while running at roughly 195–225 TPS. This is the sweet spot for interactive local inference, and the current catalog draws its Lightweight tier from this band: Ministral 3 3B (vision-capable, the Free-tier vision model and Electron default), Phi-4 Mini 3.8B (MIT-licensed text), and Llama 3.2 3B Instruct — with Llama 3.2 1B Instruct anchoring the Ultra-Light tier.
Catalog revisions. The current shipping catalog is v2026.08.13.02 (13 August 2026): LFM2.5 VL 1.6B is the recommended default on desktop and Android; LFM2.5 VL 3B replaces LFM2 VL 3B in the Lightweight tier, following the 12 August revision that added GPT-OSS 20B and retired Llama 3.2 3B Instruct and Phi-4 Mini 3.8B. All ten shipping models sit on the Pareto frontier of this sweep, and nine of the ten are vision-capable.
Models below 1B parameters show steep quality drops. Granite 4.0 350M Dense (6%), Granite 4.0 1B Dense (6%), H2O-Danube 3 500M (11%), SmolLM2 135M Instruct (16%), and SmolLM2 360M Instruct (26%) sit at or below the utility floor for general-purpose assistant tasks. They may still serve narrow roles such as classification or simple extraction, but none is viable as a primary note-taking assistant.
© 2026 IWV Digital Solutions LLC. All rights reserved.
© 2026 IWV Digital Solutions LLC. All rights reserved.
← Back to Benchmark Results