← Benchmark Results & Recommendations

All 49 Benchmarked AI Models — Complete Results

IWV Digital Solutions LLC | 75-model benchmark sweep | refreshed 12 August 2026

Reference data — full candidate sweep. This page tabulates the complete NotesXML model-selection sweep: all 76 benchmarked candidate models on Ryzen 7 7800X3D + RTX 5070 Ti 16 GB, Linux, GPU-accelerated inference, identical sampling parameters per model, sorted by HMS (descending). ★ marks the models in the current 10-model shipping catalog (v2026.08.13.02) — all 10 appear in this sweep. The 13 August 2026 revision (v2026.08.13.02) replaced LFM2 VL 3B with LFM2.5 VL 3B; the 12 August revision had added GPT-OSS 20B and retired Llama 3.2 3B Instruct and Phi-4 Mini 3.8B; retired models remain listed here as evaluated candidates. All 10 shipping models sit on the sweep’s Pareto frontier. Sorting is by HMS rather than raw quality, so fast small models are not buried beneath large slow ones.


Refreshed 13 August 2026 — catalog v2026.08.13.02, tool-calling suite updated, LFM2.5 VL 3B added. The scatter plot and the full results table on this page are regenerated from the current master benchmark catalog: 76 evaluated models, of which the 10 in the shipping NotesXML catalog are highlighted. LFM2.5 VL 1.6B (Liquid AI) enters as the recommended default on both platforms and a free-tier model — 584.8 TPS at 87% quality, vision-capable in 1.3 GB — and Gemma 4 31B joins the catalog; Llama 3.2 1B Instruct and Gemma 4 E2B were retired from the catalog and remain in the table below as evaluated candidates. The current catalog is on the AI Models page.

Performance Scatter Plot

Each point represents one of the 76 benchmarked models (refreshed 13 August 2026). The X-axis shows the token generation interval (1/TPS in milliseconds, log scale — lower is faster). The Y-axis shows the average quality score across the seven scored benchmarks (Code excluded). Models in the current 10-model catalog are highlighted in green (all 10 appear in this sweep).

Token interval vs quality, 75-model sweep Scatter plot of all 76 benchmarked models: token generation interval in milliseconds on a logarithmic X axis against average quality on the Y axis. The 10 models in shipping catalog v2026.08.13.02 are highlighted in green and labeled; other candidates are gray. Token Generation Interval vs Average Quality — 76 models (13 Aug 2026) 1ms 2ms 5ms 10ms 20ms 50ms 100ms 0% 20% 40% 60% 80% 100% Token generation interval, ms/token (log scale) — lower is faster — Ryzen 7 7800X3D + RTX 5070 Ti 16 GB Average quality (Code excluded) Gemma 4 26B-A4B Gemma 4 12B Ministral 3 14B Ministral 3 8B Gemma 4 E4B GPT-OSS 20B Ministral 3 3B LFM2.5 VL 3B LFM2.5 VL 1.6B Gemma 4 31B Shipping catalog v2026.08.13.02 (10 models) Evaluated candidate (66 models)

Complete Benchmark Table

All 76 evaluated models, sorted by HMS descending. ★ = in the v2026.08.13.02 shipping catalog (10 models) (v2026.08.13.02). A dash means the benchmark does not apply to that model.

Model Params (B) Vision TPS ms/tok Model Bench (text) HEE PLE PhD SDB LongCtx Tool Quality HMS
★ Gemma 4 26B-A4B26Yes51.419.5100%100%99%100%80%85%94%100%1.00
★ Gemma 4 12B12Yes70.714.1100%98%96%99%54%81%93%99%0.99
★ Ministral 3 14B14Yes74.613.4100%98%95%96%47%87%78%98%0.99
★ Ministral 3 8B8Yes1168.6100%96%95%96%72%86%83%97%0.99
★ Gemma 4 E4B4.5Yes118.88.4100%96%92%99%46%77%97%97%0.99
Phi-4 14B1477.113.0100%96%95%98%56%56%91%97%0.99
Phi-4 Reasoning Vision 15B15Yes73.113.7100%96%95%97%53%15%87%97%0.98
LFM 24B-A2B (MOE)2481.112.3100%96%87%98%32%14%0%95%0.98
Mistral Small 3.2 24B24Yes4820.897%96%96%100%37%75%94%97%0.97
Devstral Small 2 24B Instruct2448.420.7100%94%93%100%69%94%97%0.97
★ GPT-OSS 20B21141.57.189%100%93%96%69%82%4%94%0.97
★ Ministral 3 3B Instruct3Yes225.14.4100%88%89%94%41%88%84%94%0.97
Gemma 3 12B12Yes66.415.1100%92%91%93%82%93%94%0.97
Granite 4.1 3B3203.84.989%94%89%96%40%81%0%92%0.96
Gemma 4 E2B2.3Yes206.84.896%91%87%93%23%80%71%92%0.96
Granite 4.1 3B Instruct32114.789%94%89%95%40%81%60%92%0.96
Pixtral 12B12Yes84.811.894%82%85%96%26%89%86%91%0.96
Mistral NeMo 12B1283.612.0100%88%83%94%29%80%91%91%0.95
Granite 4.1 8B Instruct81277.986%94%89%95%29%84%92%91%0.95
OLMo-2-13B1373.213.796%84%85%97%15%84%90%0.95
Llama 3.2 8B Instruct81327.6100%86%85%89%20%81%84%90%0.95
★ LFM2.5 VL 3B3Yes279.33.692%88%83%89%41%30%0%90%0.95
LFM2 VL 3B3Yes270.33.794%84%79%93%46%31%0%90%0.95
Granite 4.0 H-Tiny (MOE)72354.389%90%86%90%38%79%84%89%0.94
Granite 3.3 8B Instruct8106.69.483%90%84%90%27%0%76%87%0.93
OLMo-2-1124 7B Instruct71377.3100%80%76%91%16%65%87%0.93
★ LFM2.5 VL 1.6B1.6Yes584.81.792%78%74%90%28%79%4%87%0.93
Gemma 3n E4B484.311.971%96%84%96%45%75%87%0.93
Phi-4 Mini 3.8B3.82294.492%86%81%86%48%84%82%86%0.93
Devstral Small 2507 24B2446.721.471%98%94%99%51%91%91%0.92
Mistral 7B Instruct v0.371208.389%84%81%88%0%76%86%0.92
OLMo-2-Specialized 7B71277.976%84%77%90%16%65%82%0.90
SmolLM3 3B (3B)3248.24.081%84%75%83%8%19%55%81%0.89
Gemma 3 4B4Yes111.98.965%90%80%86%82%31%80%0.89
Granite 3.0 2B2242.24.181%76%76%87%25%72%77%80%0.89
Gemma 3n E2B2112.18.966%86%83%84%44%67%80%0.89
Llama 3.2 3B Instruct3.22573.994%80%71%73%32%77%64%80%0.89
LFM 2.5 1.2B Thinking1.26441.667%86%81%84%41%14%0%79%0.89
Granite 3.3 2B2.52314.378%82%74%82%40%0%81%79%0.88
Granite 4.0 H-1B (Hybrid)12653.8100%78%68%69%39%18%71%79%0.88
Granite 4.1 4B4Yes182.55.583%94%87%92%32%88%81%76%0.87
SmolVLM (2B)2Yes3642.779%74%72%65%29%29%30%72%0.84
Index 1.9B Chat1.9243.44.175%72%66%71%29%0%57%71%0.83
Llama 3.2 1B Instruct1.25591.888%60%63%61%22%80%55%68%0.81
SmolLM2 1.7B Instruct1.73123.272%64%65%63%28%32%82%66%0.80
LFM 2.5 1.2B Instruct1.25171.988%56%52%68%25%18%57%66%0.79
LFM 2.5 VL 450M0.45Yes724.71.479%48%48%48%24%25%55%65%0.79
Granite 3.0 8B8105.89.585%45%17%87%34%0%86%58%0.74
Granite Vision 3.2 2B2Yes253.83.974%80%71%0%43%0%77%56%0.72
MythoMax L2 13B1369.214.579%70%68%0%0%0%54%0.70
LFM 2.5 350M0.356751.579%32%35%54%13%16%8%50%0.67
Gemma 3 1B1152.46.676%44%40%36%10%16%37%49%0.66
BioMistral 7B795.710.458%74%63%0%0%0%49%0.66
Granite 3.0 1B-A400M14442.3100%24%25%36%19%0%40%46%0.63
Qwen3.5 9B9Yes106.69.4100%10%28%17%30%49%97%39%0.56
Qwen3.5 35B-A3B (MoE)35Yes37.326.8100%10%16%18%30%47%91%44%0.55
Granite 4.0 H-Small3217.357.889%96%93%99%52%83%94%94%0.51
SmolLM2 360M Instruct0.36486.92.142%32%38%27%25%16%30%35%0.51
Qwen3.5 4B4Yes158.16.361%10%16%79%30%29%95%33%0.50
Qwen3.5 0.8B0.8Yes364.22.746%10%16%16%30%29%86%33%0.49
LFM2.5 2.6B2.63013.390%6%13%6%30%14%8%29%0.45
LFM 2.5 230M0.237241.467%10%16%19%30%18%0%28%0.44
H2O-Danube 3 500M0.5634.41.675%6%16%14%30%0%0%28%0.43
SmolLM2 135M Instruct0.135960.11.042%16%32%20%28%15%7%28%0.43
MiniCPM5 1B1583.61.769%10%12%17%30%29%29%27%0.43
LFM2.5 8B-A1B (MOE)855.917.969%6%15%14%24%28%0%26%0.41
Gemma 3 27B2712.778.761%98%93%97%0%94%87%0.39
Codestral 22B221283.394%80%68%83%0%72%81%0.37
Qwen3.5 27B27Yes14.469.4100%10%16%18%30%49%97%38%0.33
MiniCPM-V 4.60.8Yes4612.269%4%9%18%30%27%29%20%0.33
Granite 4.1 30308.9112.483%96%94%100%45%90%86%93%0.30
★ Gemma 4 31B31Yes8.5117.6100%100%100%100%51%92%91%100%0.29
MobileLLM 1.5B (R1.5)1.51666.021%8%17%13%0%0%15%0.26
Qwen3.5 2B2Yes284.73.513%10%16%18%30%29%44%11%0.20
Granite 4.0 1B (Dense)1140.37.135%0%0%0%1%10%0%9%0.16
Granite 4.0 350M (Dense)0.3515020.724%0%0%0%0%2%0%6%0.11

TPS = tokens per second on the benchmark hardware. ms/tok = 1000/TPS (token generation interval). Model Bench = Model Benchmark score (in-app task suite: six text tasks + one vision OCR, graded A–F). HEE, PLE, PhD, SDB, LongCtx, Tool = normalized scores (0–100%). Avg Score = unweighted mean of the seven normalized benchmark scores (Code excluded). HMS = 2 × min(TPS/200, 1) × AvgScore / (min(TPS/200, 1) + AvgScore).


Benchmark Descriptions

HEE (50 pts): Human Educational Equivalency — broad factual knowledge across STEM, humanities, health.
PLE (200 pts): Professional License Exam — 10 licensing domains including law, medicine, engineering.
PhD (100 pts): Doctoral-level logic and philosophy — the hardest single benchmark in the suite.
SDB-100 (10 tiers): Synthetic Deduction Benchmark — contamination-resistant structural reasoning.
LongCtx (100 pts): Long Context — retention and reasoning over extended input sequences.
Tool (100 pts): Tool Calling — structured function-call invocation and multi-step chains.
Code: Code generation and comprehension. Excluded from quality average (NotesXML is a note-taking app).
Model Bench: Model Benchmark score — the in-app task suite: six text tasks (summarization, auto-titling, grammar polish, translation, action-item extraction, Markdown/XML) plus one vision OCR task, each graded A–F.

Key Takeaways

Gemma 4 26B-A4B (MoE) leads on average quality at 94.0%, with a 100% PhD score, 100% on HEE, and the highest SDB-100 score (80%) of any model in the sweep. At 51.4 TPS it anchors the Desktop Pro tier alongside GPT-OSS 20B and Gemma 4 31B; Gemma 4 31B holds the raw quality ceiling (100%) but pays for it in speed on 16 GB VRAM hardware.

The 2–4B parameter class is remarkably competitive. Gemma 4 E2B, Phi-4 Mini 3.8B, Granite 4.1 3B Instruct, and Ministral 3 3B all deliver quality scores competitive with many 8B–14B models while running at roughly 195–225 TPS. This is the sweet spot for interactive local inference, and the current catalog draws its Lightweight tier from this band: Ministral 3 3B (vision-capable, the Free-tier vision model and Electron default), Phi-4 Mini 3.8B (MIT-licensed text), and Llama 3.2 3B Instruct — with Llama 3.2 1B Instruct anchoring the Ultra-Light tier.

Catalog revisions. The current shipping catalog is v2026.08.13.02 (13 August 2026): LFM2.5 VL 1.6B is the recommended default on desktop and Android; LFM2.5 VL 3B replaces LFM2 VL 3B in the Lightweight tier, following the 12 August revision that added GPT-OSS 20B and retired Llama 3.2 3B Instruct and Phi-4 Mini 3.8B. All ten shipping models sit on the Pareto frontier of this sweep, and nine of the ten are vision-capable.

Models below 1B parameters show steep quality drops. Granite 4.0 350M Dense (6%), Granite 4.0 1B Dense (6%), H2O-Danube 3 500M (11%), SmolLM2 135M Instruct (16%), and SmolLM2 360M Instruct (26%) sit at or below the utility floor for general-purpose assistant tasks. They may still serve narrow roles such as classification or simple extraction, but none is viable as a primary note-taking assistant.


© 2026 IWV Digital Solutions LLC. All rights reserved.


© 2026 IWV Digital Solutions LLC. All rights reserved.

← Back to Benchmark Results