← All Articles

NotesXML AI Model Benchmark Results & Recommendations

IWV Digital Solutions LLC | 10-model catalog v2026.07.26.01

10-model catalog, refreshed benchmarks. The shipping catalog now holds 10 models across four hardware tiers. 8 of the 10 sit on the Pareto efficiency frontier (speed vs quality) on the Ryzen reference rig; the two that do not — Gemma 4 E2B and Ministral 3 14B — were added for light-hardware and Intel-GPU users (explained below). The sweep covers eight benchmarks (Model Benchmark, HEE, PLE, PhD Philosophy, SDB-100, Code, Long Context, Tool Calling) on Ryzen 7 7800X3D, Linux. The HMS leaders are Ministral 3 8B and Gemma 4 12B (HMS 0.95 each); the quality leader is the desktop-only Gemma 4 26B-A4B (95%).


Overview

NotesXML's Professional tier ships with 10 local AI models. They span every hardware tier, from phones and low-RAM devices to enthusiast workstations with discrete GPUs. This article gives the full benchmark results for all 10 catalog models, the Harmonic Mean Score (HMS) ranking, and the efficiency-frontier analysis.

Catalog v2026.07.26.01 (July 2026): The catalog was built by running the full 8-benchmark suite against a broad set of candidate models, computing the Pareto frontier on TPS vs average quality, and applying a utility floor to drop models below viable capability. 8 of the 10 shipping models are on the strict Pareto frontier. Two more were added for cases the reference rig does not capture: Ministral 3 14B for Intel iGPU/GPU systems (a large, vision-capable, desktop-only model that runs on the Intel GPU, where the Gemma 4 models run CPU-only), and Gemma 4 E2B as a small, fast, vision-capable option for light hardware. Seven models are vision-capable: Ministral 3 3B, Ministral 3 8B, Ministral 3 14B, Gemma 4 E2B, Gemma 4 E4B, Gemma 4 12B (cross-platform), and Gemma 4 26B-A4B (Desktop Pro only). Free tier: Llama 3.2 1B Instruct (text, Android default) and Ministral 3 3B (vision-capable). Electron default: Ministral 3 3B.

All AI inference in NotesXML runs entirely on your device. No internet connection is required. No data is transmitted.


Benchmark Deep-Dives

Each benchmark in the suite has its own article. The article covers how the benchmark works, the full question or task list, how answers are scored, and the research behind its design. This page gives the scores; the deep-dive articles explain what the scores mean.

HEE — 50 pts
Human Educational Equivalency
50-question breadth exam across Linguistics, History, STEM, and Health. Maps score to an educational equivalency level (High School → Post-Graduate).
PLE — 200 pts
Professional License Exam
200 questions drawn from 10 real licensing exams: MBE, USMLE, PE, EPPP, NCLEX-RN, Real Estate, CPA, ARE, and CISSP.
PhD — 100 pts
Logic & Philosophy
100 doctoral-level questions in modal logic, philosophy of language, Kripke semantics, Gödel's theorems, epistemology, and metaphysics. The hardest benchmark in the suite.
SDB-100 — 100 pts
Synthetic Deduction Benchmark
Procedurally generated at runtime from 10 logic templates × 10 semantic themes. Training-data contamination impossible by design. Pure structural reasoning.
LongCtx — 100 pts
Long Context Benchmark
20 tasks across 4K–128K context tiers. Tests needle retrieval, position bias, multi-fact recall, and long-range reasoning. VRAM cascade detection built in.
ToolCall — 100 pts
Tool Calling Benchmark
25 tasks across schema compliance, tool selection, parameter extraction, multi-step sequences, error recovery, and appropriate refusal. JSON schema graded only — no tools executed.

NotesXML Benchmark Suite

NotesXML uses eight purpose-designed benchmarks to evaluate models as they actually run on local hardware — testing the combination of model capability and local inference performance that users experience in practice.

Long Context Benchmark (100 points)

Evaluates a model's ability to process, retain, and reason over extended input sequences. Passages with embedded facts at varying context depths are presented; the model is queried on details requiring genuine retention rather than positional pattern-matching. Critical for note-taking applications where users ask AI to summarize or analyze multi-page documents.

Score RangeLevel
80–100 (80–100%)Strong Retention
60–79 (60–79%)Functional Retention
40–59 (40–59%)Partial Retention
Below 40 (< 40%)Limited Retention

Tool Calling Benchmark (100 points)

Evaluates a model's ability to correctly invoke structured function calls. Tests cover simple single-tool invocations, multi-step chains, parallel calls, and error recovery — all capabilities underpinning NotesXML's AI action system.

Score RangeLevel
90–100 (90–100%)Expert Tool Use
80–89 (80–89%)Advanced Tool Use
70–79 (70–79%)Reliable Tool Use
Below 70 (< 70%)Basic Tool Use

Model Benchmark — In-App Task Suite (Grade A–F + Harmonic)

Runs the seven tasks NotesXML performs in practice and grades each A–F: six text tasks — summarization, auto-titling, grammar and spelling correction, English→Spanish translation, action-item extraction, and Markdown/XML generation — plus a vision OCR task for vision-capable models. The per-task grades are combined into a normalized quality score (exported as the AvgScore field) and an overall letter grade; the Harmonic score pairs that quality with inference speed, rewarding models that are both correct and fast. Chat-template handling, stop-token behavior, and output-format integrity are scored within each task.

HEE — Human Education Evaluation (50 points)

Evaluates broad factual knowledge across mathematics, science, history, literature, geography, and reasoning — drawn from secondary school through postgraduate level, weighted toward the upper end. The closest NotesXML benchmark to the academic MMLU family.

Score RangeLevel
45–50 (90–100%)Post-Graduate / Mastery
40–44 (80–89%)Undergraduate Level
30–39 (60–79%)High School Graduate
Below 30 (< 60%)Below Standard

PLE — Professional License Exam (200 points)

200 questions across ten professional fields: Law (MBE), Medicine (USMLE), Engineering (PE), Psychology (EPPP), Nursing (NCLEX-RN), Real Estate, Finance (CPA), Architecture (ARE), and Cybersecurity (CISSP). Tests not only factual recall but multi-step logical deduction and contextual synthesis under professional-exam constraints.

Score RangeLevel
190–200 (95–100%)Expert / Mastery
170–189 (85–94%)Expert / Mastery
160–169 (80–84%)Proficient
140–159 (70–79%)Developing
Below 140 (< 70%)Below Competency

PhD Philosophy — PhD-Level Logic & Philosophy Comprehensive Exam (100 points)

100 questions spanning metalogic, modal logic, philosophy of language and mind, epistemology, metaphysics, and continental phenomenology. Designed so that correctly answering requires understanding how theories mechanically function, not merely pattern-matching to familiar names. The most intellectually demanding single benchmark in the suite.

Score RangeLevel
90–100 (90–100%)PhD Mastery / Expert
80–89 (80–89%)Advanced
70–79 (70–79%)Proficient
Below 70 (< 70%)Below PhD Level

SDB-100 — Synthetic Deduction Benchmark (100 points)

10 questions testing multi-step deductive reasoning using synthetic ontologies that cannot exist in any training corpus — novel self-contained universes governed by arbitrary but logically absolute rules. By stripping away semantic priors, the benchmark forces genuine structural reasoning rather than pattern completion. The most discriminating and contamination-resistant benchmark in the suite.

Score RangeLevel
8–10 (80–100%)Advanced Deductive Reasoning
6–7 (60–79%)Proficient
4–5 (40–59%)Developing
Below 4 (< 40%)Below Standard

Full Results

NotesXML Benchmark Results

Full 8-benchmark sweep for all 10 catalog models, AI Benchmark Orchestrator v2.3 on Ryzen 7 7800X3D (Linux). The first eight models were run on Beta 5.141.38 (30 Jun 2026); Gemma 4 E2B and Ministral 3 14B were added and run on Beta 5.142.19 (9 Jul 2026) on the same rig. Rows are sorted by HMS (the speed/quality score). ★ marks the 8 models on the Pareto efficiency frontier; the two unmarked models (Gemma 4 E2B, Ministral 3 14B) are off the frontier on this hardware and are explained below.

ModelTierAvg TPSGradeHEE /50PLE /200PhD /100SDB-100 /100Code /100LongCtx /100Tool /100Avg ScoreHMS
Ministral 3 8B ★Enhanced71.6A48 (96%)184 (92%)94 (94%)72 (72%)95 (95%)88 (88%)83 (83%)90%0.95
Gemma 4 12B ★Enhanced66.8A49 (98%)192 (96%)99 (99%)57 (57%)97 (97%)81 (81%)93 (93%)90%0.95
Ministral 3 14BAdvanced (desktop)71.6A49 (98%)190 (95%)97 (97%)47 (47%)92 (92%)87 (87%)78 (78%)87%0.93
Gemma 4 E4B ★Enhanced119.0A48 (96%)184 (92%)96 (96%)49 (49%)84 (84%)57 (57%)94 (94%)83%0.91
Ministral 3 3B ★Lightweight220.1A42 (84%)162 (81%)81 (81%)41 (41%)88 (88%)86 (86%)84 (84%)81%0.89
Phi-4 Mini 3.8B ★Lightweight220.0A43 (86%)162 (81%)87 (87%)48 (48%)86 (86%)84 (84%)82 (82%)81%0.89
Gemma 4 E2BStandard198.3A47 (94%)173 (87%)92 (92%)23 (23%)91 (91%)80 (80%)71 (71%)79%0.88
Gemma 4 26B-A4B (MoE) ★Desktop Pro39.9A50 (100%)198 (99%)100 (100%)80 (80%)99 (99%)85 (85%)94 (94%)95%0.87
Llama 3.2 3B Instruct ★Lightweight253.0A40 (80%)142 (71%)72 (72%)32 (32%)83 (83%)78 (78%)64 (64%)73%0.84
Llama 3.2 1B Instruct ★Ultra-Light544.9A30 (60%)126 (63%)61 (61%)22 (22%)67 (67%)78 (78%)53 (53%)63%0.77

Avg TPS = tokens per second (text tasks only). Grade = Model Benchmark overall grade (all 10 models Grade A). Avg Score = qualityMean, the unweighted mean of all eight normalized benchmark scores (Model Benchmark, HEE, PLE, PhD, SDB-100, Code, LongCtx, Tool); Code is now included. HMS = harmonic mean of min(TPS/50,1) and Avg Score.

Performance vs Token Interval — The Efficiency Frontier

The plot shows each model’s quality against its token interval (1/TPS, milliseconds per generated token — lower is faster). The dashed line is the Pareto efficiency frontier: the models that no other model beats on both speed and quality at once. 8 of the 10 catalog models sit on the frontier. The two that do not — Gemma 4 E2B and Ministral 3 14B (shown as hollow markers) — are each matched by a model that is at least as fast and at least as good on this rig. They stay in the catalog for reasons the reference rig cannot show: see the note under the chart.

Efficiency Frontier: token interval (1/TPS) vs quality (10-model catalog) Scatter of the 10 NotesXML catalog models; 8 are on the Pareto efficiency frontier, and Gemma 4 E2B and Ministral 3 14B are off the frontier on this hardware. Performance vs Token Interval — 10 Catalog Models 0 5 (200 TPS) 10 (100 TPS) 15 (67 TPS) 20 (50 TPS) 25 (40 TPS) 60% 65% 70% 75% 80% 85% 90% 95% 50 TPS speed cap 1 / TPS (ms per token — lower is faster) Quality (8-benchmark qualityMean) Llama 3.2 1B Instruct Llama 3.2 3B Instruct Ministral 3 3B 👁 Phi-4 Mini 3.8B Gemma 4 E4B 👁 Ministral 3 8B 👁 Gemma 4 12B 👁 Gemma 4 26B-A4B (MoE) 👁 Gemma 4 E2B 👁 (off-frontier) Ministral 3 14B 👁 (off-frontier) 8 of 10 models on Pareto frontier Frontier model Off-frontier (dominated on this rig)
ModelAvg TPS1/TPS (ms/token)Avg ScoreFrontier
Llama 3.2 1B Instruct544.91.8463%★ YES
Llama 3.2 3B Instruct253.03.9573%★ YES
Ministral 3 3B220.14.5481%★ YES
Phi-4 Mini 3.8B220.04.5581%★ YES
Gemma 4 E2B198.35.0479%— no
Gemma 4 E4B119.08.4083%★ YES
Ministral 3 8B71.613.9790%★ YES
Ministral 3 14B71.613.9787%— no
Gemma 4 12B66.814.9790%★ YES
Gemma 4 26B-A4B (MoE)39.925.0695%★ YES

Sorted by TPS descending (fastest first). 8 of the 10 models are on the strict Pareto efficiency frontier. Gemma 4 E2B is beaten by Ministral 3 3B (faster and higher quality), and Ministral 3 14B is beaten by Ministral 3 8B (same speed, higher quality) — both on this Ryzen rig. Quality is the 8-benchmark qualityMean (Code included).

Harmonic Mean Score — Single-Number Speed/Quality Ranking

The Harmonic Mean Score (HMS) folds speed and quality into one number. Higher is better. The harmonic mean punishes imbalance: a model that is fast but weak, or strong but slow, scores lower than one that is good at both. Speed is capped at 50 TPS. Past that point a model already streams faster than most people read in a chat window, so extra speed adds nothing to the experience.

Methodology

  1. Speed component: speed_norm = min(TPS / 50, 1.0)
  2. Quality component: quality = qualityMean — the unweighted mean of all eight normalized benchmark scores. Code Benchmark is now included.
  3. Harmonic mean: HMS = 2 × speed_norm × quality / (speed_norm + quality)

HMS Ranking (all 10 catalog models)

RankModelTierAvg TPSAvg %speed_normqualityHMS
1Ministral 3 8BEnhanced71.690%1.0000.900.95
2Gemma 4 12B ★Enhanced66.890%1.0000.900.95
3Ministral 3 14BAdvanced (desktop)71.687%1.0000.870.93
4Gemma 4 E4B ★Enhanced119.083%1.0000.830.91
5Ministral 3 3B ★Lightweight220.181%1.0000.810.89
6Phi-4 Mini 3.8B ★Lightweight220.081%1.0000.810.89
7Gemma 4 E2BStandard198.379%1.0000.790.88
8Gemma 4 26B-A4B (MoE) ★Desktop Pro39.995%0.8000.950.87
9Llama 3.2 3B Instruct ★Lightweight253.073%1.0000.730.84
10Llama 3.2 1B Instruct ★Ultra-Light544.963%1.0000.630.77

Reading the HMS Ranking

Ministral 3 8B and Gemma 4 12B lead at HMS 0.95. Both clear the 50-TPS speed cap (71.6 and 66.8 TPS) and score 90% average quality, so they pair top speed with near-top quality. Ministral 3 8B is listed first as the faster of the two. Ministral 3 14B follows at 0.93, then Gemma 4 E4B at 0.91, then Ministral 3 3B and Phi-4 Mini tie at 0.89.

With the cap at 50 TPS, every model except the 26B reaches speed_norm = 1.0. For those models, HMS is driven almost entirely by quality. This is by design: once a model streams faster than you can read, quality is what separates them. The fastest model, Llama 3.2 1B (544.9 TPS), ranks last at 0.77 because its 63% quality is the lowest in the catalog.

The quality leader sits mid-pack on HMS. Gemma 4 26B-A4B (MoE) has the catalog’s highest quality (95%) but ranks eighth on HMS. Its 39.9 TPS is the only speed below the cap (speed_norm 0.80). For batch work, research synthesis, or high-stakes writing where speed matters less, read the quality column directly — the 26B is the top choice there.


Key Observations

8 of the 10 catalog models are on the Pareto frontier. On the Ryzen reference rig, two models are dominated. Gemma 4 E2B loses to Ministral 3 3B, which is both faster and higher quality. Ministral 3 14B loses to Ministral 3 8B, which runs at the same speed with higher quality. Both models stay in the catalog on purpose. Ministral 3 14B is there for Intel iGPU and GPU systems: it is a large, vision-capable, desktop-only model that runs on the Intel GPU, whereas the Gemma 4 models run CPU-only on Intel hardware. The reference rig is a Ryzen machine, so its numbers cannot show that gain. Gemma 4 E2B is a small, fast, vision-capable model for light hardware. The other eight models each win on some mix of speed, quality, memory footprint, or features.

Mid-size Enhanced models lead the HMS ranking. Ministral 3 8B and Gemma 4 12B (HMS 0.95) clear the 50-TPS cap and deliver 90% quality — the best speed/quality balance in the catalog. Among models fast enough to stream comfortably, the ranking rewards quality.

Gemma 4 26B-A4B (MoE) leads on quality. It takes the top spot on every knowledge and reasoning benchmark: 95% average quality, 100% HEE, 99% PLE, 100% PhD Philosophy, 80% SDB-100 (catalog leader), and 99% Code. With vision and a 256K-token context, it is the highest-quality model in the catalog. Desktop Pro only.

Vision now spans the catalog. Seven models are vision-capable: Ministral 3 3B, Gemma 4 E2B, Gemma 4 E4B, Ministral 3 8B, and Gemma 4 12B run cross-platform; Ministral 3 14B and Gemma 4 26B-A4B are desktop-only. Ministral 3 3B brings vision into the Free tier for the first time. Long Context retention is led by Ministral 3 8B (88%); Tool Calling by Gemma 4 E4B (94%, cross-platform).

Llama 3.2 1B’s knowledge scores reflect its size. At 1.2B parameters it is the fastest model in the catalog (544.9 TPS) and the Android Free-tier default. It is built for instant titling, summaries, and grammar polish on tight memory. For knowledge-heavy work, any 3B-or-larger model is clearly better.


Model Recommendations by Hardware Tier

Ultra-Light — Up to ~1.5 GB RAM Available for AI

Free-tier default: Llama 3.2 1B Instruct (HMS 0.77) — the Android Free-tier default and fastest model in the catalog at 544.9 TPS. 128K native context, text-only. Best for instant note titling, short summaries, transcription polish, and basic text cleanup under severe memory constraints. Llama 3.2 Community License.


Lightweight — ~1.5–4.5 GB RAM Available for AI

HMS co-leader (Lightweight): Ministral 3 3B (HMS 0.89) — Apache 2.0, 220.1 TPS, vision-capable, and the Electron default. The first vision model in the Free tier. 81% average quality (HEE 84%, PLE 81%, PhD 81%, Code 88%, Long Context 86%, Tool 84%), 256K native context. The best all-round default for note-taking with image support on Lightweight hardware.

MIT-licensed text default: Phi-4 Mini 3.8B (HMS 0.89) — MIT license, 220.0 TPS, text-only. PhD 87%, PLE 81%, Code 86%, Long Context 84%, Tool 82%. 128K native context. A strong MIT-licensed text default that ties Ministral 3 3B on HMS.

Llama family: Llama 3.2 3B Instruct (HMS 0.84) — 253.0 TPS, HEE 80%, PLE 71%, PhD 72%, Tool 64%. Llama 3.2 Community License, text-only, 128K native context.

Standard-tier vision option: Gemma 4 E2B (HMS 0.88) — Apache 2.0, 198.3 TPS, vision-capable, ~2.3B effective, 128K native context. A small, fast vision model for light hardware: 79% average quality (HEE 94%, PLE 87%, PhD 92%, Code 91%, Long Context 80%). On the reference rig Ministral 3 3B is both faster and higher quality, so E2B sits just off the frontier — but it stays a compact, vision-capable choice where its smaller footprint fits better.


Enhanced — ~5–9 GB RAM Available for AI

HMS leader: Ministral 3 8B (HMS 0.95) — Apache 2.0, 71.6 TPS, vision-capable. The catalog’s Long Context leader (88%) with 90% average quality (HEE 96%, PLE 92%, PhD 94%, Code 95%). 256K native context. The top speed/quality balance for AI Chat, structured actions, and image analysis on Enhanced hardware.

Tool-calling flagship: Gemma 4 E4B (HMS 0.91) — Apache 2.0, 119.0 TPS, Tool 94% (catalog co-leader), HEE 96%, PLE 92%, PhD 96%. Vision-capable, 128K native context. The strongest cross-platform choice for tool-driven AI actions.

Highest-quality cross-platform model: Gemma 4 12B (HMS 0.95) — Apache 2.0, 66.8 TPS, 90% average quality (tied for the catalog’s second-highest, behind only the desktop-only 26B MoE), Code 97%, PhD 99%, Tool 93%. Vision-capable, 12B dense, 256K native context. The top-quality model that still runs cross-platform — and it ties for the overall HMS lead.

Intel iGPU / GPU pick: Ministral 3 14B (HMS 0.93, Advanced tier, desktop only) — Apache 2.0, vision-capable, 14B dense, 256K native context, ~11 GB peak RAM. This is the large-model option for Intel systems: it runs on the Intel iGPU or GPU, where the Gemma 4 models fall back to CPU-only. On the Ryzen reference rig it measured 71.6 TPS at 87% average quality (HEE 98%, PLE 95%, PhD 97%, Code 92%, Long Context 87%), so on that machine Ministral 3 8B just edges it on the speed/quality frontier. On an Intel GPU, though, the 14B is how you run a large, high-quality vision model at usable speed. Desktop only (Windows and Linux).


Desktop Pro — ~22 GB RAM or 16+ GB GPU VRAM

Highest overall quality: Gemma 4 26B-A4B (MoE) (HMS 0.87) — Apache 2.0, desktop-only, 39.9 TPS. The catalog’s top quality model: 95% average, HEE 100%, PLE 99%, PhD 100%, SDB-100 80% (catalog leader), Code 99%, Tool 94%. 26B total / 4B active Mixture-of-Experts, vision-capable, 256K native context. The highest-quality choice for research, synthesis, and agentic workflows where speed is secondary.


Summary Recommendation Table

Your Hardware / Use CaseRecommended ModelHMSPrimary Strength
Phone / very low-RAM device (<1.5 GB)Llama 3.2 1B Instruct0.77Android Free-tier default; 544.9 TPS; smallest viable footprint; Llama 3.2 Community
Light device, best all-round default — vision in Free tierMinistral 3 3B0.89Electron default; vision-capable; Free-tier vision; 256K context; Apache 2.0
Light device, MIT-licensed text defaultPhi-4 Mini 3.8B0.89PhD 87%; Long Context 84%; MIT license; text-only
Light device, Llama familyLlama 3.2 3B Instruct0.84253.0 TPS; Llama-family; 128K context; Llama 3.2 Community
Light device, compact vision optionGemma 4 E2B0.88198.3 TPS; vision-capable; ~2.3B effective; small footprint; Apache 2.0
Enhanced desktop — catalog HMS co-leaderMinistral 3 8B0.95Long Context leader (88%); vision-capable; 256K context; Apache 2.0
Enhanced desktop — tool calling (Pro)Gemma 4 E4B0.91Tool 94% (co-leader); vision-capable; cross-platform; Apache 2.0
Enhanced desktop — highest quality cross-platform (vision), HMS co-leaderGemma 4 12B0.95Quality 90%; Code 97%; PhD 99%; vision-capable; Apache 2.0
Intel iGPU / GPU — large vision model on Intel hardware (desktop)Ministral 3 14B0.93Runs on the Intel GPU (Gemma 4 is CPU-only there); 14B; vision-capable; 256K context; desktop only; Apache 2.0
Desktop Pro — highest quality + visionGemma 4 26B-A4B (MoE)0.87Quality leader (95%); SDB-100 80%; vision-capable; 256K context; Apache 2.0

Complete Benchmark Data

Full benchmark results for every model evaluated in the candidate sweep — including models not selected for the catalog — are available in a dedicated reference article.

View All Benchmarked Models →

Notes on Hardware Acceleration

The benchmarks above were collected on a Ryzen 7 7800X3D system running Linux (AI Benchmark Orchestrator v2.3, Beta 5.141.38). NotesXML supports GPU acceleration on NVIDIA CUDA/Vulkan, AMD Vulkan (RADV), and Intel iGPU (Vulkan / SYCL). CPU-only inference remains fully supported on every platform.


About NotesXML AI

All AI runs locally on your device using llama.cpp. No subscription is needed beyond the one-time Professional license. Models download once and are stored on your machine. No internet is needed when the AI runs. Your notes, questions, and AI responses stay private.

© 2026 IWV Digital Solutions LLC. All rights reserved.


© 2026 IWV Digital Solutions LLC. All rights reserved.

← Back to Articles