NotesXML ships with 16 AI models you can download and run entirely on your device. No API keys, no subscriptions, no data leaving your machine. Here's how they compare to the cloud giants.
Every model listed below runs inside NotesXML via llama.cpp — the same inference engine used by researchers worldwide. Cloud models are shown for reference only.
| Model | Parameters | RAM (peak) | Context | Vision | Tier | Educational Equivalent |
|---|---|---|---|---|---|---|
| GPT-5.4 Pro Cloud | Proprietary | API | 1M+ | Yes | Cloud | PhD / Researcher |
| Claude Opus 4.6 Cloud | Proprietary | API | 1M | Yes | Cloud | PhD / Researcher |
| LFM2.5 VL 1.6B Local Free Tier Recommended | 1.6B | 2.0 GB | 32K | Yes | Ultra-Light | Basic / Elementary |
| Ministral 3 3B Local Free Tier Electron Default | 3B | 3.0 GB | 256K | Yes | Lightweight | High School |
| LFM2.5 VL 3B Local | 3.1B | 3.7 GB | 32K | Yes | Lightweight | High School |
| Gemma 4 E2B Local | 2.3B effective | 4.5 GB | 128K | Yes | Lightweight | High School |
| Gemma 4 E4B Local | 4.5B effective | 6.4 GB | 128K | Yes | Enhanced | Undergraduate |
| Ministral 3 8B Local | 8B | 7.0 GB | 256K | Yes | Enhanced | Undergraduate |
| Ministral 3 14B Desktop only Intel GPU pick | 14B | 11.0 GB | 256K | Yes | Advanced | Undergraduate |
| Gemma 4 12B Local | 12B | 10.0 GB | 256K | Yes | Enhanced | Undergraduate |
| Mellum2 12B-A2.5B (MoE, coding) Desktop only | 12B (2.5B active, MoE) | 13.0 GB | 128K | — | Desktop Pro | Post-Graduate |
| GPT-OSS 20B (MoE) Desktop only | 21B (3.6B active, MoE) | 18.0 GB | 128K | — | Desktop Pro | Post-Graduate |
| Devstral Small 2 24B (coding) Desktop only | 24B | 22.0 GB | 256K | Yes | Desktop Pro | Post-Graduate |
| Gemma 4 26B-A4B (MoE) Desktop only | 26B (4B active, MoE) | 22.0 GB | 256K | Yes | Desktop Pro | Post-Graduate |
| Gemma 4 31B Desktop only | 31B | 26.0 GB | 256K | Yes | Desktop Pro | Post-Graduate |
| Muse Glimmer 30B Desktop only | 30B (incl. 1.8B vision encoder) | 26.0 GB | 128K | Yes | Desktop Pro | Post-Graduate |
| Laguna XS 2.1 33B-A3B (MoE, coding) Desktop only | 33B (3B active, MoE) | 27.0 GB | 256K | — | Desktop Pro | Post-Graduate |
| Nemotron 3.5 Lightning 30B-A3B (MoE) Desktop only | 30B (3B active, MoE) | 32.0 GB | 128K | — | Desktop Pro | Post-Graduate |
RAM (peak) = peak memory used by the model during inference (per NX-AI-MODEL-015 RAM safety floor). This is the model’s own memory consumption, not the total device RAM required — your operating system, background apps, and the NotesXML application itself also consume RAM. As a guideline, add 4–6 GB to the peak RAM figure for Android devices or 5–8 GB for desktops to estimate the total device RAM needed. Context = native context window. MoE = Mixture of Experts (active parameters per token shown in parentheses). Vision-capable models accept images, PDFs, and screenshots as input. Free Tier models are available without a Professional license. Catalog source: notesxml-model-catalog.json v2026.08.13.02.
AI models run entirely on your device and produce results based on statistical patterns. Output quality varies by model and task. Always review AI-generated content before relying on it.
Specialty family catalogs are available for Gemma, Llama, Mistral, Phi, Granite, and GPT-OSS — 54 additional models across 6 families.
The catalog is organized into five tiers based on RAM requirements and capability. Pick the tier that matches your hardware. A sixth tier, Standard, was retired in 5.181.4 and is documented below for anyone who downloaded a model from it.
1 model · ~2.0 GB model runtime RAM · 6–8 GB total device RAM · Free Tier (LFM2.5 VL 1.6B)
LFM2.5 VL 1.6B (recommended default, both platforms). The recommended starting point on desktop and Android alike, and included in the free tier. Unusually for a model this small it is vision-capable, so image and handwriting recognition work without a Professional licence. Published by Liquid AI under the LFM Open License v1.0 — free for personal use, and for commercial use by organisations under US $10M annual revenue; see third-party licences.
3 models · 3.0–4.5 GB model runtime RAM · 8–10 GB total device RAM recommended
Ministral 3 3B (Free Tier, vision) and LFM2.5 VL 3B (vision). Ideal for quick tasks: formatting, basic summarization, general conversation and image questions, on modest hardware. Both are vision-capable. LFM2.5 VL 3B is the larger of the two and the fastest model in the catalog above the Ultra-Light tier; it is licensed under the LFM Open License v1.0 (see the licence note below). It replaced LFM2 VL 3B in NotesXML 5.188.0 after head-to-head testing; if you downloaded LFM2 VL 3B under an earlier version it remains on your device and continues to work. Llama 3.2 3B and Phi-4 Mini 3.8B were retired here in the 2026-08-12 catalog revision: both were text-only and both were outscored by the models that remain.
Tier retired · see Ultra-Light and Enhanced
Gemma 4 E2B returns to the catalog (5.198.33). Retired in 5.181.4, it was re-evaluated in September 2026 against the current suites and re-added to the Lightweight tier: vision and native audio, 128K context, Apache 2.0, 4.5 GB peak. It is not the default on any platform — the recommended model remains LFM2.5 VL 1.6B on Android and Ministral 3 3B on desktop — and on phones with limited free memory the model picker will warn before download.
3 models · 6.4–10.0 GB model runtime RAM · 16 GB total system RAM recommended · Professional
Gemma 4 E4B (vision, tool calling), Ministral 3 8B (vision), and Gemma 4 12B (vision). Gemma 4 E4B delivers expert-level tool calling and image analysis. Ministral 3 8B is the catalog’s speed/quality leader for AI Chat and structured actions. Gemma 4 12B is the highest-quality model that still runs cross-platform, second only to the desktop-only 26B MoE. All three require Professional tier.
1 model · ~11 GB model runtime RAM · 24 GB total system RAM recommended · Desktop & Professional
Ministral 3 14B (vision, Intel GPU pick). The large-model choice for Intel iGPU/GPU systems — measured on Intel hardware as the strongest large vision model of the catalog. 14B dense, vision-capable, 256K-token context, Apache 2.0. Desktop only (Windows and Linux); does not appear in the Android model picker.
8 models · 13–32 GB model runtime RAM · 32 GB total system RAM or 16+ GB VRAM · Desktop & Professional
Gemma 4 26B-A4B (MoE) — 26B total parameters with 4B active per token via Mixture-of-Experts, vision-capable, 256K-token native context. Top scores on HEE, PLE, and PhD Philosophy, plus the strongest SDB-100 deductive-reasoning result. Best with a discrete GPU with 16+ GB VRAM (e.g., RTX 4080, RTX 5070 Ti, or higher); CPU-only inference is possible but significantly slower. Apache 2.0.
GPT-OSS 20B (MoE) — 21B total parameters with 3.6B active per token, 128K-token context, Apache 2.0. Text-only, and the fastest model in the Desktop Pro tier by a wide margin: measured at roughly 2.7x the tokens per second of the 26B MoE on our 16 GB VRAM reference rig, at a lower memory ceiling. The trade-off is no vision and weaker tool calling, so it suits long-form writing, summarization and reasoning rather than image work or structured actions. Desktop only.
Gemma 4 31B (Dense) — 31B dense parameters, vision-capable, 256K-token native context. The highest raw quality score of any model we have benchmarked, and the largest in the catalog: 18.8 GB of weights plus a 1.2 GB vision projector, about 26 GB at peak. VRAM decides whether this model is fast or slow. A dense model touches every parameter on every token, so if the weights do not fit entirely in GPU memory the overflow spills to system RAM and throughput drops sharply — the effect is far more severe than on the Mixture-of-Experts models beside it, which activate only a fraction of their parameters per token. Give it a GPU that can hold roughly 26 GB and it runs as the strongest model in the catalog; run it on 16 GB VRAM or on CPU and expect it to be slow. If your GPU cannot hold it, Gemma 4 26B-A4B is the better choice at nearly the same quality. Desktop only (Windows and Linux). Apache 2.0.
Model licences. Most models in the catalog are Apache 2.0. Three are not: LFM2.5 VL 1.6B and LFM2.5 VL 3B are under the LFM Open License v1.0, which carries a commercial-use revenue condition (see the note on LFM2.5 VL above), and Nemotron 3.5 Lightning 30B-A3B and Laguna XS 2.1 33B-A3B are under the OpenMDW License Agreement v1.1 — permissive, with no revenue condition. NotesXML shows each model’s licence before you download it, and the full table is on the third-party licences page.
Cloud AI means sending your notes, documents, and ideas to someone else's server. With NotesXML, the AI runs on your device. Your data never leaves your machine — not even for processing.
Cloud AI services typically charge ongoing subscription fees — sometimes hundreds of dollars per month for top-tier models, plus per-token fees for usage. NotesXML Professional is $49.99 lifetime — and includes access to all 16 models with no per-token charges, ever.
No internet? No problem. Once you've downloaded a model, it works everywhere — airplanes, rural areas, secure facilities. Cloud models require a constant internet connection.