100% Local16 Models1.6B–33B ParametersZero Cloud

AI that runs on your hardware

NotesXML ships with 16 AI models you can download and run entirely on your device. No API keys, no subscriptions, no data leaving your machine. Here's how they compare to the cloud giants.

Local models vs. cloud services

Every model listed below runs inside NotesXML via llama.cpp — the same inference engine used by researchers worldwide. Cloud models are shown for reference only.

PhD / Researcher Post-Graduate University High School Basic
Model Parameters RAM (peak) Context Vision Tier Educational Equivalent
GPT-5.4 Pro Cloud Proprietary API 1M+ Yes Cloud PhD / Researcher
Claude Opus 4.6 Cloud Proprietary API 1M Yes Cloud PhD / Researcher
LFM2.5 VL 1.6B Local Free Tier Recommended 1.6B 2.0 GB 32K Yes Ultra-Light Basic / Elementary
Ministral 3 3B Local Free Tier Electron Default 3B 3.0 GB 256K Yes Lightweight High School
LFM2.5 VL 3B Local 3.1B 3.7 GB 32K Yes Lightweight High School
Gemma 4 E2B Local 2.3B effective 4.5 GB 128K Yes Lightweight High School
Gemma 4 E4B Local 4.5B effective 6.4 GB 128K Yes Enhanced Undergraduate
Ministral 3 8B Local 8B 7.0 GB 256K Yes Enhanced Undergraduate
Ministral 3 14B Desktop only Intel GPU pick 14B 11.0 GB 256K Yes Advanced Undergraduate
Gemma 4 12B Local 12B 10.0 GB 256K Yes Enhanced Undergraduate
Mellum2 12B-A2.5B (MoE, coding) Desktop only 12B (2.5B active, MoE) 13.0 GB 128K Desktop Pro Post-Graduate
GPT-OSS 20B (MoE) Desktop only 21B (3.6B active, MoE) 18.0 GB 128K Desktop Pro Post-Graduate
Devstral Small 2 24B (coding) Desktop only 24B 22.0 GB 256K Yes Desktop Pro Post-Graduate
Gemma 4 26B-A4B (MoE) Desktop only 26B (4B active, MoE) 22.0 GB 256K Yes Desktop Pro Post-Graduate
Gemma 4 31B Desktop only 31B 26.0 GB 256K Yes Desktop Pro Post-Graduate
Muse Glimmer 30B Desktop only 30B (incl. 1.8B vision encoder) 26.0 GB 128K Yes Desktop Pro Post-Graduate
Laguna XS 2.1 33B-A3B (MoE, coding) Desktop only 33B (3B active, MoE) 27.0 GB 256K Desktop Pro Post-Graduate
Nemotron 3.5 Lightning 30B-A3B (MoE) Desktop only 30B (3B active, MoE) 32.0 GB 128K Desktop Pro Post-Graduate

RAM (peak) = peak memory used by the model during inference (per NX-AI-MODEL-015 RAM safety floor). This is the model’s own memory consumption, not the total device RAM required — your operating system, background apps, and the NotesXML application itself also consume RAM. As a guideline, add 4–6 GB to the peak RAM figure for Android devices or 5–8 GB for desktops to estimate the total device RAM needed. Context = native context window. MoE = Mixture of Experts (active parameters per token shown in parentheses). Vision-capable models accept images, PDFs, and screenshots as input. Free Tier models are available without a Professional license. Catalog source: notesxml-model-catalog.json v2026.08.13.02.

AI models run entirely on your device and produce results based on statistical patterns. Output quality varies by model and task. Always review AI-generated content before relying on it.

Want more models?

Specialty family catalogs are available for Gemma, Llama, Mistral, Phi, Granite, and GPT-OSS — 54 additional models across 6 families.

Browse Specialty Catalogs →

Five tiers, one app

The catalog is organized into five tiers based on RAM requirements and capability. Pick the tier that matches your hardware. A sixth tier, Standard, was retired in 5.181.4 and is documented below for anyone who downloaded a model from it.

Ultra-Light

1 model · ~2.0 GB model runtime RAM · 6–8 GB total device RAM · Free Tier (LFM2.5 VL 1.6B)

LFM2.5 VL 1.6B (recommended default, both platforms). The recommended starting point on desktop and Android alike, and included in the free tier. Unusually for a model this small it is vision-capable, so image and handwriting recognition work without a Professional licence. Published by Liquid AI under the LFM Open License v1.0 — free for personal use, and for commercial use by organisations under US $10M annual revenue; see third-party licences.

Lightweight

3 models · 3.0–4.5 GB model runtime RAM · 8–10 GB total device RAM recommended

Ministral 3 3B (Free Tier, vision) and LFM2.5 VL 3B (vision). Ideal for quick tasks: formatting, basic summarization, general conversation and image questions, on modest hardware. Both are vision-capable. LFM2.5 VL 3B is the larger of the two and the fastest model in the catalog above the Ultra-Light tier; it is licensed under the LFM Open License v1.0 (see the licence note below). It replaced LFM2 VL 3B in NotesXML 5.188.0 after head-to-head testing; if you downloaded LFM2 VL 3B under an earlier version it remains on your device and continues to work. Llama 3.2 3B and Phi-4 Mini 3.8B were retired here in the 2026-08-12 catalog revision: both were text-only and both were outscored by the models that remain.

Standard

Tier retired · see Ultra-Light and Enhanced

Gemma 4 E2B returns to the catalog (5.198.33). Retired in 5.181.4, it was re-evaluated in September 2026 against the current suites and re-added to the Lightweight tier: vision and native audio, 128K context, Apache 2.0, 4.5 GB peak. It is not the default on any platform — the recommended model remains LFM2.5 VL 1.6B on Android and Ministral 3 3B on desktop — and on phones with limited free memory the model picker will warn before download.

Enhanced

3 models · 6.4–10.0 GB model runtime RAM · 16 GB total system RAM recommended · Professional

Gemma 4 E4B (vision, tool calling), Ministral 3 8B (vision), and Gemma 4 12B (vision). Gemma 4 E4B delivers expert-level tool calling and image analysis. Ministral 3 8B is the catalog’s speed/quality leader for AI Chat and structured actions. Gemma 4 12B is the highest-quality model that still runs cross-platform, second only to the desktop-only 26B MoE. All three require Professional tier.

Advanced

1 model · ~11 GB model runtime RAM · 24 GB total system RAM recommended · Desktop & Professional

Ministral 3 14B (vision, Intel GPU pick). The large-model choice for Intel iGPU/GPU systems — measured on Intel hardware as the strongest large vision model of the catalog. 14B dense, vision-capable, 256K-token context, Apache 2.0. Desktop only (Windows and Linux); does not appear in the Android model picker.

Desktop Pro

8 models · 13–32 GB model runtime RAM · 32 GB total system RAM or 16+ GB VRAM · Desktop & Professional

Gemma 4 26B-A4B (MoE) — 26B total parameters with 4B active per token via Mixture-of-Experts, vision-capable, 256K-token native context. Top scores on HEE, PLE, and PhD Philosophy, plus the strongest SDB-100 deductive-reasoning result. Best with a discrete GPU with 16+ GB VRAM (e.g., RTX 4080, RTX 5070 Ti, or higher); CPU-only inference is possible but significantly slower. Apache 2.0.

GPT-OSS 20B (MoE) — 21B total parameters with 3.6B active per token, 128K-token context, Apache 2.0. Text-only, and the fastest model in the Desktop Pro tier by a wide margin: measured at roughly 2.7x the tokens per second of the 26B MoE on our 16 GB VRAM reference rig, at a lower memory ceiling. The trade-off is no vision and weaker tool calling, so it suits long-form writing, summarization and reasoning rather than image work or structured actions. Desktop only.

Gemma 4 31B (Dense) — 31B dense parameters, vision-capable, 256K-token native context. The highest raw quality score of any model we have benchmarked, and the largest in the catalog: 18.8 GB of weights plus a 1.2 GB vision projector, about 26 GB at peak. VRAM decides whether this model is fast or slow. A dense model touches every parameter on every token, so if the weights do not fit entirely in GPU memory the overflow spills to system RAM and throughput drops sharply — the effect is far more severe than on the Mixture-of-Experts models beside it, which activate only a fraction of their parameters per token. Give it a GPU that can hold roughly 26 GB and it runs as the strongest model in the catalog; run it on 16 GB VRAM or on CPU and expect it to be slow. If your GPU cannot hold it, Gemma 4 26B-A4B is the better choice at nearly the same quality. Desktop only (Windows and Linux). Apache 2.0.

Model licences. Most models in the catalog are Apache 2.0. Three are not: LFM2.5 VL 1.6B and LFM2.5 VL 3B are under the LFM Open License v1.0, which carries a commercial-use revenue condition (see the note on LFM2.5 VL above), and Nemotron 3.5 Lightning 30B-A3B and Laguna XS 2.1 33B-A3B are under the OpenMDW License Agreement v1.1 — permissive, with no revenue condition. NotesXML shows each model’s licence before you download it, and the full table is on the third-party licences page.

Why local AI matters

Your data stays yours

Cloud AI means sending your notes, documents, and ideas to someone else's server. With NotesXML, the AI runs on your device. Your data never leaves your machine — not even for processing.

No subscriptions

Cloud AI services typically charge ongoing subscription fees — sometimes hundreds of dollars per month for top-tier models, plus per-token fees for usage. NotesXML Professional is $49.99 lifetime — and includes access to all 16 models with no per-token charges, ever.

Works offline

No internet? No problem. Once you've downloaded a model, it works everywhere — airplanes, rural areas, secure facilities. Cloud models require a constant internet connection.

Download NotesXML Free View Pricing