Why Local AI Matters for Your Finances
The question is where it reads
An assistant that answers questions about your spending has to read your spending. That is not a design flaw, it is the job. The only real question is where the reading happens, because wherever it happens, a copy of your financial records exists there.
For a cloud assistant, that place is a server. The good ones are explicit about it — Monarch publishes that transaction data goes to third-party model providers under agreements barring storage and training, and that its AI categorisation can’t be opted out of. That is a defensible arrangement, honestly described. It is also a copy of your finances on someone else’s computer.
PocketVault Finance’s answer is to do the reading on the device the records are already on.
What actually runs on your phone
- Ask answers questions about your money using a Gemma model running locally, or Apple’s built-in model on recent iPhones, where there is nothing to download at all.
- Retrieval pulls in the records a question is about, using EmbeddingGemma embeddings computed on the device, so a question about coffee reaches a café whose name never says so.
- Every tool the model can reach is read-only. It looks things up through the same calculations that produce the app’s own screens. There is no path from a conversation to a changed record.
The model downloads once — around 2.6 GB — and runs offline from then on. It is optional: tracking, budgets, bills, import and every report work without it.
What we deliberately do not use it for
This is the part most apps leave vague, so: the model does not categorise your transactions, and it does not read your bank’s messages. Both of those are done by rules.
Categorising runs in a fixed order — what the bank’s own message stated, then how you filed that merchant last time, then a merchant list shipped with the app. Reading an SMS uses a per-bank template first and a general grammar second, and anything neither can read goes to a review queue rather than being guessed at.
That is not a gap waiting for a better model. It is the more useful design for this particular job. A rule that misreads something misreads it identically every time, which means it can be found, reproduced, fixed and then tested against so it never comes back. We publish how often that import gets it wrong, and that number would be meaningless if the answer could change between two runs over the same message.
A model is the right tool for an open-ended question about your money. It is the wrong tool for a decision you need to be able to audit.
What local costs you
Worth naming plainly, because it is not free.
A model that fits on a phone is not as capable as one running in a datacentre, and it is slower. Money in particular is a weak spot: small models are prone to mangling a figure mid-sentence, which is why the app shows amounts in cards drawn from its own calculations rather than letting the model type them into a sentence. A figure it cannot match to a real calculation is held back rather than shown.
You also pay 2.6 GB of storage and a one-time download, which is a genuine obstacle on a budget phone — one reason the app never requires it for anything that matters.
What you get for that is an assistant with no server behind it. The privacy property does not rest on us keeping a promise, or on a processing agreement with a vendor, or on a policy that can be revised. It rests on there being nowhere for the data to go. You can check that yourself with a packet capture, which is the only kind of privacy claim worth making.
Where this goes
On-device models keep getting smaller and better, and the app gets the benefit without changing shape. When a better model fits on a phone, it runs on the phone.
What will not change is the division of labour: open questions to the model, auditable decisions to rules. If we ever move import or categorisation onto a model, that would be a change worth announcing loudly rather than slipping into a release note — and the error rate we publish is where you would see it first.