Skip to content
Back to Blog
Product

Why We Trained Our Own AI Instead of Renting One

William Finger4 min

When I started building Geneziz, there was a standard recipe for adding AI to an app, and everybody followed it: rent a big, powerful intelligence from a datacenter, send your users' data to it, and charge a monthly toll for the privilege.

I didn't want that recipe. Partly because I'm stubborn. Mostly because it breaks the only promise Geneziz makes: your files stay on your machine.

Renting didn't fit

Geneziz is local-first on purpose. Your bookmarks and captures become plain Markdown files on your computer — files you can open, search, and keep for twenty years. So the standard recipe came with three costs nobody puts on the pricing page:

  • Your library leaves home. Organizing "in the cloud" means shipping your saved pages to someone else's servers to be sorted. That's the exact opposite of local-first.
  • It dies without internet. A rented brain lives far away. If your connection doesn't stay up, neither does your organizer.
  • The meter never stops. Every organized item, every question asked — someone, somewhere, is paying rent. And rent always gets passed on.

None of that needed to exist. It only existed because training your own was supposed to be too hard.

So we trained our own

Geneziz AI is ours. Not licensed, not borrowed, not white-labeled — trained by us, in our own lab, on hardware we own. Every phase, from early experiments to the version that ships, ran on one Mac and one PC. We never rented a GPU. Not once.

I want to be honest about what that means. It's the slow way. Big labs throw money at this; we built ours on hardware we already owned. Some lessons took weeks I didn't plan for.

But slow bought something: because we own every piece, there's nobody upstream to pay. No monthly bill for me, no subscription for you.

Teaching it to be a good archivist

The organizer inside Geneziz had to learn one thing really well: the judgment of a good archivist. Not my folder structure, not yours — the craft of deciding where things belong, so it adapts to whatever library it meets.

We made our own study material for that. We graded its homework ourselves, again and again, against one honest bar: it ships when it's actually good, not when it looks smart. And to be clear about something that matters — it learned from material we created for the purpose. Your library was never in the lesson plan. Nobody's was.

Four things made a tiny brain punch far above its size:

  • A shaped question, not an essay prompt. We never ask "which folder and why?" We ask a multiple-choice question and read the model's confidence over the options. One decision, one token — discrimination is trainable in a small brain; prose is not.
  • The right data, at the right dose. We measured a law: synthetic material plateaus no matter how much you add, and strangers' data is noise. What moved the needle was real, curated material — a few hundred examples took the organizer from guessing to landing the right folder in the top three 94% of the time.
  • One frozen recipe. The training method never changes between attempts — the dataset is the only variable. One generation costs about 55 minutes on a single consumer GPU. That's how a home lab out-iterates a budget.
  • Measured honesty. Every comparison runs through leak checks and pre-committed reading rules — which caught our own baseline lying twice. Painful. That's exactly why the wins we publish are bankable.

The scoreboard from our own library: a model that takes 0.4GB on disk beat a model 18× larger by more than 30 accuracy points — answering in about 8 milliseconds, offline, for free.

That honesty didn't stop at launch, either. The AI Hub shows the organizer's real accuracy on your own library against an 85% quality bar — measured on your machine, not in some lab.

What that buys you

Here's the part you actually feel:

  • It's already inside. The organizer ships in the installer — the whole app, brains included, adds up to about 1.6GB, smaller than one HD movie. On first launch, it's already working. Nothing to download first, nothing to configure.
  • It works offline. On a plane, on flaky hotel wifi, years from now. The internet is optional; your library isn't tied to ours.
  • It's free. Local AI never touches your credits. Organizing your library costs you exactly nothing. That's not a promo — it's the design.
  • It learns you. Every folder you confirm teaches your local AI, on your machine, in seconds. And when it's not sure, it says so: three best guesses, with honest confidence numbers.
  • Bigger, when you want it. For chat, you can add a bigger brain from inside the app: about 2.5GB to download, around 3.5GB of memory while it runs, pausable anytime — and adding it never blocks a single question: a small built-in assistant answers offline from day one.
  • No leftovers. Updates clean out old brains automatically, so your disk doesn't collect gigabyte-sized dust.

One small thing I'm proud of: you will never see a model name anywhere in the app. Just plain words — local, fast, yours.

The slow way, on purpose

The obvious path — rent — would have shipped faster. It would also have made Geneziz a little less yours, and a little less free to keep.

So we did it the slow way. Our machines, our pace, our bar. When Geneziz quietly suggests the right folder for that thing you saved three weeks ago, that suggestion was earned the hard way — and it costs you exactly nothing.

Your files. Your machine. Your intelligence. That's the whole idea.

Related Posts

Table of Contents