Local LLM - powered by Gemma

Moose can think on
your machine.

Want AI that runs on your own computer? Moose runs a real Gemma model right on your laptop and writes its answers there - unlimited and free. And it's a feature, not a requirement: go managed or bring your own OpenRouter key whenever you want frontier models doing the thinking.

Runs offline. Works on Windows & macOS.

Hi, Moose · Chat Gemma 4B · local
draft me a tight FAQ answer on "is local AI private"
on it. writing this with Gemma right here on your machine.
Running on your machine
$0
cost
0
tokens billed
100%
of the answer, on device
$0
Per query, always
The local model has no per-token bill. Chat, brief, and draft as much as you like.
Unlimited
Chats, with memory
Moose keeps your context and projects in mind across every conversation.
Offline
Works without a connection
Once downloaded, the local model runs on your computer and writes its answers there. You control your data in Settings.
Pick your size

A model that fits the laptop you already have.

Gemma comes in three sizes. Start small on an older machine, or load the big one on a workstation for the deepest analysis. The model downloads once, then runs fully offline - and you can switch any time.

Downloads in the background, runs offline forever after.
Uses your GPU when it can, your CPU when it can't.
Switch sizes whenever - the cost stays exactly the same.
Choose your model size
Light
Gemma 2B
8 GB RAM · ~3.5 GB
Balanced
Gemma 4B
16 GB RAM · ~5.4 GB
Max
Gemma 26B
32 GB RAM · ~17 GB
Gemma 4B≈1.2s a reply

The everyday default. Strong briefs and insights with no perceptible wait. Most people stay here.

Best for
Everyday work
Needs
16 GB RAM
You are in control

Local means local. Your work stays on your disk.

When you run the local model, the AI lives next to your files and writes its answers on your machine. You control what leaves it, and the controls are in Settings. The bigger cloud steps are ones you set up yourself: a managed plan, an OpenRouter key you add, or a CMS you connect to publish.

Stays on your machine
Every answer the local model writes
Your context, brand voice, and memory
Drafts, briefs, and your library
Search Console data you connect
What goes to the cloud

Things go out when you choose cloud: a managed plan where we run the frontier models for you, your own OpenRouter key, or a CMS you connect to publish. Each one is opt-in. You can review what leaves your machine in Settings.

A managed-plan task - our cloud, no keys to manage
A frontier task - through the OpenRouter key you added
A publish - only to a CMS you connected
Bring your own key

Want a frontier model in the loop? Plug in your key.

When a job calls for ChatGPT, Claude, or Gemini, connect your OpenRouter key and Moose routes that one task through it. You pay OpenRouter at the provider's rate - never a markup to us. Prefer zero setup? The managed plan runs frontier models for you, no key needed. Either way, the local model is always there, free.

No markup, ever. Your OpenRouter key, billed at provider rates.
Per-task routing. Send out only the jobs that need frontier power.
Your key stays local. Stored on your machine, never on ours.
Connected models
Hi, Moose - Local AI - Gemma 4
Always on · free · on your machine
Default
OpenRouter · your key
One key · billed by OpenRouter
Connected
Routes toChatGPTClaudeGemini+ more
Common questions

Local AI Privacy and Models FAQs

How does Hi, Moose protect my data?

When you use the local model, its answers are generated on your machine, and your context, drafts, brand voice, memory, library, and connected Search Console data stay there. The model runs alongside your files, and its downloaded model can operate fully offline. You control what leaves your machine, and you can adjust it in Settings.

What information leaves my computer when I use local AI?

With the local model, the answer is written on your computer. You control what leaves your computer, and you can adjust it in Settings. Information also goes out when you opt into a cloud action, such as using a managed plan, routing a task through an OpenRouter key you added, or publishing to a CMS you connected.

Can I use Hi, Moose AI without per-query costs?

Yes. The local Gemma model has no per-token or per-query bill. You can chat, create briefs, and draft without local usage limits. If you choose to route specific tasks through OpenRouter, those tasks are billed by OpenRouter at the provider's rate.

Which local AI model sizes are available?

Hi, Moose offers Gemma in three local sizes: Light Gemma 2B, Balanced Gemma 4B, and Max Gemma 26B. Light is designed for older machines with 8 GB RAM, Balanced is the everyday option with 16 GB RAM, and Max is for deeper analysis on workstations with 32 GB RAM. You can switch sizes whenever you need.

Can I connect ChatGPT, Claude, or Gemini to Hi, Moose?

Yes. Add your own OpenRouter key to route selected tasks to frontier models such as ChatGPT, Claude, or Gemini. Your key is stored on your machine, not on Hi, Moose servers. You can also choose a managed plan for frontier-model access without managing a key, while keeping the local model available for free.

Local LLM - powered by Gemma

Your laptop can do
the thinking.

Free to download, unlimited local use, and you control what leaves your machine. Add your OpenRouter key or a managed plan whenever you want frontier power.

Download for MacDownload for Windows