Local LLM - powered by Gemma

Moose can think on
your machine.

Want AI that never leaves your desk? Moose runs a real Gemma model right on your laptop, entirely offline - unlimited, private, and free. And it's a feature, not a requirement: go managed or bring your own OpenRouter key whenever you want frontier models doing the thinking.

Runs offline. Works on Windows & macOS.

Hi, Moose · Chat Gemma 4B · local
draft me a tight FAQ answer on "is local AI private"
on it. running this through Gemma right here - nothing's going to the cloud.
Running on your machine
$0
cost
0
tokens sent out
100%
on device
$0
Per query, always
The local model has no per-token bill. Chat, brief, and draft as much as you like.
Unlimited
Chats, with memory
Moose keeps your context and projects in mind across every conversation.
Nothing
Leaves your desk
On the local model, prompts, context, and drafts stay on disk. Cloud is always your explicit choice.
Pick your size

A model that fits the laptop you already have.

Gemma comes in three sizes. Start small on an older machine, or load the big one on a workstation for the deepest analysis. The model downloads once, then runs fully offline - and you can switch any time.

Downloads in the background, runs offline forever after.
Uses your GPU when it can, your CPU when it can't.
Switch sizes whenever - the cost stays exactly the same.
Choose your model size
Light
Gemma 2B
8 GB RAM · ~3.5 GB
Balanced
Gemma 4B
16 GB RAM · ~5.4 GB
Max
Gemma 26B
32 GB RAM · ~17 GB
Gemma 4B≈1.2s a reply

The everyday default. Strong briefs and insights with no perceptible wait. Most people stay here.

Best for
Everyday work
Needs
16 GB RAM
Private by default

Local means local. Your work stays on your disk.

When you run the local model, the AI lives next to your files: prompts, drafts, and context never leave the machine. The cloud only comes into it through something you set up yourself: a managed plan, an OpenRouter key you add, or a CMS you connect to publish.

Stays on your machine
Every prompt you type to the local model
Your context, brand voice, and memory
Drafts, briefs, and your library
Search Console data you connect
Leaves only if you say so

Things go out when you choose cloud: a managed plan where we run the frontier models for you, your own OpenRouter key, or a CMS you connect to publish. Each one is opt-in.

A managed-plan task - our cloud, no keys to manage
A frontier task - through the OpenRouter key you added
A publish - only to a CMS you connected
Bring your own key

Want a frontier model in the loop? Plug in your key.

When a job calls for ChatGPT, Claude, or Gemini, connect your OpenRouter key and Moose routes that one task through it. You pay OpenRouter at the provider's rate - never a markup to us. Prefer zero setup? The managed plan runs frontier models for you, no key needed. Either way, the local model is always there, free.

No markup, ever. Your OpenRouter key, billed at provider rates.
Per-task routing. Send out only the jobs that need frontier power.
Your key stays local. Stored on your machine, never on ours.
Connected models
Hi, Moose - Local AI - Gemma 4
Always on · free · private
Default
OpenRouter · your key
One key · billed by OpenRouter
Connected
Routes toChatGPTClaudeGemini+ more
Common questions

Local AI Privacy and Models FAQs

How does Hi, Moose protect my data?

When you use the local model, your prompts, context, drafts, brand voice, memory, library, and connected Search Console data stay on your machine. The model runs alongside your files, and its downloaded model can operate fully offline. Cloud services are used only when you explicitly choose them.

What information leaves my computer when I use local AI?

Nothing leaves your computer for local-model tasks. Information is sent out only when you opt into a cloud action, such as using a managed plan, routing a task through an OpenRouter key you added, or publishing to a CMS you connected.

Can I use Hi, Moose AI without per-query costs?

Yes. The local Gemma model has no per-token or per-query bill. You can chat, create briefs, and draft without local usage limits. If you choose to route specific tasks through OpenRouter, those tasks are billed by OpenRouter at the provider's rate.

Which local AI model sizes are available?

Hi, Moose offers Gemma in three local sizes: Light Gemma 2B, Balanced Gemma 4B, and Max Gemma 26B. Light is designed for older machines with 8 GB RAM, Balanced is the everyday option with 16 GB RAM, and Max is for deeper analysis on workstations with 32 GB RAM. You can switch sizes whenever you need.

Can I connect ChatGPT, Claude, or Gemini to Hi, Moose?

Yes. Add your own OpenRouter key to route selected tasks to frontier models such as ChatGPT, Claude, or Gemini. Your key is stored on your machine, not on Hi, Moose servers. You can also choose a managed plan for frontier-model access without managing a key, while keeping the local model available for free.

Local LLM - powered by Gemma

Your laptop can do
the thinking.

Free to download, unlimited local use, private by default. Add your OpenRouter key or a managed plan whenever you want frontier power.

Download for MacDownload for Windows