Control your Mac with your voice. Completely offline.
Your voice, your screen, your commands — nothing leaves your machine. No cloud APIs. No API keys. No subscriptions.
Every cloud voice assistant makes the same tradeoff: you get convenience, they get your data. Your audio gets uploaded. Your screen gets sent to a server. Your commands get logged.
LocalClicky breaks that tradeoff. Everything runs on your hardware:
- Whisper.cpp — transcription, runs locally
- Ollama (qwen3, gemma4) — AI reasoning and vision, runs locally
- macOS say — text-to-speech, built into your Mac (default)
- PyAutoGUI — cursor and click control
No data leaves your machine by default. Not your voice. Not your screenshots. Not your commands.
Optionally, you can swap the built-in say voice for Fish Audio's free S2.1-pro TTS model for more natural-sounding speech — this does send the response text to Fish Audio's API. See Text-to-speech engine below.
- Sits in the menubar — no Dock icon, stays out of the way
- Say "Hey Jarvis" → starts a session — stays active until you say goodbye
- Voice Activity Detection — auto-stops recording when you stop talking (no fixed timeout)
- Sees your screen on demand — vision model (gemma4:e4b) takes a screenshot when needed
- Describes video clips — samples frames from a video file and summarizes what's in it
- Moves your cursor and clicks based on what it sees on screen
- Controls your Mac: open/quit apps, adjust volume, control Spotify, manage files, run shell commands, inject JS into Chrome
- Edits videos: trim, mute, merge, speed up, resize, add text — all via ffmpeg, no upload
- Creates reminders with natural language dates
- Multi-round tool calling — runs commands, checks results, confirms or retries
- Conversation memory across the session (last 10 exchanges)
- Session mode — chain commands back-to-back without repeating the wake word
| Icon | State |
|---|---|
| 🎙️ | Idle / ready |
| 👂 | Listening for "Computer" |
| 🔴 | Recording your voice |
| 🔄 | Transcribing |
| 🤔 | Thinking (Ollama) |
| 🔊 | Speaking response |
| Error |
brew install whisper-cpp
# Download the base English model
mkdir -p /opt/homebrew/share/whisper-cpp/models
curl -L -o /opt/homebrew/share/whisper-cpp/models/ggml-base.en.bin \
"https://huggingface.co/ggerganov/whisper.cpp/resolve/main/ggml-base.en.bin"brew install ollama
# Start Ollama
ollama serve
# Pull the models
ollama pull qwen3:8b # command model — tool calling, Mac control
ollama pull gemma4:e4b # vision model — sees your screen when neededbrew install ffmpegRequired for any video editing commands. Skip if you don't need video editing.
cd PyClicky
python3 -m venv venv
source venv/bin/activate
pip install -r requirements.txt
python -c "import openwakeword; openwakeword.utils.download_models()"pip install webrtcvad-wheelsWithout this, recording falls back to a 30-second hard cap instead of stopping when you stop talking.
cd PyClicky
source venv/bin/activate
ollama serve & # if not already running
python main.pyThe app appears in your menubar. No Dock icon.
LocalClicky needs three macOS permissions for the python3 binary inside your venv:
/path/to/PyClicky/venv/bin/python3
| Permission | Why | Where to grant |
|---|---|---|
| Microphone | Voice recording | Prompted automatically on first run |
| Screen Recording | Screenshot for vision | System Settings → Privacy & Security → Screen Recording |
| Accessibility | Cursor movement & clicks | System Settings → Privacy & Security → Accessibility |
Tip: If
python3is not selectable in the file picker, add Terminal instead — Python inherits Terminal's permissions when launched from it.
Say "Hey Jarvis" — the icon turns 🔴 and recording starts. When you stop talking, it automatically processes your command and responds.
After responding, it stays active and listens for your next command immediately — no need to say "Computer" again.
Say "bye", "goodbye", "stop listening", "go to sleep", or "that's all" — the assistant says goodbye and returns to wake word mode.
The session also auto-expires after 25 seconds of silence.
| You say | What happens |
|---|---|
| "Open Spotify and play hip hop" | Opens Spotify, searches and plays |
| "Set Spotify volume to 30 percent" | AppleScript sets Spotify's internal volume |
| "Set volume to 50 percent" | Sets macOS system volume |
| "Click the notification bell" | Takes screenshot, finds the bell, clicks it |
| "What's on my screen?" | Takes screenshot, describes what it sees |
| "Create a reminder to call John tomorrow at 9am" | Creates reminder in macOS Reminders |
| "Open a new tab in Chrome" | AppleScript opens a new Chrome tab |
| "Play next track" | AppleScript skips to next Spotify track |
| "Make a folder called Projects on my Desktop" | mkdir ~/Desktop/Projects |
| "What is the capital of France?" | Answers directly, no tools needed |
| "Trim the video on my desktop from 10 seconds to 30 seconds" | ffmpeg cuts the clip, saves to Desktop |
| "Mute the audio in intro dot mp4" | ffmpeg strips audio track |
| "Speed up the video to 2x" | ffmpeg applies setpts + atempo filters |
| "Merge video dot mp4 and clip dot mp4" | ffmpeg concat filter, saves to Desktop |
| "What's in the vacation video?" | Finds the file, samples 6 frames, describes the content |
| "Describe intro dot mp4" | Samples frames from the named file and summarizes it |
When you ask to click or find something, the assistant calls look_at_screen — it takes a clean screenshot, sends it to the vision model (gemma4:e4b), and gets back a bounding box for the target element. The center of that box is computed and clicked automatically.
The model decides on its own when it needs to see the screen — you don't have to phrase commands any special way.
Wake word ("Computer")
↓
AudioRecorder.start() ← opens sounddevice InputStream
↓ (VAD auto-stop on silence, 30s hard cap)
AudioRecorder.stop() → WAV file
↓
WhisperTranscriber.transcribe() → runs whisper-cli → transcript text
↓
Dismissal check ("bye" etc.) → end session / OllamaClient.chat()
↓
OllamaClient.chat() — always qwen3:8b with think mode + tools:
├─ run_shell_command → zsh → output
├─ query_system → read-only zsh → output
├─ look_at_screen → screencapture → gemma4:e4b → [CLICK:x1,y1,x2,y2]
├─ describe_video → ffmpeg samples frames → gemma4:e4b → summary
└─ create_reminder → Python builds correct AppleScript → osascript
(up to 5 tool rounds, streaming)
↓
CursorControl.extract_action() → parse [CLICK/POINT/RCLICK:x1,y1,x2,y2]
CursorControl.execute() → compute center → pyautogui moves/clicks
↓
SpeechOutput.speak() → macOS `say` (default) or Fish Audio API speaks the response
↓
Session active: wait 0.4s → start recording again
Session idle 25s: return to WakeWordDetector
PyClicky/
├── main.py # rumps menubar app — icons, menu, state display
├── companion.py # state machine — session management, full pipeline
├── ollama_client.py # qwen3 with tools, gemma4 vision via look_at_screen
├── wake_word.py # offline wake word via openWakeWord (hey_jarvis pretrained model)
├── audio_recorder.py # sounddevice mic capture + VAD silence detection → WAV
├── whisper_transcriber.py # calls whisper-cli subprocess, returns transcript
├── screen_capture.py # screencapture → resize to 1280px → base64 JPEG
├── video_analyzer.py # ffmpeg/ffprobe → samples evenly-spaced frames → base64 JPEGs
├── cursor_control.py # parses [CLICK/POINT/RCLICK:x1,y1,x2,y2], clicks center
├── speech_output.py # TTS: macOS `say` (default) or Fish Audio API, switched via TTS_ENGINE
├── shell_executor.py # zsh subprocess runner, cwd=~
└── requirements.txt
Edit ollama_client.py:
VISION_MODEL = "gemma4:e4b" # called by look_at_screen tool for visual tasks
COMMAND_MODEL = "qwen3:8b" # main model — tool calling, reasoning, Mac controlThe command model must support reliable tool calling. The vision model must be multimodal.
| Vision | Command | Notes |
|---|---|---|
gemma4:e4b |
qwen3:8b |
Default — good balance of speed and capability |
gemma4:e4b |
qwen3:14b |
Better reasoning, needs ~16GB RAM |
gemma4:27b |
qwen3:8b |
Better vision accuracy, needs ~32GB RAM |
qwen2.5vl:7b |
qwen3:8b |
Alternative vision model |
Edit wake_word.py:
# Use a different pretrained model (e.g. "alexa", "hey_mycroft"):
WAKE_MODEL = "hey_jarvis"
# Point to a custom trained .onnx or .tflite file instead:
WAKE_MODEL_PATH = "/path/to/your/computer.onnx" # overrides WAKE_MODEL when set
# Lower = more sensitive (more false positives), higher = stricter:
DETECTION_THRESHOLD = 0.5To train a custom "computer" model, follow the
openWakeWord training guide,
then set WAKE_MODEL_PATH to the output .onnx file.
Edit companion.py:
SESSION_IDLE_TIMEOUT = 25.0 # seconds of silence before returning to wake word modeEdit screen_capture.py:
MAX_WIDTH = 1280 # resize screenshot to this width before sending to vision model
JPEG_QUALITY = 75 # compression qualityLower MAX_WIDTH = faster responses, slightly less visual detail. Higher = more detail, larger payload.
Edit ollama_client.py:
OLLAMA_URL = "http://localhost:11434/api/chat"By default, LocalClicky speaks with the built-in macOS say command — fully offline. You can switch to Fish Audio's free S2.1-pro model for a more natural voice by setting a few variables in .env:
# .env
TTS_ENGINE=fish # "say" (default) or "fish"
FISH_AUDIO_API_KEY=your_api_key_here
FISH_AUDIO_MODEL=s2.1-pro-free # free tier model id
FISH_AUDIO_REFERENCE_ID= # optional: use a specific Fish Audio voiceGet a free API key from fish.audio — see their S2.1-pro free API announcement for details.
Notes:
- If
TTS_ENGINE=fishbut noFISH_AUDIO_API_KEYis set, it automatically falls back tosay. - If a Fish Audio request fails at runtime (network issue, bad key, etc.), it also falls back to
sayfor that utterance rather than staying silent. - Using Fish Audio sends the assistant's spoken responses to their API — this is the one opt-in exception to LocalClicky's fully-offline design.
"No speech detected" every time
- Check microphone permission
- Speak louder or closer to the mic
- Whisper model path may be wrong — check logs for the
model:line
Recording never stops / runs too long
- Install webrtcvad:
pip install webrtcvad-wheelsfor VAD silence detection - Without it, recording stops after 30 seconds
Screenshot always fails
- Grant Screen Recording to Terminal in System Settings → Privacy & Security
- Test:
screencapture -x -t jpg /tmp/test.jpg && echo OK
Cursor doesn't move
- Grant Accessibility to Terminal in System Settings → Privacy & Security → Accessibility
Wake word never triggers
- Wake word detection runs fully offline via openWakeWord — no internet needed
- Default keyword is "hey Jarvis" (not "Computer") — say that phrase to trigger
- To use a different keyword, change
WAKE_MODELinwake_word.py(see Configuration) - Check logs for
WAKE triggered:lines; lowerDETECTION_THRESHOLDif it's not firing - Speak clearly and at normal pace — very fast or whispered speech may score below threshold
Mic error when OBS or Zoom is running
- The app will retry 5 times automatically
- If it still fails, close the other app briefly then restart the session
Model says "I can't see your screen"
- Ensure Screen Recording permission is granted
- Try rephrasing: "look at my screen and click..."
Ollama 400 error
- Check
ollama list— ensure both models are pulled - Restart Ollama:
ollama serve
"Too many steps" response
- The model hit the 5-round tool call limit
- Check shell_executor logs for the underlying command error
- macOS 12+
- Python 3.11+
- Homebrew
- ~8GB RAM free (for both models)
- Ollama running locally
| Package | Purpose |
|---|---|
rumps |
macOS menubar app framework |
sounddevice |
Mic input stream |
soundfile |
Write WAV files |
numpy |
Audio buffer manipulation |
httpx |
Streaming HTTP to Ollama |
openwakeword |
Offline wake word detection |
pyautogui |
Cursor movement and clicks |
Pillow |
Screenshot resize |
webrtcvad-wheels |
Voice activity detection (optional) |
LocalClicky is early. Meaningful areas to improve:
- Custom "computer" wake word — train a personal openWakeWord model using the training guide and swap
WAKE_MODEL_PATHinwake_word.py - App-specific skills — context-aware commands for Terminal, Xcode, Figma, VS Code
- Packaging — proper
.appbundle so users don't need to run from terminal - Windows / Linux ports — the core pipeline is cross-platform; the menubar layer isn't
- Better click accuracy — the vision model (gemma4) has limited spatial precision; a GUI-specific model would help significantly
If you want to work on any of these, open an issue first. PRs welcome.
MIT