Private by default
Keep everyday conversations and inference on-device. Cloud services are an optional handoff for work that genuinely needs them.
A quiet, local-first companion for Windows. She stays out of the way until you need a hand—then helps you think, find, and get things done.
The idea
Josephine is envisioned as a collaborative advisor—not another tab demanding attention. A lightweight local model keeps everyday help close, while optional cloud handoffs can take on tasks that need more compute.
Keep everyday conversations and inference on-device. Cloud services are an optional handoff for work that genuinely needs them.
A quiet tray presence, quick activation, and automatic sleep help keep the assistant responsive without wasting laptop resources.
Josephine is designed to connect natural-language requests to real Windows actions, apps, research, and focused-work routines.
What Jo is designed to do
Summon the dashboard with a global shortcut. A compact glass-style overlay gives clear listening, processing, and network feedback without taking over the screen.
Push-to-talk voice input keeps the microphone under your control. Local speech recognition can turn requests into text, with screen and clipboard context available when you choose.
A small quantized model handles routing and everyday conversation. Larger, network-dependent tasks can be handed off to a cloud provider with visible status and cancellation.
Launch verified apps, organize a deep-work setup, adjust supported system controls, set reminders, or gather information from the web and research sources.
Under the hood
Each request moves through a small, understandable pipeline. The router chooses a local tool or model first, and makes network-dependent work visible.
Quiet entry point
User-controlled input
Choose the right path
Local actions or inference
HUD + offline voice
Online: local requests stay local; eligible heavy tasks can be handed off with clear status.
Offline: local features remain available; queued cloud work shows that it is waiting and can be cancelled.
The building blocks
Tray menu, overlay states, hotkey activation, idle sleep, and safe shutdown.
PyQt6 · pynput
Push-to-talk transcription, optional clipboard context, and foreground-window awareness.
Vosk / Whisper · pywin32
Quantized model inference, recent-turn context, intent classification, and controlled unload/reload.
llama-cpp-python · GGUF
Verified app launching, supported system operations, reminders, and explicit failure feedback.
Python · Windows APIs
Non-blocking local speech playback synchronized with the assistant's response.
Piper · sounddevice
Check connectivity before cloud work; show waiting state and let the user cancel.
HTTP client · network check
Made for modest hardware
The target is a 16 GB dual-channel laptop: use a quantized 3B model, avoid unnecessary background work, and release model memory after a short idle period.
Model acceleration and Windows hardware integrations depend on the target machine and will be validated during implementation.
Roadmap
Start with a dependable local core, then add the interface and senses around it.
Set up the Python application, run local GGUF inference, and establish a small command router.
Build the core
Add reliable app discovery and launching, then introduce carefully scoped system controls and reminders.
Connect useful tools
Wire in local speech recognition and offline text-to-speech without blocking the interface.
Make it hands-free
Build the tray-first overlay, connect its states to the backend, and add hotkey activation.
Bring Jo to life
Measure resource use, unload the model when idle, harden fallbacks, and package for Windows startup.
Ready for daily use
Build log
Follow the latest progress and notes from building Josephine.
Josephine is a work in progress—an experiment in making on-device AI feel personal, practical, and respectful of your attention.
More projects