d3ro-voice/docs/map/04-desktop-app.md

35 KiB
Raw Permalink Blame History

04 — Desktop App (Electron) Map

Surface: apps/desktop Stack: Electron 33 + React 19 + MUI 7 + Vite (electron-vite) + better-sqlite3/drizzle + uiohook-napi + nut-js Source root: apps/desktop/src (main/, preload/, renderer/)


1. Process architecture

Layer Path Contents
Main src/main/ Services, IPC handlers, windows, bootstrap/lifecycle, DB
Preload src/preload/ index.ts exposes window.electronAPI; popup.ts exposes window.popupAPI
Renderer src/renderer/ React app: AppLayout + 7 pages + modals + 5 vanilla popups

Main entry src/main/index.ts: sets app name/AppUserModelId, disables GPU acceleration, EPIPE/uncaught handlers, registers d3ro-voice:// deep-link protocol (Supabase OAuth implicit + PKCE), single-instance lock, then bootstrap() + setupLifecycle().

Bootstrap src/main/bootstrap.ts: ordered BootstrapStep[] — logger, config, database (critical), license, create-windows (critical), tray, ipc-handlers (critical), custom-instructions, voice-commands, sound-effects, auto-launch, popup-preload, key-bindings, voice-mode, stt-warmup, llm-polling, meeting-summary-wiring, meeting-mode, cloud-sync, auto-update. Wires VoiceMode events to sound + history persistence, and subscribes to KeyBindingService triggered for the history-popup / command-popup actions (bootstrap.ts:159) — those two were hardcoded accelerators before and are now rebindable like everything else.


2. Main services (src/main/services/)

Singleton + EventEmitter pattern (getXService() accessors).

Core voice pipeline

Service Purpose
VoiceModeService Orchestrator: 9-state RecognitionState + 4-state AudioState, dual-condition flush, action queue. Events: session-started/completed/cancelled, transcription-update, audio-level, recognition/audio-state-changed, premium-llm-fallback, error
AudioCaptureService Mic PCM16 16kHz mono (bundled SoX on Windows, node-record-lpcm16 elsewhere). Spawns hidden (windowsHide); a missing SoX fails with the exact fix command
LocalSTTService faster-whisper Python sidecar manager (state machine, dual-flush, model download/cancel, background warm-up, live partial transcription). Connects over IPv4 loopback (getSidecarBaseUrl) and fails fast with an actionable message when the bundled engine or virtualenv is missing
KeyBindingService uiohook-napi global hooking for keyboard and mouse, driven by the @d3ro/core/keybinding contract: 6 rebindable actions (dictation, hands-free, command, caption, history-popup, command-popup), several bindings per action, structural reserved-combo checks. Events: triggered (in-process payload carries actionId, type (pressed/released), isDoublePress, holdMode, timestamp; the renderer-facing keybinding:triggered event is the narrower KeyBindingTriggeredEvent, keybinding.ts:1092), changed, error. globalShortcut is used only to mute the macOS system beep, and only for accelerators it registered itself. Mouse events cannot be suppressed by uiohook, so a bound button also performs its native action
TextInsertService Clipboard save→set→Ctrl+V→restore via nut-js
SoundEffectService Preloaded WAV feedback (start/stop/error/cancel/chime)

STT engine layer (services/stt/)

File Purpose
STTManager Dispatcher across local + 6 cloud providers, auto-fallback (events provider-changed, config-changed, fallback-to-local). transcribePartial/warmUpLocal route to the local engine only
types.ts ISTTDriver contract
audio-utils.ts pcmToWav, createProbeWav
`drivers/OpenAI Groq

LLM layer

Service Purpose
LocalLLMService Ollama REST (models, pull w/ progress, server start, NDJSON streaming). The 2026-09-22 runaway guard gives each request its own AbortController (external caller signals are relayed and detached on completion), requires done before an NDJSON stream succeeds, and clears incomplete streams. Generation/stream requests are bounded to 2048 tokens / 120 s; chat is bounded to 512 / 60 s. Ollama server spawn and polling are deduplicated, and lifecycle dispose terminates owned work.
PremiumLLMService Claude via Supabase llm-proxy, local fallback
OnlineLLMService JWT-authenticated .NET backend client
llm-prompts.ts SSOT for prompt resolution, placeholder substitution, and argument placement. resolveSystemPrompt (:118) maps an LLMAction to its base prompt and handles custom explicitly instead of dropping silently to refine. renderInstructionPrompt (:78) substitutes {{text}} / {{userPrompt}} / {{targetLanguage}} and warns by name for any placeholder left standing rather than letting it reach the model. buildInstructionInvocation (:101) decides where an instruction goes in processText(text, action, targetLanguage, customPrompt): the instruction becomes the system prompt and the transcript the text, except for instructions that spell out {{text}}, which keep the old meaning for backward compatibility. resolveTargetLanguage (:50) is the one place translate targets are decided (still English, see 11 GAP-LLM-01). All three LLM entry paths call the same functions — VoiceModeService (:838), ChainService (:196), and the LLM.PROCESS IPC handler (llm-handlers.ts:96) — so no caller re-implements the rules

Memory & knowledge

Service Purpose
HistoryService SQLite history CRUD/search/stats
DictionaryService Custom vocabulary CRUD/search + cloud sync hooks + JSON/CSV import/export (dictionary:import/export, save/open dialogs)
MemoService Memo tags over history (memo_tags)
RAGService Local RAG: nomic-embed-text embeddings, cosine search over rag_chunks
CustomInstructionService User LLM commands (5 built-ins)
VoiceCommandService Keyword → command rule matching
ChainService Multi-step LLM pipelines (LLMChain). Each step resolves its instruction through llm-prompts.ts (ChainService.ts:196); before that, chain steps sent placeholders through unsubstituted
ScreenContextService Active-window + selected-text context

Phase 10+ features

Service Purpose
CaptionService Live captions from system/loopback audio; caption overlay (events segment, state-changed, session-saved, error)
FileTranscriptionService Audio/video file → ffmpeg → 30s chunks → STT merge (events progress, complete, error, state-changed)
MeetingSummaryService Post-caption LLM summary
DictationTemplateService Field-by-field voice form filling
VoiceConversationService STT→LLM→TTS loop, 10-turn memory. Local conversations are single-flight and pass their request signal through to local LLM chat, so a voice cancel aborts its own active request rather than only changing UI/watchdog state.
TTSPlaybackService Platform TTS (macOS say, Windows SAPI), sentence queue
VoiceActionService Voice → LLM JSON action plan → OS execution (dangerous blocked)

Phase 12–15

Service Purpose
MeetingModeService Meeting recording: live transcript, timestamp memos, doc generation/export, diarization
MeetingDocTemplateService Meeting-doc templates (built-ins + CRUD)

Input intelligence (2026-09-21; Local Flow Intelligence extension 2026-09-22)

Service / file Purpose
InputTelemetryService Global input telemetry: keystroke/click/scroll counters, mouse travel (Manhattan sum of per-axis deltas), active time, per-hour-per-app buckets flushed every 5 s; typed-text learning via UIA snapshot diffs. It extends the existing telemetry store with Flow Radar (hourly activity density + chars + edit stability; a potential-focus score, not a real session), Edit Friction (chars/backspaces; edits per 100 chars), App Quality (per-app suggestion total/accepted/accept rate/average latency), Privacy Receipt (actual row counts and retention), and Smart Exclusion evidence. Raw key events, key codes and content streams are never retained; separately, opt-in learned text may be retained in typing_samples and personal_phrases under the stated receipt policy. Consent is opt-in (inputTelemetryEnabled, inputLearnTypedText).
UiaContextService Client for the sidecar UIA bridge (GET /uia/focus): focused text, caret rect/offset, IsPassword, IME composition state and hasSelection. The sidecar derives hasSelection only by comparing the TextPattern selection range Start/End endpoints: it does not call GetText on that range or add a separate selected-text payload (the existing focused-field text can still include a selection). Fail-closed on password/unavailable; transient failures back off for 1.5/3/6/12/30 s and permanent failures for 5 min.
SuggestionService Next-sentence ghost text: 600 ms debounce, then buildSuggestionPrompt → LocalLLMService.streamGenerate with a caller-owned AbortSignal → candidate ranking and accept (via TextInsertService). Presentation-active includes candidates plus generating, warmingUp and partialText; clear/dismiss aborts the active request, invalidates its generation token, clears TTL state, then emits cleared/hide so a spinner-only popup closes. A successful final candidate payload resets generating=false and partialText=null. App DNA extends phrase ranking with appName context and a 1.75× same-app bonus; Memory Decay uses a 30-day half-life plus frequency. Why This Suggestion exposes only local-model/local-memory provenance and continuation/related/phrase/appPhrase counts to the overlay, never raw evidence text. Instant Recall uses local-memory provenance without another model-budget spend only when the model is unavailable, raises a non-cancellation error, truly times out, or returns empty; it never publishes after dismiss, new typing, token/context mismatch or staleness. Its runaway limits are a 5 s minimum interval, 6 requests/min default (hard-config maximum 12), 3 candidates, 64 output tokens, 12-character regeneration growth and an 8 s timeout; app startup does not warm the model and suggestion requests use keep_alive: 2m.
global-input-hook.ts Ref-counted owner of the single process-wide uiohook hook so KeyBindingService and telemetry can both attach without one stopping the other.
utils/win32-foreground.ts Foreground window title/pid/exe/bounds via koffi FFI into user32/kernel32 (chosen over get-windows, which needs an install script this repo does not run).
sidecar/uia_bridge.py Windows UIA snapshot (uiautomation 2.0.29, comtypes) on a dedicated COM-initialised thread with a 1.5 s budget; never walks the UIA tree (Chrome/VS Code tree walks take 10-30 s). Warm calls measure 0-15 ms.

The existing input-telemetry-handlers and suggestion-handlers IPC extensions expose the receipt and the flow/suggestion summaries. Receipt reads or deletes report an IPC error when storage access fails; they do not return invented counts or claim a purge succeeded. The receipt is local-only: input_activity, typing_samples, and suggestions retain 30 days; personal_phrases has no age-based automatic expiry and is removed by individual deletion, delete-all, or consent withdrawal. Smart Exclusion never records password evidence and recommends only the current app after at least four observations, zero readable results, and a problematic ratio of at least 75%; its one-click action adds an exclusion and never auto-excludes. Shortcut Safety Audit is core auditKeyBindingMap applied by the Settings UI; it surfaces invalid/conflict issues while preserving the existing hold/double-press exceptions.

Account / infra / monetization

Service Purpose
ConfigService electron-store AppConfig (configGet/Set, defaults)
LicenseService Freemium tiers, quotas (daily_usage), activation, upgrade prompts
CloudSyncService Supabase auth + lifecycle: per-user DB switching, local-mode import on first sign-in, Realtime (ws transport) + 5-min heartbeat, device check-in, debounced flush/pull triggers. The sync itself lives in services/sync/
services/sync/ SyncEngine (backfill once per user DB → push outbox → pull by server-clock keyset cursor → tombstones → memo-tag reconcile), sync-outbox (sync_outbox/sync_state tables), sync-adapters (history, dictionary, meetings, meeting memos/documents, custom commands, templates), memo-tag-sync, supabase-sync-remote, device-registration, realtime-transport, settings-sync (mobile user_settings ↔ language/theme/defaultLLMAction/activeInstructionId), audio-sync (recording upload to the audio bucket, remote cleanup on delete), knowledge adapter (source-text chunks), transcript-sync (meeting segments from the desktop transcript), builtin-instruction-sync (preset prompt edits ↔ server preset rows). Services record changes with getCloudSyncService().pushOne/pushDelete(entity, id); pulled rows are applied without re-queuing
CloudSTTService Thin cloud STT wrapper over D3ROCloudDriver
UpdateService electron-updater (canonical Forgejo feed, channels, mandatory/full-vs-delta policy, staged rollout, restart dialog)
AutoLaunchService OS login-item auto-start
LoggerService electron-log wrapper + category loggers
upgrade / billing No in-app checkout. license-handlers.ts LICENSE.OPEN_BILLING opens billingUrl({ tier }) (web Payple); tier returns via license:tierChanged. Stripe payment-handlers.ts·CheckoutModal·payment:* IPC removed 2026-09-26

Ads (services/ads/)

File Purpose
AdMediationEngine Multi-ad mediation + header bidding
AdSettlementService Revenue settlement, withholding, payout ledger
BaseAdAdapter / UnavailableAdAdapter Adapter contract + fail-closed base
DirectHouseSponsorAdapter Real configurable adapter: bids/reports against an operator HTTPS endpointUrl (AdNetworkConfig.endpointUrl), validates creatives, fail-closed (adapter_not_configured) when unconfigured
9 placeholder adapters (AppLovin, Carbon, EthicalAds, GoogleAdManager, InMobi, Mintegral, Playwire, PubMatic, Unity) Extend UnavailableAdAdapter — registered, no live bids (provider_not_integrated)

3. IPC layer

Registry: src/main/ipc/index.ts calls 31 registerXHandlers() in fixed order. Channel SSOT: packages/core/src/ipc-channels.ts.

Handler Channel group(s)
ads-handlers ADS
audio-handlers AUDIO
caption-handlers CAPTION + SYSTEM_AUDIO
chain-handlers CHAIN
cloud-sync-handlers CLOUD_SYNC; forwards engine data-changed as app:dataChanged { type: 'cloud-sync', entities } so History/Dashboard/Dictionary/Commands/Meetings/Templates reload
config-handlers CONFIG
context-handlers CONTEXT
dictionary-handlers DICTIONARY
file-transcription-handlers FILE_TRANSCRIPTION
history-handlers HISTORY (incl. history:setFavorite, history:getAudio → local bytes or signed URL) + stats:getSummary
input-telemetry-handlers INPUT_TELEMETRY
instruction-handlers INSTRUCTION
keybinding-handlers KEYBINDING
license-handlers LICENSE
llm-handlers LLM + llm:premium:* + ONLINE_AUTH
meeting-doc-template-handlers MEETING_DOC_TEMPLATE
meeting-mode-handlers MEETING_MODE + MEETING_CHAT
meeting-summary-handlers MEETING_SUMMARY
memo-handlers MEMO
rag-handlers RAG
stt-handlers STT
suggestion-handlers SUGGESTION + POPUP_SUGGESTION
support-handlers SUPPORT
system-handlers SYSTEM
template-handlers DICTATION_TEMPLATE
voice-action-handlers VOICE_ACTION
voice-command-handlers VOICE_COMMAND
voice-conversation-handlers VOICE_CONVERSATION
voice-handlers VOICE
window-handlers WINDOW + SYSTEM.OPEN_EXTERNAL

The KEYBINDING group replaced the old per-action HOTKEY group. HOTKEY had 14 channels — a get/set pair per action plus three that were never implemented — so every new action meant new channels. KEYBINDING is 9 channels that take the action as a parameter: getMap, setBindings, resetAction, resetAll, validate, isEnabled, setEnabled, plus the triggered / changed events (packages/core/src/ipc-channels.ts:104). Adding an action now costs zero channels.

LLM.PROCESS normalizes at the IPC boundary. The handler runs buildInstructionInvocation itself when action === 'custom' with a customPrompt (llm-handlers.ts:94-108), so the renderer passes the raw instruction text and never duplicates the substitution or argument-placement rules. This is what makes VoiceModeService, ChainService, and LLM.PROCESS literally share one implementation. No channel or type changed for this; LLMProcessParams is unchanged.

Preload exposes window.electronAPI with 35 namespaces: platform, audio, config, voice, stt, keybinding, llm (incl. premium), history, dictionary, stats, window, system, instruction, app, memo, voiceCommand, context, chain, caption, license, fileTranscription, meetingSummary, dictationTemplate, rag, voiceAction, voiceConversation, meetingMode, meetingChat, meetingDocTemplate, cloudSync, onlineAuth, ads, support, inputTelemetry, suggestion. The keybinding bridge is 9 methods mirroring the channels above (src/preload/index.ts:323), replacing the 11-method hotkey bridge. Envelope: IPCResult<T> (success/error); app.onDataChanged is the global refresh channel.


4. Windows & popups

windows/WindowManager.ts creates 7 windows: main (borderless, custom TitleBar; macOS hiddenInset), recording-tip, result-popup, history-popup, command-popup, caption-overlay, suggestion-overlay. Injects popup theme CSS + i18n strings; 2-phase resize. windows/TrayManager.ts — tray icon + menu + double-click show.

Popup invariants (each shipped broken once — do not regress):

  • 팝업 HTML의 스크립트는 반드시 <script type="module">로 선언한다. Vite는 모듈 스크립트만 번들에 포함하므로 classic <script src="./script.js">는 dev에서만 로드되고 패키징 산출물에서는 파일 자체가 사라진다(오버레이가 정적 HTML로 멈춘 원인). scripts/ci/verify-desktop-renderer-bundles.mjs가 빌드 HTML이 참조하는 모든 로컬 asset의 존재를 검사한다.
  • 렌더러 로드 전의 webContents.send는 조용히 버려진다. 팝업 전송은 sendToPopupWindow를 쓰고, 이 함수가 did-finish-load까지 메시지를 보관했다가 전달한다. attachPopupLifecycle이 로드 상태 추적·테마 주입·팝업 렌더러 진단 로그를 한 곳에서 묶는다.
  • 팝업 표시는 presentPopup으로 통일한다(showInactive + topmost 재선언 + moveTop + webContents.invalidate). 한 번 hide()된 팝업이 두 번째 표시에서 z-order/repaint를 잃어 보이지 않던 문제를 막는다.
  • suggestion-overlay는 비축소 UIA 선택이 생기면 selection-active로 즉시 clear한다. 실제 표시 중일 때만 5초 이후에도 UIA를 주기 검증하고 mouse-up 뒤 120 ms에 다시 검증한다. 입력 focus가 사라져 available/editable이 false가 되거나 선택이 생기면 요청을 abort하고 hide한다. X는 renderer panel을 먼저 즉시 숨긴 뒤 main IPC가 BrowserWindow.hide()를 직접 호출하고 service dismiss를 수행한다. 늦은 결과는 generation token으로 재표시할 수 없다.

Vanilla popups (src/renderer/popups/):

Popup Purpose
recording-tip 9-bar waveform indicator, partial transcript
result-popup Transcription result + copy, auto-close with hover pause
history-popup Recent transcriptions; ↑↓/Enter/1-9/ESC. Opened by the history-popup action (default Ctrl+Shift+V, rebindable)
command-popup Command selection. Opened by the command-popup action (default Ctrl+Shift+C, rebindable)
caption-overlay Live caption overlay (font/opacity/maxLines)
suggestion-overlay Next-sentence ghost text: caret-anchored (anchorFloatingPanel), non-focusable, click-through unless suggestionOverlayInteractive; accept/next/dismiss come from global key bindings (the window never owns focus). Its narrow layout shows a one-line local provenance/count summary only, never raw memory evidence. Presentation includes candidate, generating, warm-up and partial-text states; its X hides the panel before main-process dismissal.

5. Renderer IA

Routing is state-based in AppLayout.tsx (Route union + NAV_ITEMS), no react-router.

Page Route Feature
DashboardPage dashboard Voice cockpit: hero, bento tiles, multi-engine hub (STT/LLM), telemetry, recent history, file drop
HistoryPage history History & memory timeline; search, tag filter, pagination, export/delete
DictionaryPage dictionary Custom vocabulary editor
CommandsPage commands Custom instructions + voice keyword rules + LLM chains + dictation templates
VoiceConversationPage conversation Duplex voice assistant (local pipeline vs OpenAI Realtime)
KnowledgeBasePage knowledge Local RAG: add/index docs, semantic query, reindex/remove
MeetingModePage meeting Meeting studio: live transcript, memos, doc generation/export, diarization

Modals/components: SettingsModal (tabs General/Audio/STT/LLM/Input/License/Cloud/About), LicenseModal, LicenseTab, CloudSyncSection, InputInsightsPanel (consent + suggestion policy + weekly insights + learned phrases), InputConsentPanel (receipt + current-app exclusion recommendation), InputInsightsView (flow, friction and app-quality summaries), OnboardingModal, UpgradePromptModal, ProBadge, TemplateSection, FileDropZone, OllamaGuideModal, CodexOAuthGuideModal, TitleBar, StatusBar, meeting components (9), voice-conversation, payment (CheckoutModal, checkout-flow.ts), support (SupportModal), ads (AdBanner, RewardedQuotaModal), shared cards.

Key-binding UI lives in components/keybinding/ (Keycap, KeyBindingPicker, KeyBindingField, translation-key), embedded in the Settings General tab (SettingsModal.tsx:239) — one field per action plus a global on/off switch. The picker offers both key recording and a searchable grouped dropdown (MUI Autocomplete over KEY_CATALOG, KeyBindingPicker.tsx:536). It replaced HotkeyRecordModal. renderer/utils/format-hotkey.ts is now a 17-line platform adapter only; key names, modifier glyphs, and join rules come from @d3ro/core/keybinding.

Hooks: useRealtimeConversation (OpenAI Realtime WebRTC), useLicenseState, useProFeature, useKeyBindingMap (subscribes to keybinding:changed; the dashboard renders the live dictation binding through BindingKeycaps), useInputInsights (telemetry + suggestion state + weekly summary + phrases; the Dashboard shows a weekly input card when collection is on).

DB schema (src/main/db/schema.ts, drizzle SQLite): history, dictionary, stats, memo_tags, daily_usage, rag_documents, rag_chunks, meeting_sessions, meeting_memos, meeting_documents, input_activity (hour × app counters), typing_samples, personal_phrases, suggestions.


6. Desktop status summary

  • Core dictation/LLM/history pipeline: implemented + tested. The vitest case count in apps/desktop is 1465 after input intelligence and the Local Flow Intelligence extension; playwright e2e is separate. Read the pass numbers together with the better-sqlite3 ABI the tree is built for (11 GAP-INFRA-06) — they are not comparable across configurations:
    • Host Node ABI (2026-09-21, before the LLM fix): 1311 / 1314 passing. The three failures are environment-dependent rather than regressions — two need a local sidecar venv or embedding server, one pins an error message that has since changed (11 GAP-QA-02). This configuration has not been re-measured since the LLM fix.
    • Electron ABI (2026-09-21, after input intelligence): 366 failed | 1042 passed (1408), against a clean-tree baseline of 366 failed | 994 passed (1360) in the same configuration — identical failure count, +48 passed, zero new failures. 365 of those 366 are tests/red/*.usecase.test.ts files dying at DB creation because of the ABI mismatch, not assertions.
    • Native-ABI mismatch run (2026-09-23, 1.5.0 release verification): 366 failed | 1099 passed (1465). The better-sqlite3 module is still built for Electron ABI 130 while the host Node is ABI 131, so this is the same mismatch configuration. The failure count is unchanged (365 ABI + 1 stale sidecar error-message assertion, 11 GAP-QA-02) and the passed count rose with the new tests, so there are zero new failures. Full strict/full suite and Electron GUI verification remain open gates, not evidence of green.
  • Input intelligence (2026-09-21): telemetry capture, weekly insights, next-sentence ghost text and phrase learning are implemented; real typing has verified capture → UIA snapshot → policy decisions, while overlay position/appearance, accept-insert, password blocking and weekly numbers remain manual gates (11 GAP-INPUT-01). The 24.7 s cold / 4.9 s warm figures are a historical latency diagnosis, not current behaviour. 2026-09-22 operating verification: at 19:14:12 app startup warmup plus keep_alive: 30m held gemma4:e4b resident through 19:44:12 at VRAM 3,226,342,521 bytes and context 4096. Windows GPU Engine PID sampling found no active Ollama compute then, so this was forced residency rather than infinite inference. Separately, 19:11:17–19:11:58 logs show automatic suggestions repeatedly generated under the old 900 ms / 12 per min / 5-candidate / 128-token / one-character-growth policy: a permitted burst, not proof of a single stuck request. The replacement guard removes boot warmup, uses keep_alive: 2m, and sets 600 ms debounce, 5 s minimum interval, 6 per min (hard maximum 12), 3 candidates, 64 tokens, 12-character growth and 8 s timeout. LocalLLMService adds per-request cancellation, bounded generate/stream/chat calls, required done frames and incomplete cleanup; voice cancellation is single-flight and reaches the request; server spawn/polling are deduplicated and disposed. Independent targeted verification passed 6 test files / 69 tests with 0 failures; changed code/tests ESLint and git diff --check exited 0. Raw Ollama proof: a cold bounded request hit the client hard timeout at 15.044 s then left /api/ps empty and /api/version recovered in 80 ms; explicit warmup returned HTTP 200 in 16.639 s; the subsequent num_predict=1, keep_alive='2m' request returned HTTP 200 in 553 ms with done:true, eval_count:1, response=OK, done_reason:length, and an observed /api/ps expiry of about 119.9 s. At 19:48:59 +09:00, with no intervening generate/unload/kill/retry, a single /api/ps returned HTTP 200 in 45.8 ms with {models:[]} and /api/version returned HTTP 200 in 7.3 ms with 0.32.13: raw API expiry/unload evidence only. This does not prove an app restart, GUI overlay or real automatic-typing path, so those remain [~] runtime gates (11 GAP-LLM-04, GAP-INPUT-06).
  • 2026-09-23 overlay lifecycle / Windows child-process audit: focused five test files passed 80 tests; core had 131 passing tests in an earlier independent verification; desktop typecheck/lint, Python py_compile, and git diff --check exited 0. Code and automated-test audit covers windowsHide:true on TTS PowerShell, VoiceAction cmd/PowerShell/general exec, audio-device and active-window PowerShell, plus all three ffmpeg paths; with existing SoX/STT/Ollama/sound-effect coverage, no Windows-capable desktop-main child-process call is known to be omitted. This is not Electron GUI runtime proof. Keep INPUT-07 and GAP-INPUT runtime gates [~] until an external-terminal run-desktop.bat restart confirms dismissal under generation/selection/focus loss and no cmd/PowerShell window recurrence for TTS, voice action, audio enumeration, screen context and file transcription.
  • Cross-platform packaging: Windows NSIS (signed, forceCodeSigning), macOS DMG/ZIP arm64 (ad-hoc signing); auto-update via canonical Forgejo feed with update policy (release/update-policy.json).
  • Local-first AI (SoX + faster-whisper sidecar + bundled Ollama) and cloud paths both present.
  • Local STT is packaged (1.3.0): electron-builder.yml extraResources copies sidecar-dist/sidecar → resources/sidecar and resources/ffmpeg → resources/ffmpeg; scripts/ci/verify-sidecar-bundle.mjs gates packaging. Build locally with npm --prefix apps/desktop run sidecar:setup && npm --prefix apps/desktop run sidecar:build. The sidecar stays in console mode so stdout/stderr reach the app log (UTF-8, line-buffered); a packaged sidecar must exist or startup fails loudly instead of silently falling back to a system Python.
  • All local engine URLs (LocalSTTService, LocalLLMService, RAGService, OnlineLLMService, STTManager) pass through src/main/utils/loopback.ts, which rewrites localhost to 127.0.0.1, because some Windows hosts resolve localhost to IPv6 only and local engines bind IPv4.
  • Meeting intelligence, RAG, voice conversation (local + Realtime), captions, file transcription: implemented.
  • LLM instruction prompts: fixed 2026-09-21 (9c2b4d4), not yet verified in a running app. Running a custom instruction inserted the instruction's own wording instead of the processed result. Two faults stacked: the instruction was passed in the text argument with the system-prompt argument left empty, and BASE_SYSTEM_PROMPTS has no custom key so resolution fell back to refine silently — the model polished the instruction and the transcript never reached it; separately, only {{text}} was substituted and none of the five built-in presets use it ({{targetLanguage}}, {{userPrompt}}, or no placeholder), so the substitution was a no-op from the day it was written. Introduced in fea923d (2026-04-05) and present in every release v0.1.0-alpha..v1.4.0 — the path never worked; this is not a regression. Plain actions (refine/summarize/grammar/expand) were unaffected and are now pinned by regression cases. The fix routes all three entry paths through llm-prompts.ts (see §2) and additionally corrects two things found alongside it: a voice shortcut naming an instruction was nullified by the defaultLLMAction === 'none' gate (VoiceModeService.ts:779), and the commands-page pipeline bench called llm.generate, which preload does not expose, so every run threw and the catch displayed the input as if it had succeeded — a fail-closed violation that is the reason the bug went unnoticed for five months (CommandsPage.tsx:180-205, now on llm.process with failures rendered as failures).
    • Verification limits — do not read this as verified. Unit tests pass (llm-prompts.test.ts 21, llm-handlers.test.ts 7, VoiceModeService.test.ts 22, ChainService.test.ts 5), and each of the four fixes was reverted individually to confirm the tests actually fail without it. npm run lint (apps/desktop scope) passes; tsconfig.check.json errors went 36 → 35 (the llm.generate error is gone) with no errors in the touched files. But there is no running-app run, and tests/red/{instruction,chain,voice,config}.usecase.test.ts — precisely the related paths — never executed because of the better-sqlite3 ABI mismatch. That range is neither passing nor failing; it is untested (11 GAP-LLM-02, GAP-INFRA-06).
  • Key bindings: implemented and verified on Windows. Every global shortcut now comes from one contract (@d3ro/core/keybinding) with multiple bindings per action, mouse-button support, and no hardcoded accelerators left in bootstrap.ts. A manual run on 2026-09-21 confirmed legacy migration (custom values preserved), 6 actions loaded, the uiohook keyboard and mouse hook active with zero boot errors, and multi-binding working; contract side is packages/core 117 tests GREEN with no type errors in the key-binding files (11 GAP-KEY-01 [x]). Two things remain open: KeyBindingService has no unit test of its own, and macOS/Linux mouse behavior is unconfirmed (11 GAP-KEY-02). The rewrite also fixed a dead hands-free double-press path, an order-dependent reserved-combo check, a globalShortcut.unregisterAll() that wiped the popup accelerators, and a setEnabled(true) that re-enabled hooking with an empty binding set.
  • The same pass fixed an unrelated pre-existing dashboard bug: caption.onStateChanged delivers { state }, but DashboardPage passed the whole object into setCaptionState, so the caption status readout never showed the right value (DashboardPage.tsx:148).
  • Ad mediation: DirectHouseSponsorAdapter performs real configurable REST bids; the other 9 adapters remain fail-closed stubs pending official SDKs (see 11-gap-backlog.md GAP-ADS-01/02).
  • Tier resolution now routes through @d3ro/core/entitlement (resolveEntitlement, normalizeEntitlementTier); useLicenseState.isPro includes pro_plus.
  • No TODO/FIXME markers found in src (grep clean). src/main/types/ is an empty directory.

7. Key file anchors

Thing Path
App entry / deep links src/main/index.ts
Bootstrap order src/main/bootstrap.ts
IPC registry src/main/ipc/index.ts
IPC channel SSOT packages/core/src/ipc-channels.ts
Key-binding contract SSOT packages/core/src/keybinding.ts (catalog, actions, validation, conflicts, formatting, parsing)
Input intelligence domain SSOT packages/core/src/input-intelligence.ts (key classification, typed delta, suggestion policy, overlay anchoring, flow/friction aggregation, phrase decay/ranking, provenance and exclusion recommendation)
Input intelligence services src/main/services/InputTelemetryService.ts, SuggestionService.ts, UiaContextService.ts, global-input-hook.ts
Foreground window FFI src/main/utils/win32-foreground.ts (koffi → user32/kernel32)
UIA bridge (sidecar) sidecar/uia_bridge.py + GET /uia/focus in sidecar/main.py
Input UI src/renderer/components/input-insights/{InputConsentPanel,InputInsightsView}.tsx, src/renderer/hooks/useInputInsights.ts, src/renderer/popups/suggestion-overlay/
Local Flow Intelligence tests tests/main/services/input-flow-domain.test.ts (9), input-flow-services.test.ts (12); targeted suite 21/21 at the documented handoff point
Key-binding service / IPC / UI src/main/services/KeyBindingService.ts, src/main/ipc/keybinding-handlers.ts, src/renderer/components/keybinding/
Preload API src/preload/index.ts
Windows src/main/windows/WindowManager.ts
Voice orchestrator src/main/services/VoiceModeService.ts
LLM prompt / placeholder SSOT src/main/services/llm-prompts.ts (shared by VoiceModeService, ChainService, ipc/llm-handlers.ts)
DB schema src/main/db/schema.ts
Renderer shell / routes src/renderer/components/AppLayout.tsx
Update feed SSOT src/main/update-feed.ts
Update policy SSOT release/update-policy.json + src/main/update-policy.ts
Path/loopback resolution src/main/utils/paths.ts, src/main/utils/loopback.ts
Sidecar source / packaging sidecar/main.py, scripts/setup-sidecar.mjs, scripts/build-sidecar.mjs, scripts/ci/verify-sidecar-bundle.mjs