Some checks failed
ci / 정본·보안·린트·타입·테스트 (push) Successful in 50s
ci / 모바일 린트·타입·Jest (push) Successful in 38s
ci / Supabase Edge Functions + Cloudflare Worker (push) Successful in 21s
ci / .NET API 서버 테스트 (push) Successful in 11s
deploy-site / deploy (push) Successful in 33s
ci / 워크스페이스 빌드 검증 (push) Has been cancelled
258 lines
35 KiB
Markdown
258 lines
35 KiB
Markdown
# 04 — Desktop App (Electron) Map
|
||
|
||
> Surface: `apps/desktop`
|
||
> Stack: Electron 33 + React 19 + MUI 7 + Vite (electron-vite) + better-sqlite3/drizzle + uiohook-napi + nut-js
|
||
> Source root: `apps/desktop/src` (`main/`, `preload/`, `renderer/`)
|
||
|
||
---
|
||
|
||
## 1. Process architecture
|
||
|
||
| Layer | Path | Contents |
|
||
|---|---|---|
|
||
| Main | `src/main/` | Services, IPC handlers, windows, bootstrap/lifecycle, DB |
|
||
| Preload | `src/preload/` | `index.ts` exposes `window.electronAPI`; `popup.ts` exposes `window.popupAPI` |
|
||
| Renderer | `src/renderer/` | React app: `AppLayout` + 7 pages + modals + 5 vanilla popups |
|
||
|
||
**Main entry** `src/main/index.ts`: sets app name/AppUserModelId, disables GPU acceleration, EPIPE/uncaught handlers, registers `d3ro-voice://` deep-link protocol (Supabase OAuth implicit + PKCE), single-instance lock, then `bootstrap()` + `setupLifecycle()`.
|
||
|
||
**Bootstrap** `src/main/bootstrap.ts`: ordered `BootstrapStep[]` — logger, config, **database (critical)**, license, create-windows (critical), tray, **ipc-handlers (critical)**, custom-instructions, voice-commands, sound-effects, auto-launch, popup-preload, key-bindings, voice-mode, stt-warmup, llm-polling, meeting-summary-wiring, meeting-mode, cloud-sync, auto-update. Wires VoiceMode events to sound + history persistence, and subscribes to `KeyBindingService` `triggered` for the `history-popup` / `command-popup` actions (`bootstrap.ts:159`) — those two were hardcoded accelerators before and are now rebindable like everything else.
|
||
|
||
---
|
||
|
||
## 2. Main services (`src/main/services/`)
|
||
|
||
Singleton + `EventEmitter` pattern (`getXService()` accessors).
|
||
|
||
### Core voice pipeline
|
||
| Service | Purpose |
|
||
|---|---|
|
||
| `VoiceModeService` | Orchestrator: 9-state `RecognitionState` + 4-state `AudioState`, dual-condition flush, action queue. Events: session-started/completed/cancelled, transcription-update, audio-level, recognition/audio-state-changed, premium-llm-fallback, error |
|
||
| `AudioCaptureService` | Mic PCM16 16kHz mono (bundled SoX on Windows, node-record-lpcm16 elsewhere). Spawns hidden (`windowsHide`); a missing SoX fails with the exact fix command |
|
||
| `LocalSTTService` | faster-whisper Python sidecar manager (state machine, dual-flush, model download/cancel, background warm-up, live partial transcription). Connects over IPv4 loopback (`getSidecarBaseUrl`) and fails fast with an actionable message when the bundled engine or virtualenv is missing |
|
||
| `KeyBindingService` | uiohook-napi global hooking for **keyboard and mouse**, driven by the `@d3ro/core/keybinding` contract: 6 rebindable actions (dictation, hands-free, command, caption, history-popup, command-popup), several bindings per action, structural reserved-combo checks. Events: `triggered` (in-process payload carries `actionId`, `type` (`pressed`/`released`), `isDoublePress`, `holdMode`, `timestamp`; the renderer-facing `keybinding:triggered` event is the narrower `KeyBindingTriggeredEvent`, `keybinding.ts:1092`), `changed`, `error`. `globalShortcut` is used only to mute the macOS system beep, and only for accelerators it registered itself. Mouse events cannot be suppressed by uiohook, so a bound button also performs its native action |
|
||
| `TextInsertService` | Clipboard save→set→Ctrl+V→restore via nut-js |
|
||
| `SoundEffectService` | Preloaded WAV feedback (start/stop/error/cancel/chime) |
|
||
|
||
### STT engine layer (`services/stt/`)
|
||
| File | Purpose |
|
||
|---|---|
|
||
| `STTManager` | Dispatcher across local + 6 cloud providers, auto-fallback (events provider-changed, config-changed, fallback-to-local). `transcribePartial`/`warmUpLocal` route to the local engine only |
|
||
| `types.ts` | `ISTTDriver` contract |
|
||
| `audio-utils.ts` | `pcmToWav`, `createProbeWav` |
|
||
| `drivers/OpenAI|Groq|Deepgram|AssemblyAI|Google|Custom|D3ROCloud` | Provider drivers; `D3ROCloudDriver` uses Supabase access token |
|
||
|
||
### LLM layer
|
||
| Service | Purpose |
|
||
|---|---|
|
||
| `LocalLLMService` | Ollama REST (models, pull w/ progress, server start, NDJSON streaming). The 2026-09-22 runaway guard gives each request its own `AbortController` (external caller signals are relayed and detached on completion), requires `done` before an NDJSON stream succeeds, and clears incomplete streams. Generation/stream requests are bounded to 2048 tokens / 120 s; chat is bounded to 512 / 60 s. Ollama server spawn and polling are deduplicated, and lifecycle dispose terminates owned work. |
|
||
| `PremiumLLMService` | Claude via Supabase `llm-proxy`, local fallback |
|
||
| `OnlineLLMService` | JWT-authenticated .NET backend client |
|
||
| `llm-prompts.ts` | **SSOT for prompt resolution, placeholder substitution, and argument placement.** `resolveSystemPrompt` (`:118`) maps an `LLMAction` to its base prompt and handles `custom` explicitly instead of dropping silently to `refine`. `renderInstructionPrompt` (`:78`) substitutes `{{text}}` / `{{userPrompt}}` / `{{targetLanguage}}` and **warns by name** for any placeholder left standing rather than letting it reach the model. `buildInstructionInvocation` (`:101`) decides where an instruction goes in `processText(text, action, targetLanguage, customPrompt)`: the instruction becomes the **system prompt** and the transcript the **text**, except for instructions that spell out `{{text}}`, which keep the old meaning for backward compatibility. `resolveTargetLanguage` (`:50`) is the one place translate targets are decided (still `English`, see `11` GAP-LLM-01). **All three LLM entry paths call the same functions** — `VoiceModeService` (`:838`), `ChainService` (`:196`), and the `LLM.PROCESS` IPC handler (`llm-handlers.ts:96`) — so no caller re-implements the rules |
|
||
|
||
### Memory & knowledge
|
||
| Service | Purpose |
|
||
|---|---|
|
||
| `HistoryService` | SQLite history CRUD/search/stats |
|
||
| `DictionaryService` | Custom vocabulary CRUD/search + cloud sync hooks + JSON/CSV import/export (`dictionary:import`/`export`, save/open dialogs) |
|
||
| `MemoService` | Memo tags over history (`memo_tags`) |
|
||
| `RAGService` | Local RAG: `nomic-embed-text` embeddings, cosine search over `rag_chunks` |
|
||
| `CustomInstructionService` | User LLM commands (5 built-ins) |
|
||
| `VoiceCommandService` | Keyword → command rule matching |
|
||
| `ChainService` | Multi-step LLM pipelines (LLMChain). Each step resolves its instruction through `llm-prompts.ts` (`ChainService.ts:196`); before that, chain steps sent placeholders through unsubstituted |
|
||
| `ScreenContextService` | Active-window + selected-text context |
|
||
|
||
### Phase 10+ features
|
||
| Service | Purpose |
|
||
|---|---|
|
||
| `CaptionService` | Live captions from system/loopback audio; caption overlay (events segment, state-changed, session-saved, error) |
|
||
| `FileTranscriptionService` | Audio/video file → ffmpeg → 30s chunks → STT merge (events progress, complete, error, state-changed) |
|
||
| `MeetingSummaryService` | Post-caption LLM summary |
|
||
| `DictationTemplateService` | Field-by-field voice form filling |
|
||
| `VoiceConversationService` | STT→LLM→TTS loop, 10-turn memory. Local conversations are single-flight and pass their request signal through to local LLM chat, so a voice cancel aborts its own active request rather than only changing UI/watchdog state. |
|
||
| `TTSPlaybackService` | Platform TTS (macOS `say`, Windows SAPI), sentence queue |
|
||
| `VoiceActionService` | Voice → LLM JSON action plan → OS execution (dangerous blocked) |
|
||
|
||
### Phase 12–15
|
||
| Service | Purpose |
|
||
|---|---|
|
||
| `MeetingModeService` | Meeting recording: live transcript, timestamp memos, doc generation/export, diarization |
|
||
| `MeetingDocTemplateService` | Meeting-doc templates (built-ins + CRUD) |
|
||
|
||
### Input intelligence (2026-09-21; Local Flow Intelligence extension 2026-09-22)
|
||
| Service / file | Purpose |
|
||
|---|---|
|
||
| `InputTelemetryService` | Global input telemetry: keystroke/click/scroll counters, mouse travel (Manhattan sum of per-axis deltas), active time, per-hour-per-app buckets flushed every 5 s; typed-text learning via UIA snapshot diffs. It extends the existing telemetry store with Flow Radar (hourly activity density + chars + edit stability; a potential-focus score, **not** a real session), Edit Friction (chars/backspaces; edits per 100 chars), App Quality (per-app suggestion total/accepted/accept rate/average latency), Privacy Receipt (actual row counts and retention), and Smart Exclusion evidence. Raw key events, key codes and content streams are never retained; separately, opt-in learned text may be retained in `typing_samples` and `personal_phrases` under the stated receipt policy. Consent is opt-in (`inputTelemetryEnabled`, `inputLearnTypedText`). |
|
||
| `UiaContextService` | Client for the sidecar UIA bridge (`GET /uia/focus`): focused text, caret rect/offset, `IsPassword`, IME composition state and `hasSelection`. The sidecar derives `hasSelection` only by comparing the TextPattern selection range Start/End endpoints: it does not call `GetText` on that range or add a separate selected-text payload (the existing focused-field text can still include a selection). Fail-closed on password/unavailable; transient failures back off for 1.5/3/6/12/30 s and permanent failures for 5 min. |
|
||
| `SuggestionService` | Next-sentence ghost text: 600 ms debounce, then `buildSuggestionPrompt` → `LocalLLMService.streamGenerate` with a caller-owned `AbortSignal` → candidate ranking and accept (via `TextInsertService`). Presentation-active includes candidates plus `generating`, `warmingUp` and `partialText`; clear/dismiss aborts the active request, invalidates its generation token, clears TTL state, then emits `cleared`/hide so a spinner-only popup closes. A successful final candidate payload resets `generating=false` and `partialText=null`. App DNA extends phrase ranking with `appName` context and a 1.75× same-app bonus; Memory Decay uses a 30-day half-life plus frequency. Why This Suggestion exposes only `local-model`/`local-memory` provenance and continuation/related/phrase/appPhrase **counts** to the overlay, never raw evidence text. Instant Recall uses local-memory provenance without another model-budget spend only when the model is unavailable, raises a non-cancellation error, truly times out, or returns empty; it never publishes after dismiss, new typing, token/context mismatch or staleness. Its runaway limits are a 5 s minimum interval, 6 requests/min default (hard-config maximum 12), 3 candidates, 64 output tokens, 12-character regeneration growth and an 8 s timeout; app startup does not warm the model and suggestion requests use `keep_alive: 2m`. |
|
||
| `global-input-hook.ts` | Ref-counted owner of the single process-wide `uiohook` hook so `KeyBindingService` and telemetry can both attach without one stopping the other. |
|
||
| `utils/win32-foreground.ts` | Foreground window title/pid/exe/bounds via `koffi` FFI into user32/kernel32 (chosen over `get-windows`, which needs an install script this repo does not run). |
|
||
| `sidecar/uia_bridge.py` | Windows UIA snapshot (`uiautomation` 2.0.29, comtypes) on a dedicated COM-initialised thread with a 1.5 s budget; never walks the UIA tree (Chrome/VS Code tree walks take 10-30 s). Warm calls measure 0-15 ms. |
|
||
|
||
The existing `input-telemetry-handlers` and `suggestion-handlers` IPC extensions expose the receipt and the flow/suggestion summaries. Receipt reads or deletes report an IPC error when storage access fails; they do not return invented counts or claim a purge succeeded. The receipt is local-only: `input_activity`, `typing_samples`, and `suggestions` retain 30 days; `personal_phrases` has no age-based automatic expiry and is removed by individual deletion, delete-all, or consent withdrawal. Smart Exclusion never records password evidence and recommends only the current app after at least four observations, zero readable results, and a problematic ratio of at least 75%; its one-click action adds an exclusion and never auto-excludes. Shortcut Safety Audit is core `auditKeyBindingMap` applied by the Settings UI; it surfaces invalid/conflict issues while preserving the existing hold/double-press exceptions.
|
||
|
||
### Account / infra / monetization
|
||
| Service | Purpose |
|
||
|---|---|
|
||
| `ConfigService` | electron-store `AppConfig` (`configGet/Set`, defaults) |
|
||
| `LicenseService` | Freemium tiers, quotas (`daily_usage`), activation, upgrade prompts |
|
||
| `CloudSyncService` | Supabase auth + lifecycle: per-user DB switching, local-mode import on first sign-in, Realtime (`ws` transport) + 5-min heartbeat, device check-in, debounced flush/pull triggers. The sync itself lives in `services/sync/` |
|
||
| `services/sync/` | `SyncEngine` (backfill once per user DB → push outbox → pull by server-clock keyset cursor → tombstones → memo-tag reconcile), `sync-outbox` (`sync_outbox`/`sync_state` tables), `sync-adapters` (history, dictionary, meetings, meeting memos/documents, custom commands, templates), `memo-tag-sync`, `supabase-sync-remote`, `device-registration`, `realtime-transport`. Services record changes with `getCloudSyncService().pushOne/pushDelete(entity, id)`; pulled rows are applied without re-queuing |
|
||
| `CloudSTTService` | Thin cloud STT wrapper over `D3ROCloudDriver` |
|
||
| `UpdateService` | electron-updater (canonical Forgejo feed, channels, mandatory/full-vs-delta policy, staged rollout, restart dialog) |
|
||
| `AutoLaunchService` | OS login-item auto-start |
|
||
| `LoggerService` | electron-log wrapper + category loggers |
|
||
| upgrade / billing | No in-app checkout. `license-handlers.ts` `LICENSE.OPEN_BILLING` opens `billingUrl({ tier })` (web Payple); tier returns via `license:tierChanged`. Stripe `payment-handlers.ts`·`CheckoutModal`·`payment:*` IPC removed 2026-09-26 |
|
||
|
||
### Ads (`services/ads/`)
|
||
| File | Purpose |
|
||
|---|---|
|
||
| `AdMediationEngine` | Multi-ad mediation + header bidding |
|
||
| `AdSettlementService` | Revenue settlement, withholding, payout ledger |
|
||
| `BaseAdAdapter` / `UnavailableAdAdapter` | Adapter contract + fail-closed base |
|
||
| `DirectHouseSponsorAdapter` | **Real configurable adapter**: bids/reports against an operator HTTPS `endpointUrl` (`AdNetworkConfig.endpointUrl`), validates creatives, fail-closed (`adapter_not_configured`) when unconfigured |
|
||
| 9 placeholder adapters (AppLovin, Carbon, EthicalAds, GoogleAdManager, InMobi, Mintegral, Playwire, PubMatic, Unity) | Extend `UnavailableAdAdapter` — registered, no live bids (`provider_not_integrated`) |
|
||
|
||
---
|
||
|
||
## 3. IPC layer
|
||
|
||
Registry: `src/main/ipc/index.ts` calls 31 `registerXHandlers()` in fixed order. Channel SSOT: `packages/core/src/ipc-channels.ts`.
|
||
|
||
| Handler | Channel group(s) |
|
||
|---|---|
|
||
| `ads-handlers` | ADS |
|
||
| `audio-handlers` | AUDIO |
|
||
| `caption-handlers` | CAPTION + SYSTEM_AUDIO |
|
||
| `chain-handlers` | CHAIN |
|
||
| `cloud-sync-handlers` | CLOUD_SYNC; forwards engine `data-changed` as `app:dataChanged` `{ type: 'cloud-sync', entities }` so History/Dashboard/Dictionary/Commands/Meetings/Templates reload |
|
||
| `config-handlers` | CONFIG |
|
||
| `context-handlers` | CONTEXT |
|
||
| `dictionary-handlers` | DICTIONARY |
|
||
| `file-transcription-handlers` | FILE_TRANSCRIPTION |
|
||
| `history-handlers` | HISTORY (incl. `history:setFavorite`) + `stats:getSummary` |
|
||
| `input-telemetry-handlers` | INPUT_TELEMETRY |
|
||
| `instruction-handlers` | INSTRUCTION |
|
||
| `keybinding-handlers` | KEYBINDING |
|
||
| `license-handlers` | LICENSE |
|
||
| `llm-handlers` | LLM + `llm:premium:*` + ONLINE_AUTH |
|
||
| `meeting-doc-template-handlers` | MEETING_DOC_TEMPLATE |
|
||
| `meeting-mode-handlers` | MEETING_MODE + MEETING_CHAT |
|
||
| `meeting-summary-handlers` | MEETING_SUMMARY |
|
||
| `memo-handlers` | MEMO |
|
||
| `rag-handlers` | RAG |
|
||
| `stt-handlers` | STT |
|
||
| `suggestion-handlers` | SUGGESTION + POPUP_SUGGESTION |
|
||
| `support-handlers` | SUPPORT |
|
||
| `system-handlers` | SYSTEM |
|
||
| `template-handlers` | DICTATION_TEMPLATE |
|
||
| `voice-action-handlers` | VOICE_ACTION |
|
||
| `voice-command-handlers` | VOICE_COMMAND |
|
||
| `voice-conversation-handlers` | VOICE_CONVERSATION |
|
||
| `voice-handlers` | VOICE |
|
||
| `window-handlers` | WINDOW + `SYSTEM.OPEN_EXTERNAL` |
|
||
|
||
The **`KEYBINDING`** group replaced the old per-action `HOTKEY` group. `HOTKEY` had 14 channels — a get/set pair per action plus three that were never implemented — so every new action meant new channels. `KEYBINDING` is 9 channels that take the action **as a parameter**: `getMap`, `setBindings`, `resetAction`, `resetAll`, `validate`, `isEnabled`, `setEnabled`, plus the `triggered` / `changed` events (`packages/core/src/ipc-channels.ts:104`). Adding an action now costs zero channels.
|
||
|
||
**`LLM.PROCESS` normalizes at the IPC boundary.** The handler runs `buildInstructionInvocation` itself when `action === 'custom'` with a `customPrompt` (`llm-handlers.ts:94-108`), so the renderer passes the **raw instruction text** and never duplicates the substitution or argument-placement rules. This is what makes `VoiceModeService`, `ChainService`, and `LLM.PROCESS` literally share one implementation. No channel or type changed for this; `LLMProcessParams` is unchanged.
|
||
|
||
Preload exposes **`window.electronAPI`** with 35 namespaces: `platform, audio, config, voice, stt, keybinding, llm (incl. premium), history, dictionary, stats, window, system, instruction, app, memo, voiceCommand, context, chain, caption, license, fileTranscription, meetingSummary, dictationTemplate, rag, voiceAction, voiceConversation, meetingMode, meetingChat, meetingDocTemplate, cloudSync, onlineAuth, ads, support, inputTelemetry, suggestion`. The `keybinding` bridge is 9 methods mirroring the channels above (`src/preload/index.ts:323`), replacing the 11-method `hotkey` bridge. Envelope: `IPCResult<T>` (success/error); `app.onDataChanged` is the global refresh channel.
|
||
|
||
---
|
||
|
||
## 4. Windows & popups
|
||
|
||
`windows/WindowManager.ts` creates 7 windows: main (borderless, custom TitleBar; macOS `hiddenInset`), recording-tip, result-popup, history-popup, command-popup, caption-overlay, suggestion-overlay. Injects popup theme CSS + i18n strings; 2-phase resize. `windows/TrayManager.ts` — tray icon + menu + double-click show.
|
||
|
||
**Popup invariants** (each shipped broken once — do not regress):
|
||
|
||
- 팝업 HTML의 스크립트는 반드시 `<script type="module">`로 선언한다. Vite는 모듈 스크립트만 번들에 포함하므로 classic `<script src="./script.js">`는 dev에서만 로드되고 패키징 산출물에서는 파일 자체가 사라진다(오버레이가 정적 HTML로 멈춘 원인). `scripts/ci/verify-desktop-renderer-bundles.mjs`가 빌드 HTML이 참조하는 모든 로컬 asset의 존재를 검사한다.
|
||
- 렌더러 로드 전의 `webContents.send`는 조용히 버려진다. 팝업 전송은 `sendToPopupWindow`를 쓰고, 이 함수가 `did-finish-load`까지 메시지를 보관했다가 전달한다. `attachPopupLifecycle`이 로드 상태 추적·테마 주입·팝업 렌더러 진단 로그를 한 곳에서 묶는다.
|
||
- 팝업 표시는 `presentPopup`으로 통일한다(`showInactive` + topmost 재선언 + `moveTop` + `webContents.invalidate`). 한 번 `hide()`된 팝업이 두 번째 표시에서 z-order/repaint를 잃어 보이지 않던 문제를 막는다.
|
||
- `suggestion-overlay`는 비축소 UIA 선택이 생기면 `selection-active`로 즉시 clear한다. 실제 표시 중일 때만 5초 이후에도 UIA를 주기 검증하고 mouse-up 뒤 120 ms에 다시 검증한다. 입력 focus가 사라져 `available`/`editable`이 false가 되거나 선택이 생기면 요청을 abort하고 hide한다. X는 renderer panel을 먼저 즉시 숨긴 뒤 main IPC가 `BrowserWindow.hide()`를 직접 호출하고 service dismiss를 수행한다. 늦은 결과는 generation token으로 재표시할 수 없다.
|
||
|
||
Vanilla popups (`src/renderer/popups/`):
|
||
| Popup | Purpose |
|
||
|---|---|
|
||
| `recording-tip` | 9-bar waveform indicator, partial transcript |
|
||
| `result-popup` | Transcription result + copy, auto-close with hover pause |
|
||
| `history-popup` | Recent transcriptions; ↑↓/Enter/1-9/ESC. Opened by the `history-popup` action (default `Ctrl+Shift+V`, rebindable) |
|
||
| `command-popup` | Command selection. Opened by the `command-popup` action (default `Ctrl+Shift+C`, rebindable) |
|
||
| `caption-overlay` | Live caption overlay (font/opacity/maxLines) |
|
||
| `suggestion-overlay` | Next-sentence ghost text: caret-anchored (`anchorFloatingPanel`), non-focusable, click-through unless `suggestionOverlayInteractive`; accept/next/dismiss come from global key bindings (the window never owns focus). Its narrow layout shows a one-line local provenance/count summary only, never raw memory evidence. Presentation includes candidate, generating, warm-up and partial-text states; its X hides the panel before main-process dismissal. |
|
||
|
||
---
|
||
|
||
## 5. Renderer IA
|
||
|
||
Routing is state-based in `AppLayout.tsx` (`Route` union + `NAV_ITEMS`), no react-router.
|
||
|
||
| Page | Route | Feature |
|
||
|---|---|---|
|
||
| `DashboardPage` | dashboard | Voice cockpit: hero, bento tiles, multi-engine hub (STT/LLM), telemetry, recent history, file drop |
|
||
| `HistoryPage` | history | History & memory timeline; search, tag filter, pagination, export/delete |
|
||
| `DictionaryPage` | dictionary | Custom vocabulary editor |
|
||
| `CommandsPage` | commands | Custom instructions + voice keyword rules + LLM chains + dictation templates |
|
||
| `VoiceConversationPage` | conversation | Duplex voice assistant (local pipeline vs OpenAI Realtime) |
|
||
| `KnowledgeBasePage` | knowledge | Local RAG: add/index docs, semantic query, reindex/remove |
|
||
| `MeetingModePage` | meeting | Meeting studio: live transcript, memos, doc generation/export, diarization |
|
||
|
||
Modals/components: `SettingsModal` (tabs General/Audio/STT/LLM/Input/License/Cloud/About), `LicenseModal`, `LicenseTab`, `CloudSyncSection`, `InputInsightsPanel` (consent + suggestion policy + weekly insights + learned phrases), `InputConsentPanel` (receipt + current-app exclusion recommendation), `InputInsightsView` (flow, friction and app-quality summaries), `OnboardingModal`, `UpgradePromptModal`, `ProBadge`, `TemplateSection`, `FileDropZone`, `OllamaGuideModal`, `CodexOAuthGuideModal`, `TitleBar`, `StatusBar`, meeting components (9), voice-conversation, payment (`CheckoutModal`, `checkout-flow.ts`), support (`SupportModal`), ads (`AdBanner`, `RewardedQuotaModal`), shared cards.
|
||
|
||
Key-binding UI lives in `components/keybinding/` (`Keycap`, `KeyBindingPicker`, `KeyBindingField`, `translation-key`), embedded in the Settings **General** tab (`SettingsModal.tsx:239`) — one field per action plus a global on/off switch. The picker offers both key recording and a searchable grouped dropdown (MUI `Autocomplete` over `KEY_CATALOG`, `KeyBindingPicker.tsx:536`). It replaced `HotkeyRecordModal`. `renderer/utils/format-hotkey.ts` is now a 17-line platform adapter only; key names, modifier glyphs, and join rules come from `@d3ro/core/keybinding`.
|
||
|
||
Hooks: `useRealtimeConversation` (OpenAI Realtime WebRTC), `useLicenseState`, `useProFeature`, `useKeyBindingMap` (subscribes to `keybinding:changed`; the dashboard renders the live `dictation` binding through `BindingKeycaps`), `useInputInsights` (telemetry + suggestion state + weekly summary + phrases; the Dashboard shows a weekly input card when collection is on).
|
||
|
||
DB schema (`src/main/db/schema.ts`, drizzle SQLite): `history`, `dictionary`, `stats`, `memo_tags`, `daily_usage`, `rag_documents`, `rag_chunks`, `meeting_sessions`, `meeting_memos`, `meeting_documents`, `input_activity` (hour × app counters), `typing_samples`, `personal_phrases`, `suggestions`.
|
||
|
||
---
|
||
|
||
## 6. Desktop status summary
|
||
|
||
- Core dictation/LLM/history pipeline: **implemented + tested**. The vitest case count in `apps/desktop` is **1465** after input intelligence and the Local Flow Intelligence extension; playwright e2e is separate. Read the pass numbers together with the `better-sqlite3` ABI the tree is built for (`11` GAP-INFRA-06) — they are not comparable across configurations:
|
||
- **Host Node ABI** (2026-09-21, before the LLM fix): 1311 / 1314 passing. The three failures are environment-dependent rather than regressions — two need a local sidecar venv or embedding server, one pins an error message that has since changed (`11` GAP-QA-02). **This configuration has not been re-measured since the LLM fix.**
|
||
- **Electron ABI** (2026-09-21, after input intelligence): `366 failed | 1042 passed (1408)`, against a clean-tree baseline of `366 failed | 994 passed (1360)` in the same configuration — identical failure count, +48 passed, **zero new failures**. 365 of those 366 are `tests/red/*.usecase.test.ts` files dying at DB creation because of the ABI mismatch, not assertions.
|
||
- **Native-ABI mismatch run** (2026-09-23, `1.5.0` release verification): `366 failed | 1099 passed (1465)`. The `better-sqlite3` module is still built for Electron ABI 130 while the host Node is ABI 131, so this is the same mismatch configuration. The failure count is unchanged (365 ABI + 1 stale sidecar error-message assertion, `11` GAP-QA-02) and the passed count rose with the new tests, so there are **zero new failures**. Full strict/full suite and Electron GUI verification remain open gates, not evidence of green.
|
||
- Input intelligence (`2026-09-21`): telemetry capture, weekly insights, next-sentence ghost text and phrase learning are implemented; real typing has verified capture → UIA snapshot → policy decisions, while overlay position/appearance, accept-insert, password blocking and weekly numbers remain manual gates (`11` GAP-INPUT-01). The 24.7 s cold / 4.9 s warm figures are a **historical latency diagnosis**, not current behaviour. **2026-09-22 operating verification:** at 19:14:12 app startup warmup plus `keep_alive: 30m` held `gemma4:e4b` resident through 19:44:12 at VRAM 3,226,342,521 bytes and context 4096. Windows GPU Engine PID sampling found no active Ollama compute then, so this was forced residency rather than infinite inference. Separately, 19:11:17–19:11:58 logs show automatic suggestions repeatedly generated under the old 900 ms / 12 per min / 5-candidate / 128-token / one-character-growth policy: a permitted burst, not proof of a single stuck request. The replacement guard removes boot warmup, uses `keep_alive: 2m`, and sets 600 ms debounce, 5 s minimum interval, 6 per min (hard maximum 12), 3 candidates, 64 tokens, 12-character growth and 8 s timeout. `LocalLLMService` adds per-request cancellation, bounded generate/stream/chat calls, required `done` frames and incomplete cleanup; voice cancellation is single-flight and reaches the request; server spawn/polling are deduplicated and disposed. Independent targeted verification passed 6 test files / 69 tests with 0 failures; changed code/tests ESLint and `git diff --check` exited 0. Raw Ollama proof: a cold bounded request hit the client hard timeout at 15.044 s then left `/api/ps` empty and `/api/version` recovered in 80 ms; explicit warmup returned HTTP 200 in 16.639 s; the subsequent `num_predict=1`, `keep_alive='2m'` request returned HTTP 200 in 553 ms with `done:true`, `eval_count:1`, `response=OK`, `done_reason:length`, and an observed `/api/ps` expiry of about 119.9 s. At 19:48:59 +09:00, with no intervening generate/unload/kill/retry, a single `/api/ps` returned HTTP 200 in 45.8 ms with `{models:[]}` and `/api/version` returned HTTP 200 in 7.3 ms with `0.32.13`: raw API expiry/unload evidence only. This does not prove an app restart, GUI overlay or real automatic-typing path, so those remain `[~]` runtime gates (`11` GAP-LLM-04, GAP-INPUT-06).
|
||
- **2026-09-23 overlay lifecycle / Windows child-process audit:** focused five test files passed 80 tests; core had 131 passing tests in an earlier independent verification; desktop typecheck/lint, Python `py_compile`, and `git diff --check` exited 0. Code and automated-test audit covers `windowsHide:true` on TTS PowerShell, VoiceAction cmd/PowerShell/general exec, audio-device and active-window PowerShell, plus all three ffmpeg paths; with existing SoX/STT/Ollama/sound-effect coverage, no Windows-capable desktop-main child-process call is known to be omitted. This is not Electron GUI runtime proof. Keep INPUT-07 and GAP-INPUT runtime gates `[~]` until an external-terminal `run-desktop.bat` restart confirms dismissal under generation/selection/focus loss and no cmd/PowerShell window recurrence for TTS, voice action, audio enumeration, screen context and file transcription.
|
||
- Cross-platform packaging: Windows NSIS (signed, `forceCodeSigning`), macOS DMG/ZIP arm64 (ad-hoc signing); auto-update via canonical Forgejo feed with update policy (`release/update-policy.json`).
|
||
- Local-first AI (SoX + faster-whisper sidecar + bundled Ollama) and cloud paths both present.
|
||
- **Local STT is packaged** (`1.3.0`): `electron-builder.yml` `extraResources` copies `sidecar-dist/sidecar` → `resources/sidecar` and `resources/ffmpeg` → `resources/ffmpeg`; `scripts/ci/verify-sidecar-bundle.mjs` gates packaging. Build locally with `npm --prefix apps/desktop run sidecar:setup && npm --prefix apps/desktop run sidecar:build`. The sidecar stays in console mode so `stdout`/`stderr` reach the app log (UTF-8, line-buffered); a packaged sidecar **must** exist or startup fails loudly instead of silently falling back to a system Python.
|
||
- All local engine URLs (`LocalSTTService`, `LocalLLMService`, `RAGService`, `OnlineLLMService`, `STTManager`) pass through `src/main/utils/loopback.ts`, which rewrites `localhost` to `127.0.0.1`, because some Windows hosts resolve `localhost` to IPv6 only and local engines bind IPv4.
|
||
- Meeting intelligence, RAG, voice conversation (local + Realtime), captions, file transcription: implemented.
|
||
- **LLM instruction prompts: fixed 2026-09-21 (`9c2b4d4`), not yet verified in a running app.** Running a custom instruction inserted the instruction's own wording instead of the processed result. Two faults stacked: the instruction was passed in the `text` argument with the system-prompt argument left empty, and `BASE_SYSTEM_PROMPTS` has no `custom` key so resolution fell back to `refine` **silently** — the model polished the instruction and the transcript never reached it; separately, only `{{text}}` was substituted and none of the five built-in presets use it (`{{targetLanguage}}`, `{{userPrompt}}`, or no placeholder), so the substitution was a no-op from the day it was written. Introduced in `fea923d` (2026-04-05) and present in every release `v0.1.0-alpha`..`v1.4.0` — **the path never worked; this is not a regression.** Plain actions (`refine`/`summarize`/`grammar`/`expand`) were unaffected and are now pinned by regression cases. The fix routes all three entry paths through `llm-prompts.ts` (see §2) and additionally corrects two things found alongside it: a voice shortcut naming an instruction was nullified by the `defaultLLMAction === 'none'` gate (`VoiceModeService.ts:779`), and the commands-page pipeline bench called `llm.generate`, which preload does not expose, so every run threw and the `catch` displayed the **input** as if it had succeeded — a fail-closed violation that is the reason the bug went unnoticed for five months (`CommandsPage.tsx:180-205`, now on `llm.process` with failures rendered as failures).
|
||
- **Verification limits — do not read this as verified.** Unit tests pass (`llm-prompts.test.ts` 21, `llm-handlers.test.ts` 7, `VoiceModeService.test.ts` 22, `ChainService.test.ts` 5), and each of the four fixes was reverted individually to confirm the tests actually fail without it. `npm run lint` (apps/desktop scope) passes; `tsconfig.check.json` errors went 36 → 35 (the `llm.generate` error is gone) with no errors in the touched files. But there is **no running-app run**, and `tests/red/{instruction,chain,voice,config}.usecase.test.ts` — precisely the related paths — never executed because of the `better-sqlite3` ABI mismatch. That range is neither passing nor failing; it is untested (`11` GAP-LLM-02, GAP-INFRA-06).
|
||
- **Key bindings: implemented and verified on Windows.** Every global shortcut now comes from one contract (`@d3ro/core/keybinding`) with multiple bindings per action, mouse-button support, and no hardcoded accelerators left in `bootstrap.ts`. A manual run on 2026-09-21 confirmed legacy migration (custom values preserved), 6 actions loaded, the uiohook keyboard **and** mouse hook active with zero boot errors, and multi-binding working; contract side is `packages/core` 117 tests GREEN with no type errors in the key-binding files (`11` GAP-KEY-01 `[x]`). Two things remain open: `KeyBindingService` has no unit test of its own, and macOS/Linux mouse behavior is unconfirmed (`11` GAP-KEY-02). The rewrite also fixed a dead hands-free double-press path, an order-dependent reserved-combo check, a `globalShortcut.unregisterAll()` that wiped the popup accelerators, and a `setEnabled(true)` that re-enabled hooking with an empty binding set.
|
||
- The same pass fixed an unrelated pre-existing dashboard bug: `caption.onStateChanged` delivers `{ state }`, but `DashboardPage` passed the whole object into `setCaptionState`, so the caption status readout never showed the right value (`DashboardPage.tsx:148`).
|
||
- **Ad mediation**: `DirectHouseSponsorAdapter` performs real configurable REST bids; the other 9 adapters remain fail-closed stubs pending official SDKs (see `11-gap-backlog.md` GAP-ADS-01/02).
|
||
- Tier resolution now routes through `@d3ro/core/entitlement` (`resolveEntitlement`, `normalizeEntitlementTier`); `useLicenseState.isPro` includes `pro_plus`.
|
||
- No `TODO`/`FIXME` markers found in `src` (grep clean). `src/main/types/` is an empty directory.
|
||
|
||
---
|
||
|
||
## 7. Key file anchors
|
||
|
||
| Thing | Path |
|
||
|---|---|
|
||
| App entry / deep links | `src/main/index.ts` |
|
||
| Bootstrap order | `src/main/bootstrap.ts` |
|
||
| IPC registry | `src/main/ipc/index.ts` |
|
||
| IPC channel SSOT | `packages/core/src/ipc-channels.ts` |
|
||
| Key-binding contract SSOT | `packages/core/src/keybinding.ts` (catalog, actions, validation, conflicts, formatting, parsing) |
|
||
| Input intelligence domain SSOT | `packages/core/src/input-intelligence.ts` (key classification, typed delta, suggestion policy, overlay anchoring, flow/friction aggregation, phrase decay/ranking, provenance and exclusion recommendation) |
|
||
| Input intelligence services | `src/main/services/InputTelemetryService.ts`, `SuggestionService.ts`, `UiaContextService.ts`, `global-input-hook.ts` |
|
||
| Foreground window FFI | `src/main/utils/win32-foreground.ts` (koffi → user32/kernel32) |
|
||
| UIA bridge (sidecar) | `sidecar/uia_bridge.py` + `GET /uia/focus` in `sidecar/main.py` |
|
||
| Input UI | `src/renderer/components/input-insights/{InputConsentPanel,InputInsightsView}.tsx`, `src/renderer/hooks/useInputInsights.ts`, `src/renderer/popups/suggestion-overlay/` |
|
||
| Local Flow Intelligence tests | `tests/main/services/input-flow-domain.test.ts` (9), `input-flow-services.test.ts` (12); targeted suite 21/21 at the documented handoff point |
|
||
| Key-binding service / IPC / UI | `src/main/services/KeyBindingService.ts`, `src/main/ipc/keybinding-handlers.ts`, `src/renderer/components/keybinding/` |
|
||
| Preload API | `src/preload/index.ts` |
|
||
| Windows | `src/main/windows/WindowManager.ts` |
|
||
| Voice orchestrator | `src/main/services/VoiceModeService.ts` |
|
||
| LLM prompt / placeholder SSOT | `src/main/services/llm-prompts.ts` (shared by `VoiceModeService`, `ChainService`, `ipc/llm-handlers.ts`) |
|
||
| DB schema | `src/main/db/schema.ts` |
|
||
| Renderer shell / routes | `src/renderer/components/AppLayout.tsx` |
|
||
| Update feed SSOT | `src/main/update-feed.ts` |
|
||
| Update policy SSOT | `release/update-policy.json` + `src/main/update-policy.ts` |
|
||
| Path/loopback resolution | `src/main/utils/paths.ts`, `src/main/utils/loopback.ts` |
|
||
| Sidecar source / packaging | `sidecar/main.py`, `scripts/setup-sidecar.mjs`, `scripts/build-sidecar.mjs`, `scripts/ci/verify-sidecar-bundle.mjs` |
|