feat(desktop): make local speech transcription work end to end

Local dictation had never produced a transcript on an installed build. The
engine itself was healthy; every connection to it was broken.

Installed builds shipped no speech engine at all: the packaging config had no
entry for the faster-whisper sidecar and no pipeline step built one, so the app
always fell back to a system Python without the runtime. Development was broken
too, because the sidecar and SoX paths were resolved against the Vite output
directory instead of the app root, which also meant recording failed with a SoX
ENOENT. On hosts where localhost resolves only to IPv6, every local request was
refused outright, which silently disabled both local transcription and the local
LLM.

The sidecar is now built and bundled (including the Silero VAD data it needs),
gated by a packaging check that fails when the engine or its data is missing.
Paths are discovered from the app root and fail loudly when the engine is
absent. Local engine URLs are normalized to the IPv4 loopback, decoding is tuned
so repeated hallucinations cannot compound (the same transcript now takes about
a fifth of the time), the engine is warmed up at startup, and holding the hotkey
now shows the text forming live in the recording tip.
This commit is contained in:
Yun Chan 2026-09-18 00:48:47 +09:00
parent 359b244dc9
commit 2d585bfc29
52 changed files with 1450 additions and 3861 deletions

View file

@ -22,6 +22,7 @@ import { AssemblyAIDriver } from './drivers/AssemblyAIDriver'
import { GoogleDriver } from './drivers/GoogleDriver'
import { CustomDriver } from './drivers/CustomDriver'
import { D3ROCloudDriver } from './drivers/D3ROCloudDriver'
import { normalizeLoopbackUrl } from '../../utils/loopback'
const logger = getLogger('STTManager')
@ -43,7 +44,7 @@ export const STT_PROVIDERS_META: STTProviderInfo[] = [
badge: 'Cloud · Zero Config',
requiresApiKey: false,
defaultModel: 'default',
defaultBaseUrl: 'http://localhost:5000',
defaultBaseUrl: 'http://127.0.0.1:5000',
models: ['default', 'whisper-large-v3-turbo', 'nova-3', 'gemini-2.0-flash'],
isCloud: true,
},
@ -109,7 +110,7 @@ export const STT_PROVIDERS_META: STTProviderInfo[] = [
badge: 'Self-Hosted / Proxy',
requiresApiKey: false,
defaultModel: 'whisper-1',
defaultBaseUrl: 'http://localhost:8000/v1',
defaultBaseUrl: 'http://127.0.0.1:8000/v1',
models: ['whisper-1', 'custom'],
isCloud: true,
},
@ -155,7 +156,8 @@ export class STTManager extends EventEmitter {
return {
apiKey: specificConfig.apiKey ?? '',
baseUrl: specificConfig.baseUrl ?? meta?.defaultBaseUrl ?? '',
// 저장된 값이 localhost일 수 있다(IPv6 해석 실패) → IPv4 루프백으로 정규화.
baseUrl: normalizeLoopbackUrl(specificConfig.baseUrl ?? meta?.defaultBaseUrl ?? ''),
modelId: specificConfig.modelId ?? meta?.defaultModel ?? '',
temperature: specificConfig.temperature ?? 0,
}
@ -245,6 +247,25 @@ export class STTManager extends EventEmitter {
}
}
/**
* 녹음 중 실시간 미리보기 전사(최종 삽입과 무관).
* 로컬 Whisper에서만 지원한다 — 클라우드 공급자는 요청 비용/지연이 커서 사용하지 않는다.
* 실패는 빈 문자열로 흡수된다.
*/
async transcribePartial(audioBuffer: Buffer, options?: TranscribeOptions): Promise<string> {
if (this.getActiveProvider() !== 'local') return ''
return getLocalSTTService().transcribePartial(audioBuffer, options)
}
/**
* 로컬 STT 엔진(sidecar + 모델)을 백그라운드로 미리 데운다.
* 첫 받아쓰기 지연을 없애는 것이 목적이며 실패해도 조용히 넘어간다.
*/
async warmUpLocal(): Promise<boolean> {
if (this.getActiveProvider() !== 'local') return false
return getLocalSTTService().warmUp()
}
getStatus(): STTStatus {
const provider = this.getActiveProvider()
if (provider === 'local') {