아바타 v3 세션 연결 — P1 서연 리노컷 아바타와 음성 립싱크
Some checks are pending
API contract / OpenAPI type drift (push) Waiting to run
Some checks are pending
API contract / OpenAPI type drift (push) Waiting to run
- ClientAvatar가 리그 레지스트리(v3/rigs)로 P1을 판정해 그림 영역을 ClientAvatarV3로 그린다(로드 중·실패 시 기존 SVG, VITE_AVATAR_V3=0이면 끔) - 발화 구동 공용 모듈 speechDriver: Lab과 세션이 함께 쓰고, TTS AudioBuffer를 오디오 시계로 표본해 립싱크·발화 동반층·지문 cue를 돌린다 - Session: TTS 재생 시작 시 발화(원문·버퍼·시작 시각) 전달, 음성 실패 시 텍스트 타이밍 발화, 정지 경로에서 해제, 개방도·겉표정 강도 전달 - 정지 화면(시작 전·썸네일)도 한 프레임을 그리고, 박스 실측 폭으로 상반신/얼굴 크롭을 고른다 - 세션 발화 E2E(오디오·실패 경로), P1 레이아웃 게이트, P1 표정 시험을 v3 기대로 갱신
This commit is contained in:
parent
ecb36d123f
commit
613bcb603e
14 changed files with 1000 additions and 182 deletions
|
|
@ -96,6 +96,9 @@ src/
|
||||||
- UI 프리미티브는 `components/ui` 배럴에서 가져온다. props/타입이 안정 계약이다.
|
- UI 프리미티브는 `components/ui` 배럴에서 가져온다. props/타입이 안정 계약이다.
|
||||||
- `ClientAvatar` props(`persona`/`state`/`affect`/`analyser`)는 확정 인터페이스.
|
- `ClientAvatar` props(`persona`/`state`/`affect`/`analyser`)는 확정 인터페이스.
|
||||||
avatar 에이전트는 이 파일 내부 SVG/모션만 고도화하고 시그니처는 유지.
|
avatar 에이전트는 이 파일 내부 SVG/모션만 고도화하고 시그니처는 유지.
|
||||||
|
리노컷 리그가 있는 페르소나(`components/avatar/v3/rigs`, 지금은 P1)는 내부에서 v3로 그리며,
|
||||||
|
선택 prop `openness`·`surfaceIntensity`·`speech`(TTS 발화 구동)를 더 받는다(결정문
|
||||||
|
`docs/decisions/avatar-expression-engine-v3.md` §8.5). 빌드 플래그 `VITE_AVATAR_V3=0`이면 모두 기존 SVG.
|
||||||
- 세션 데이터는 `lib/api.ts`의 `sessionApi`(start/turn/end/stream) 사용.
|
- 세션 데이터는 `lib/api.ts`의 `sessionApi`(start/turn/end/stream) 사용.
|
||||||
SSE 토큰 수신은 `openSessionStream(sessionId, { onToken, onDone, ... })`.
|
SSE 토큰 수신은 `openSessionStream(sessionId, { onToken, onDone, ... })`.
|
||||||
- 페이지는 `default export`. `AppShell`로 감싸면 톱바/네비/역할 accent가 자동 적용.
|
- 페이지는 `default export`. `AppShell`로 감싸면 톱바/네비/역할 accent가 자동 적용.
|
||||||
|
|
|
||||||
|
|
@ -337,15 +337,16 @@ test.describe("persona avatar expression rig", () => {
|
||||||
await expect(page.locator(".sx-stage__now")).toContainText("온화함");
|
await expect(page.locator(".sx-stage__now")).toContainText("온화함");
|
||||||
});
|
});
|
||||||
|
|
||||||
test("uses the original SVG parameter rig for P1 Seoyeon", async ({ page }) => {
|
test("uses the v3 linocut avatar for P1 Seoyeon", async ({ page }) => {
|
||||||
await page.goto("/learn/session/P1");
|
await page.goto("/learn/session/P1");
|
||||||
|
|
||||||
const avatar = page.locator('.vg-avatar[data-persona-code="P1"]').first();
|
const avatar = page.locator('.vg-avatar[data-persona-code="P1"]').first();
|
||||||
await expect(avatar).toBeVisible();
|
await expect(avatar).toBeVisible();
|
||||||
await expect(avatar).toHaveAttribute("data-render-mode", "svg");
|
|
||||||
await expect(avatar).toHaveAttribute("data-affect", "sad");
|
await expect(avatar).toHaveAttribute("data-affect", "sad");
|
||||||
await expect(avatar.locator(".vg-raster")).toHaveCount(0);
|
|
||||||
await expect(avatar.locator(".vg-avatar__svg")).toBeVisible();
|
const v3 = avatar.locator('[data-avatar-renderer="linocut"]');
|
||||||
await expect(avatar.locator('[data-avatar-neck="true"]')).toBeVisible();
|
await expect(v3).toHaveAttribute("data-load-state", "ready", { timeout: 15_000 });
|
||||||
|
await expect(v3.locator(".linocut-avatar")).toBeVisible();
|
||||||
|
await expect(v3).toHaveAttribute("data-viseme", "X");
|
||||||
});
|
});
|
||||||
});
|
});
|
||||||
|
|
|
||||||
300
apps/web/e2e/avatar-session-speech.spec.ts
Normal file
300
apps/web/e2e/avatar-session-speech.spec.ts
Normal file
|
|
@ -0,0 +1,300 @@
|
||||||
|
/* =====================================================================
|
||||||
|
avatar-session-speech.spec.ts — P1 세션 TTS 재생 ↔ v3(리노컷) 아바타
|
||||||
|
발화 연결 E2E(결정문 §8.5 "세션 연결" 테스트 계약).
|
||||||
|
|
||||||
|
목적: TTS 응답을 테스트 안에서 만든 WAV(유성 구간이 있는 톤버스트
|
||||||
|
여러 개)로 모킹해, 재생 중 v3 래퍼의 data-speech-source="audio"·
|
||||||
|
data-viseme 순환을 확인하고, 재생이 끝나면 발화 속성이 정리되는지
|
||||||
|
본다. TTS 실패(500)면 data-speech-source="text"로 텍스트 타이밍
|
||||||
|
발화로 떨어지는지도 본다.
|
||||||
|
|
||||||
|
근거: apps/web/src/pages/Session.tsx playTtsAudio/speakTextClientTurn,
|
||||||
|
apps/web/src/components/avatar/v3/ClientAvatarV3.tsx.
|
||||||
|
주의: 모든 API는 route fixture로 모킹한다(실제 AI 엔진·TTS 제공자 없음).
|
||||||
|
헤드리스 자동재생 정책 때문에 이 파일만 autoplay-policy를 느슨하게
|
||||||
|
연다(세션 범위, 다른 스펙에 영향 없음).
|
||||||
|
===================================================================== */
|
||||||
|
|
||||||
|
import { expect, test, type Locator, type Page, type Route } from "@playwright/test";
|
||||||
|
|
||||||
|
test.use({ launchOptions: { args: ["--autoplay-policy=no-user-gesture-required"] } });
|
||||||
|
|
||||||
|
const SESSION_ID = "66666666-6666-4666-8666-666666666666";
|
||||||
|
const LEARNER_TEXT = "요즘 많이 힘들었어요. 어떤 마음이 가장 크게 남아 있나요?";
|
||||||
|
const CLIENT_REPLY = "그냥요. 잠을 잘 못 자요. 아무 것도 하고 싶지 않아요.";
|
||||||
|
|
||||||
|
function jsonRoute(body: unknown, status = 200) {
|
||||||
|
return { status, contentType: "application/json", body: JSON.stringify(body) };
|
||||||
|
}
|
||||||
|
|
||||||
|
function sseTurnBody(tokens: string[], done: Record<string, unknown>): string {
|
||||||
|
const lines: string[] = [];
|
||||||
|
for (const token of tokens) lines.push("event: token", `data: ${token}`, "");
|
||||||
|
lines.push("event: done", `data: ${JSON.stringify(done)}`, "");
|
||||||
|
return lines.join("\n");
|
||||||
|
}
|
||||||
|
|
||||||
|
async function fulfillTurn(route: Route, tokens: string[], done: Record<string, unknown>) {
|
||||||
|
await route.fulfill({ status: 200, contentType: "text/event-stream", body: sseTurnBody(tokens, done) });
|
||||||
|
}
|
||||||
|
|
||||||
|
function writeAsciiString(view: DataView, offset: number, text: string): void {
|
||||||
|
for (let i = 0; i < text.length; i++) view.setUint8(offset + i, text.charCodeAt(i));
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* 유성 구간이 있는 톤버스트 여러 개로 테스트용 WAV를 만든다(16비트 PCM 모노).
|
||||||
|
* decodeAudioData가 바로 디코드할 수 있다. 무음 구간(0)과 사인파 구간을 번갈아
|
||||||
|
* 넣어 speechEnvelope.ts의 유성 구간 검출(포락선 > max(0.02, 0.12·P95))이
|
||||||
|
* 서로 다른 구간 여러 개를 집어내게 한다.
|
||||||
|
*/
|
||||||
|
function buildToneBurstWav(): Buffer {
|
||||||
|
const sampleRate = 16000;
|
||||||
|
const amplitude = 0.7;
|
||||||
|
const segments: Array<{ freq: number; durMs: number }> = [
|
||||||
|
{ freq: 0, durMs: 150 },
|
||||||
|
{ freq: 220, durMs: 320 },
|
||||||
|
{ freq: 0, durMs: 180 },
|
||||||
|
{ freq: 420, durMs: 280 },
|
||||||
|
{ freq: 0, durMs: 180 },
|
||||||
|
{ freq: 320, durMs: 300 },
|
||||||
|
{ freq: 0, durMs: 220 },
|
||||||
|
];
|
||||||
|
|
||||||
|
const samples: number[] = [];
|
||||||
|
for (const seg of segments) {
|
||||||
|
const count = Math.round((seg.durMs / 1000) * sampleRate);
|
||||||
|
for (let i = 0; i < count; i++) {
|
||||||
|
samples.push(seg.freq > 0 ? amplitude * Math.sin((2 * Math.PI * seg.freq * i) / sampleRate) : 0);
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
const dataLength = samples.length * 2;
|
||||||
|
const buffer = Buffer.alloc(44 + dataLength);
|
||||||
|
const view = new DataView(buffer.buffer, buffer.byteOffset, buffer.byteLength);
|
||||||
|
|
||||||
|
writeAsciiString(view, 0, "RIFF");
|
||||||
|
view.setUint32(4, 36 + dataLength, true);
|
||||||
|
writeAsciiString(view, 8, "WAVE");
|
||||||
|
writeAsciiString(view, 12, "fmt ");
|
||||||
|
view.setUint32(16, 16, true);
|
||||||
|
view.setUint16(20, 1, true); // PCM
|
||||||
|
view.setUint16(22, 1, true); // mono
|
||||||
|
view.setUint32(24, sampleRate, true);
|
||||||
|
view.setUint32(28, sampleRate * 2, true); // byte rate(mono·16비트)
|
||||||
|
view.setUint16(32, 2, true); // block align
|
||||||
|
view.setUint16(34, 16, true); // bits per sample
|
||||||
|
writeAsciiString(view, 36, "data");
|
||||||
|
view.setUint32(40, dataLength, true);
|
||||||
|
|
||||||
|
for (let i = 0; i < samples.length; i++) {
|
||||||
|
const clamped = Math.max(-1, Math.min(1, samples[i]));
|
||||||
|
view.setInt16(44 + i * 2, Math.round(clamped * 32767), true);
|
||||||
|
}
|
||||||
|
return buffer;
|
||||||
|
}
|
||||||
|
|
||||||
|
async function routeSessionScreen(page: Page) {
|
||||||
|
const startedAt = new Date(Date.now() - 60_000);
|
||||||
|
|
||||||
|
// catch-all을 먼저 등록한다 — Playwright는 나중에 등록한 route가 이긴다.
|
||||||
|
await page.route("**/api/**", (route) => route.fulfill(jsonRoute({ detail: "not part of this fixture" }, 404)));
|
||||||
|
|
||||||
|
await page.route("**/api/auth/me", (route) =>
|
||||||
|
route.fulfill(
|
||||||
|
jsonRoute({
|
||||||
|
user_id: "00000000-0000-0000-0000-0avatarspeech",
|
||||||
|
email: "avatar-speech.learner@hs.ac.kr",
|
||||||
|
display_name: "학습자",
|
||||||
|
role: "learner",
|
||||||
|
admin_access: false,
|
||||||
|
super_admin: false,
|
||||||
|
account_status: "approved",
|
||||||
|
approval_required: false,
|
||||||
|
cohort_ids: [],
|
||||||
|
consent_at: Math.floor(Date.now() / 1000),
|
||||||
|
onboarding_completed_at: Math.floor(Date.now() / 1000),
|
||||||
|
nickname: "학습자",
|
||||||
|
self_introduction: "아바타 v3 발화 연결 E2E 검증용 학습자입니다.",
|
||||||
|
avatar_url: "",
|
||||||
|
}),
|
||||||
|
),
|
||||||
|
);
|
||||||
|
|
||||||
|
await page.route("**/api/users/me/prepost-measures**", (route) =>
|
||||||
|
route.fulfill(
|
||||||
|
jsonRoute({
|
||||||
|
pilot_id: "phase3-pilot-draft",
|
||||||
|
instrument_version: "pilot-prepost-scaffold-2026-06-28",
|
||||||
|
measures: [],
|
||||||
|
complete_pre_count: 0,
|
||||||
|
complete_post_count: 0,
|
||||||
|
updated_at: null,
|
||||||
|
}),
|
||||||
|
),
|
||||||
|
);
|
||||||
|
|
||||||
|
await page.route("**/api/personas", (route) =>
|
||||||
|
route.fulfill(
|
||||||
|
jsonRoute([
|
||||||
|
{
|
||||||
|
code: "P1",
|
||||||
|
display_name: "서연(가명) · 고2 · 우울/자살사고",
|
||||||
|
difficulty: "hard",
|
||||||
|
theory_target: ["humanistic", "cbt"],
|
||||||
|
demographics: { age_band: "16-18", sex: "female", grade: "고2", status: "재학" },
|
||||||
|
presenting_summary: "우울감과 자살사고 위험",
|
||||||
|
voice_preset: null,
|
||||||
|
source: "database",
|
||||||
|
degraded: false,
|
||||||
|
},
|
||||||
|
]),
|
||||||
|
),
|
||||||
|
);
|
||||||
|
|
||||||
|
await page.route("**/api/voice/health", (route) =>
|
||||||
|
route.fulfill(jsonRoute({ available: false, reason: "e2e fixture" })),
|
||||||
|
);
|
||||||
|
|
||||||
|
await page.route("**/api/sessions", async (route) => {
|
||||||
|
if (route.request().method() !== "POST") {
|
||||||
|
await route.fallback();
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
await route.fulfill(
|
||||||
|
jsonRoute(
|
||||||
|
{
|
||||||
|
session_id: SESSION_ID,
|
||||||
|
case_id: "avatar-speech-case-001",
|
||||||
|
session_no: 1,
|
||||||
|
stage: "라포",
|
||||||
|
effective_openness: 0.3,
|
||||||
|
recall_summary: null,
|
||||||
|
degraded: false,
|
||||||
|
},
|
||||||
|
201,
|
||||||
|
),
|
||||||
|
);
|
||||||
|
});
|
||||||
|
|
||||||
|
await page.route(`**/api/sessions/${SESSION_ID}`, (route) =>
|
||||||
|
route.fulfill(
|
||||||
|
jsonRoute({
|
||||||
|
session_id: SESSION_ID,
|
||||||
|
case_id: "avatar-speech-case-001",
|
||||||
|
persona_code: "P1",
|
||||||
|
persona_name: "서연",
|
||||||
|
session_no: 1,
|
||||||
|
status: "active",
|
||||||
|
stage: "라포",
|
||||||
|
theory_mode: "humanistic",
|
||||||
|
effective_openness: 0.3,
|
||||||
|
started_at: startedAt.toISOString(),
|
||||||
|
ended_at: null,
|
||||||
|
review_ready: false,
|
||||||
|
turns: [],
|
||||||
|
}),
|
||||||
|
),
|
||||||
|
);
|
||||||
|
|
||||||
|
// 회기 전 자기점검(pre)은 이미 원장에 잠긴 상태로 제공한다 — 이 스펙의 대상이 아니다.
|
||||||
|
await page.route(`**/api/sessions/${SESSION_ID}/alliance-pulses`, async (route) => {
|
||||||
|
if (route.request().method() === "POST") {
|
||||||
|
await route.fulfill(jsonRoute({ pulse_id: "aaaaaaaa-aaaa-4aaa-8aaa-aaaaaaaaaaaa", status: "awaiting_agents" }, 202));
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
await route.fulfill(
|
||||||
|
jsonRoute({
|
||||||
|
items: [
|
||||||
|
{
|
||||||
|
pulse_id: "aaaaaaaa-aaaa-4aaa-8aaa-aaaaaaaaaaaa",
|
||||||
|
checkpoint: "pre",
|
||||||
|
status: "ready",
|
||||||
|
learner_locked_at: startedAt.toISOString(),
|
||||||
|
revealed_at: startedAt.toISOString(),
|
||||||
|
error_code: null,
|
||||||
|
self_scores: { goal: 0.5, task: 0.5, bond: 0.5 },
|
||||||
|
measurements: [],
|
||||||
|
},
|
||||||
|
],
|
||||||
|
}),
|
||||||
|
);
|
||||||
|
});
|
||||||
|
|
||||||
|
await page.route(`**/api/sessions/${SESSION_ID}/live-coach`, (route) =>
|
||||||
|
route.fulfill(jsonRoute({ source: "database", quota: { remaining: 3, max: 3 }, credit_events: [], events: [] })),
|
||||||
|
);
|
||||||
|
}
|
||||||
|
|
||||||
|
function doneEvent(overrides: Record<string, unknown> = {}) {
|
||||||
|
return {
|
||||||
|
session_id: SESSION_ID,
|
||||||
|
stage: "라포",
|
||||||
|
effective_openness: 0.42,
|
||||||
|
turn_seq: 1,
|
||||||
|
safety_flagged: false,
|
||||||
|
...overrides,
|
||||||
|
};
|
||||||
|
}
|
||||||
|
|
||||||
|
async function sendLearnerTurn(page: Page) {
|
||||||
|
const input = page.getByLabel("학습자 발화 입력");
|
||||||
|
await input.fill(LEARNER_TEXT);
|
||||||
|
await page.getByRole("button", { name: "보내기" }).click();
|
||||||
|
}
|
||||||
|
|
||||||
|
/** 재생 중 data-viseme가 거친 서로 다른 값의 집합(관찰 상한 5초, 50ms 간격 표본). */
|
||||||
|
async function collectDistinctVisemes(locator: Locator, minDistinct: number): Promise<Set<string>> {
|
||||||
|
const seen = new Set<string>();
|
||||||
|
const deadline = Date.now() + 5_000;
|
||||||
|
while (Date.now() < deadline) {
|
||||||
|
const viseme = await locator.getAttribute("data-viseme");
|
||||||
|
if (viseme) seen.add(viseme);
|
||||||
|
if (seen.size >= minDistinct) break;
|
||||||
|
await new Promise((resolve) => setTimeout(resolve, 50));
|
||||||
|
}
|
||||||
|
return seen;
|
||||||
|
}
|
||||||
|
|
||||||
|
test.describe("P1 세션 — v3 아바타 TTS 발화 연결", () => {
|
||||||
|
test("TTS 오디오 재생 중 립싱크가 돌고, 끝나면 발화 속성이 정리된다", async ({ page }) => {
|
||||||
|
await routeSessionScreen(page);
|
||||||
|
await page.route("**/api/voice/speech", (route) =>
|
||||||
|
route.fulfill({ status: 200, contentType: "audio/wav", body: buildToneBurstWav() }),
|
||||||
|
);
|
||||||
|
await page.route(`**/api/sessions/${SESSION_ID}/stream`, (route) => fulfillTurn(route, [CLIENT_REPLY], doneEvent()));
|
||||||
|
|
||||||
|
await page.goto("/learn/session/P1");
|
||||||
|
await page.getByRole("button", { name: "회기 시작" }).click();
|
||||||
|
await expect(page.locator(".sx-page--active")).toBeVisible({ timeout: 15_000 });
|
||||||
|
|
||||||
|
await sendLearnerTurn(page);
|
||||||
|
|
||||||
|
const v3 = page.locator(".sx-page--active .vg-avatar [data-avatar-renderer='linocut']");
|
||||||
|
await expect(v3).toHaveAttribute("data-speech-source", "audio", { timeout: 15_000 });
|
||||||
|
|
||||||
|
const visemes = await collectDistinctVisemes(v3, 2);
|
||||||
|
expect(visemes.size, `관찰한 비짐: ${[...visemes].join(",")}`).toBeGreaterThanOrEqual(2);
|
||||||
|
|
||||||
|
await expect(v3).not.toHaveAttribute("data-speech-source", "audio", { timeout: 10_000 });
|
||||||
|
await expect(v3).toHaveAttribute("data-viseme", "X", { timeout: 10_000 });
|
||||||
|
});
|
||||||
|
|
||||||
|
test("TTS 실패(500)면 텍스트 타이밍 발화로 떨어진다", async ({ page }) => {
|
||||||
|
await routeSessionScreen(page);
|
||||||
|
await page.route("**/api/voice/speech", (route) => route.fulfill(jsonRoute({ detail: "tts failure fixture" }, 500)));
|
||||||
|
await page.route(`**/api/sessions/${SESSION_ID}/stream`, (route) => fulfillTurn(route, [CLIENT_REPLY], doneEvent()));
|
||||||
|
|
||||||
|
await page.goto("/learn/session/P1");
|
||||||
|
await page.getByRole("button", { name: "회기 시작" }).click();
|
||||||
|
await expect(page.locator(".sx-page--active")).toBeVisible({ timeout: 15_000 });
|
||||||
|
|
||||||
|
await sendLearnerTurn(page);
|
||||||
|
|
||||||
|
const v3 = page.locator(".sx-page--active .vg-avatar [data-avatar-renderer='linocut']");
|
||||||
|
await expect(v3).toHaveAttribute("data-speech-source", "text", { timeout: 15_000 });
|
||||||
|
|
||||||
|
const visemes = await collectDistinctVisemes(v3, 2);
|
||||||
|
expect(visemes.size, `관찰한 비짐: ${[...visemes].join(",")}`).toBeGreaterThanOrEqual(2);
|
||||||
|
});
|
||||||
|
});
|
||||||
|
|
@ -666,6 +666,62 @@ test.describe("learner session full-screen layout", () => {
|
||||||
}
|
}
|
||||||
});
|
});
|
||||||
|
|
||||||
|
// P1 서연은 v3(리노컷) 아바타로 그린다(결정문 §8.5). 위 "dense viewport" 스윕을 P1로도
|
||||||
|
// 돌려 v3 래퍼(절대 위치·원형 클립)가 기존 SVG와 같은 레이아웃 게이트를 통과하는지 본다.
|
||||||
|
test("keeps critical session controls visible across dense viewport sizes for the v3 avatar (P1)", async ({
|
||||||
|
page,
|
||||||
|
}) => {
|
||||||
|
await signInAsLearner(page);
|
||||||
|
|
||||||
|
const viewports = [
|
||||||
|
{ width: 1366, height: 768 },
|
||||||
|
{ width: 1180, height: 768 },
|
||||||
|
{ width: 1024, height: 768 },
|
||||||
|
{ width: 881, height: 768 },
|
||||||
|
{ width: 820, height: 1180 },
|
||||||
|
{ width: 390, height: 844 },
|
||||||
|
{ width: 320, height: 568 },
|
||||||
|
];
|
||||||
|
|
||||||
|
await page.setViewportSize(viewports[0]);
|
||||||
|
await page.goto("/learn/session/P1");
|
||||||
|
await page.getByRole("button", { name: "회기 시작" }).click();
|
||||||
|
await completeAlliancePreCheckpoint(page);
|
||||||
|
await expect(page.locator(".sx-page.sx-page--active")).toBeVisible({ timeout: 15_000 });
|
||||||
|
|
||||||
|
for (const viewport of viewports) {
|
||||||
|
await page.setViewportSize(viewport);
|
||||||
|
await page.evaluate(() => new Promise(requestAnimationFrame));
|
||||||
|
|
||||||
|
await expect(page.locator(".sx-page.sx-page--active")).toBeVisible({ timeout: 15_000 });
|
||||||
|
await expectNoDocumentOverflow(page);
|
||||||
|
await expectNoHorizontalOverflow(page);
|
||||||
|
await expectSessionControlsInsideViewport(page);
|
||||||
|
await expectNoVisibleSessionPanelOverlap(page);
|
||||||
|
await expectMainControlsUnclipped(page);
|
||||||
|
await expectSessionPageHeightToMatchViewport(page);
|
||||||
|
await expectActiveSessionUsableLayout(page);
|
||||||
|
|
||||||
|
/* 320×568 저높이 폰은 session.css가 아바타 오브(.sx-orb-wrap)를 숨긴다(기존 SVG도 같다).
|
||||||
|
보일 때는 v3 래퍼가 스테이지 원 자리를 정확히 차지해야 한다. */
|
||||||
|
const orbShown = !(viewport.width === 320 && viewport.height === 568);
|
||||||
|
const stage = page.locator(".sx-page--active .vg-avatar__stage").first();
|
||||||
|
if (orbShown) {
|
||||||
|
await expect(stage).toBeVisible();
|
||||||
|
const v3 = stage.locator('[data-avatar-renderer="linocut"]');
|
||||||
|
await expect(v3).toHaveAttribute("data-load-state", "ready", { timeout: 15_000 });
|
||||||
|
const [stageBox, v3Box] = await Promise.all([stage.boundingBox(), v3.boundingBox()]);
|
||||||
|
expect(stageBox, "아바타 스테이지 좌표").not.toBeNull();
|
||||||
|
expect(v3Box, "v3 래퍼 좌표").not.toBeNull();
|
||||||
|
for (const key of ["x", "y", "width", "height"] as const) {
|
||||||
|
expect(Math.abs(stageBox![key] - v3Box![key]), `v3 래퍼 ${key}`).toBeLessThanOrEqual(1);
|
||||||
|
}
|
||||||
|
} else {
|
||||||
|
await expect(page.locator(".sx-page--active .sx-orb-wrap")).toBeHidden();
|
||||||
|
}
|
||||||
|
}
|
||||||
|
});
|
||||||
|
|
||||||
test("does not leave an unsaved local transcript when a text turn is rejected", async ({ page }) => {
|
test("does not leave an unsaved local transcript when a text turn is rejected", async ({ page }) => {
|
||||||
await signInAsLearner(page);
|
await signInAsLearner(page);
|
||||||
const persona = await fetchAvailablePersona(page, 1);
|
const persona = await fetchAvailablePersona(page, 1);
|
||||||
|
|
|
||||||
|
|
@ -35,6 +35,8 @@ import { Brows } from "./Brows";
|
||||||
import { Mouth } from "./Mouth";
|
import { Mouth } from "./Mouth";
|
||||||
import { live2dModel3Path, live2dModelForPersonaCode, live2dMotionForExpression } from "./live2dModel";
|
import { live2dModel3Path, live2dModelForPersonaCode, live2dMotionForExpression } from "./live2dModel";
|
||||||
import { useExpressionTransition } from "./useExpressionTransition";
|
import { useExpressionTransition } from "./useExpressionTransition";
|
||||||
|
import ClientAvatarV3, { type AvatarSpeech } from "./v3/ClientAvatarV3";
|
||||||
|
import { linocutRigFor } from "./v3/rigs";
|
||||||
import "./client-avatar.css";
|
import "./client-avatar.css";
|
||||||
|
|
||||||
/* ── 공개 타입 재노출 (기존 import 경로 호환) ──────────────────────────
|
/* ── 공개 타입 재노출 (기존 import 경로 호환) ──────────────────────────
|
||||||
|
|
@ -48,6 +50,7 @@ export {
|
||||||
type AvatarPersona,
|
type AvatarPersona,
|
||||||
} from "./persona";
|
} from "./persona";
|
||||||
export type { Live2DExpressionMotion, Live2DPersonaModel } from "./live2dModel";
|
export type { Live2DExpressionMotion, Live2DPersonaModel } from "./live2dModel";
|
||||||
|
export type { AvatarSpeech } from "./v3/ClientAvatarV3";
|
||||||
|
|
||||||
export interface ClientAvatarProps {
|
export interface ClientAvatarProps {
|
||||||
persona: AvatarPersona;
|
persona: AvatarPersona;
|
||||||
|
|
@ -65,6 +68,15 @@ export interface ClientAvatarProps {
|
||||||
* null/미지정이면 speaking 동안 차분한 의사 발화 모션.
|
* null/미지정이면 speaking 동안 차분한 의사 발화 모션.
|
||||||
*/
|
*/
|
||||||
speakingProgress?: number | null;
|
speakingProgress?: number | null;
|
||||||
|
/**
|
||||||
|
* 개방도 0~1(세션 effective_openness). v3(리노컷) 리그가 있는 페르소나에서만 쓴다.
|
||||||
|
* 미지정이면 rapport로 대신한다(결정문 §8.5).
|
||||||
|
*/
|
||||||
|
openness?: number;
|
||||||
|
/** 겉표정 강도 0~1(performance.ts surfaceIntensityFor). 미지정이면 0.5. v3 전용. */
|
||||||
|
surfaceIntensity?: number;
|
||||||
|
/** TTS 발화 구동(결정문 §8.5). v3 리그가 있는 페르소나에서만 립싱크·발화 동반층을 돈다. */
|
||||||
|
speech?: AvatarSpeech | null;
|
||||||
/** px 지름 (기본 220 — §5.2 아바타 220px) */
|
/** px 지름 (기본 220 — §5.2 아바타 220px) */
|
||||||
size?: number;
|
size?: number;
|
||||||
className?: string;
|
className?: string;
|
||||||
|
|
@ -215,6 +227,9 @@ export function ClientAvatar({
|
||||||
analyser = null,
|
analyser = null,
|
||||||
rapport = 0,
|
rapport = 0,
|
||||||
speakingProgress = null,
|
speakingProgress = null,
|
||||||
|
openness,
|
||||||
|
surfaceIntensity,
|
||||||
|
speech = null,
|
||||||
size = 220,
|
size = 220,
|
||||||
className,
|
className,
|
||||||
animated = true,
|
animated = true,
|
||||||
|
|
@ -222,7 +237,29 @@ export function ClientAvatar({
|
||||||
showMeta = true,
|
showMeta = true,
|
||||||
}: ClientAvatarProps) {
|
}: ClientAvatarProps) {
|
||||||
const reduced = useReducedMotion();
|
const reduced = useReducedMotion();
|
||||||
const motionEnabled = animated && !reduced;
|
|
||||||
|
/* v3(리노컷) 분기(결정문 §8.5) — persona.code로 판정(지금은 P1만). 리그가 있으면
|
||||||
|
그림 영역을 ClientAvatarV3로 그린다. 로드가 끝날 때까지 기존 SVG를 그대로 보이고,
|
||||||
|
ready가 되면 300ms 불투명도 전환으로 v3를 위에 올린 뒤 기존 SVG를 언마운트해
|
||||||
|
rAF를 멈춘다. error면 v3를 언마운트하고 기존 SVG로 남는다(§8.4 로드 실패 규칙). */
|
||||||
|
const rig = useMemo(() => linocutRigFor(persona.code), [persona.code]);
|
||||||
|
const [v3LoadState, setV3LoadState] = useState<"loading" | "ready" | "error">("loading");
|
||||||
|
const [legacyMounted, setLegacyMounted] = useState(true);
|
||||||
|
|
||||||
|
useEffect(() => {
|
||||||
|
setV3LoadState("loading");
|
||||||
|
setLegacyMounted(true);
|
||||||
|
}, [rig]);
|
||||||
|
|
||||||
|
useEffect(() => {
|
||||||
|
if (v3LoadState !== "ready") return;
|
||||||
|
const transitionMs = reduced ? 0 : 300;
|
||||||
|
const timer = window.setTimeout(() => setLegacyMounted(false), transitionMs);
|
||||||
|
return () => window.clearTimeout(timer);
|
||||||
|
}, [v3LoadState, reduced]);
|
||||||
|
|
||||||
|
const v3Active = rig !== null && v3LoadState !== "error";
|
||||||
|
const motionEnabled = animated && !reduced && legacyMounted;
|
||||||
|
|
||||||
// 외형 안전값
|
// 외형 안전값
|
||||||
const skin = persona.skinTone;
|
const skin = persona.skinTone;
|
||||||
|
|
@ -275,6 +312,10 @@ export function ClientAvatar({
|
||||||
const shoulderRotate = params.shoulderTurn * 0.4;
|
const shoulderRotate = params.shoulderTurn * 0.4;
|
||||||
const mouthOpen = Math.max(mouth, params.mouthOpen);
|
const mouthOpen = Math.max(mouth, params.mouthOpen);
|
||||||
|
|
||||||
|
// v3 전용 보정값(결정문 §8.5): 없으면 openness는 rapport로, 강도는 0.5로 둔다.
|
||||||
|
const v3Openness = openness ?? rapport ?? 0.5;
|
||||||
|
const v3SurfaceIntensity = surfaceIntensity ?? 0.5;
|
||||||
|
|
||||||
return (
|
return (
|
||||||
<figure
|
<figure
|
||||||
className={"vg-avatar" + (className ? " " + className : "")}
|
className={"vg-avatar" + (className ? " " + className : "")}
|
||||||
|
|
@ -307,66 +348,84 @@ export function ClientAvatar({
|
||||||
) : null}
|
) : null}
|
||||||
|
|
||||||
<div className="vg-avatar__stage" style={{ height: size }}>
|
<div className="vg-avatar__stage" style={{ height: size }}>
|
||||||
{/* 호흡하는 광배 */}
|
{legacyMounted ? (
|
||||||
<AuraLayer
|
<>
|
||||||
hue={params.auraHue}
|
{/* 호흡하는 광배 */}
|
||||||
opacity={params.auraOpacity}
|
<AuraLayer
|
||||||
saturation={age.auraSaturation}
|
hue={params.auraHue}
|
||||||
state={state}
|
opacity={params.auraOpacity}
|
||||||
mouth={mouth}
|
saturation={age.auraSaturation}
|
||||||
reduced={reduced}
|
state={state}
|
||||||
/>
|
mouth={mouth}
|
||||||
|
reduced={reduced}
|
||||||
|
/>
|
||||||
|
|
||||||
<svg
|
<svg
|
||||||
className="vg-avatar__svg"
|
className="vg-avatar__svg"
|
||||||
viewBox="0 0 200 200"
|
viewBox="0 0 200 200"
|
||||||
width={size}
|
width={size}
|
||||||
height={size}
|
height={size}
|
||||||
role="img"
|
role="img"
|
||||||
aria-hidden="true"
|
aria-hidden="true"
|
||||||
>
|
>
|
||||||
{/* 흉상 그룹: 어깨 호흡(translateY) + 저항 시 미세 회전 */}
|
{/* 흉상 그룹: 어깨 호흡(translateY) + 저항 시 미세 회전 */}
|
||||||
<g transform={`translate(0 ${-breath}) rotate(${shoulderRotate} 100 150)`}>
|
<g transform={`translate(0 ${-breath}) rotate(${shoulderRotate} 100 150)`}>
|
||||||
{/* 어깨/상반신 실루엣 */}
|
{/* 어깨/상반신 실루엣 */}
|
||||||
<BodySilhouette color={outfitColor} shoulderTurn={params.shoulderTurn} />
|
<BodySilhouette color={outfitColor} shoulderTurn={params.shoulderTurn} />
|
||||||
|
|
||||||
<g transform={`rotate(${params.headTilt} 100 101)`}>
|
<g transform={`rotate(${params.headTilt} 100 101)`}>
|
||||||
<HairBack style={hairStyle} color={hairColor} jawWidth={age.jawWidth} />
|
<HairBack style={hairStyle} color={hairColor} jawWidth={age.jawWidth} />
|
||||||
<NeckBridge skin={skin} />
|
<NeckBridge skin={skin} />
|
||||||
|
|
||||||
{/* 머리 (양식화 — 코·주름·모공 없음) */}
|
{/* 머리 (양식화 — 코·주름·모공 없음) */}
|
||||||
<ellipse cx="100" cy="92" rx={42 * age.jawWidth} ry="46" fill={skin} />
|
<ellipse cx="100" cy="92" rx={42 * age.jawWidth} ry="46" fill={skin} />
|
||||||
|
|
||||||
<HairFront style={hairStyle} color={hairColor} />
|
<HairFront style={hairStyle} color={hairColor} />
|
||||||
|
|
||||||
{/* 얼굴 그룹: 시선 회피/saccade (translate) */}
|
{/* 얼굴 그룹: 시선 회피/saccade (translate) */}
|
||||||
<g transform={`translate(${gazeX} ${gazeY})`}>
|
<g transform={`translate(${gazeX} ${gazeY})`}>
|
||||||
<Brows
|
<Brows
|
||||||
browTilt={params.browTilt}
|
browTilt={params.browTilt}
|
||||||
browLift={params.browLift}
|
browLift={params.browLift}
|
||||||
browPinch={params.browPinch}
|
browPinch={params.browPinch}
|
||||||
color={hairColor}
|
color={hairColor}
|
||||||
/>
|
/>
|
||||||
<Eyes
|
<Eyes
|
||||||
eyeSize={age.eyeSize}
|
eyeSize={age.eyeSize}
|
||||||
blink={blink}
|
blink={blink}
|
||||||
eyeOpen={params.eyeOpen}
|
eyeOpen={params.eyeOpen}
|
||||||
eyelidDrop={params.eyelidDrop}
|
eyelidDrop={params.eyelidDrop}
|
||||||
pupilScale={params.pupilScale}
|
pupilScale={params.pupilScale}
|
||||||
irisColor={irisColor}
|
irisColor={irisColor}
|
||||||
/>
|
/>
|
||||||
<Mouth
|
<Mouth
|
||||||
open={mouthOpen}
|
open={mouthOpen}
|
||||||
curve={params.mouthCurve}
|
curve={params.mouthCurve}
|
||||||
width={params.mouthWidth}
|
width={params.mouthWidth}
|
||||||
tension={params.mouthTension}
|
tension={params.mouthTension}
|
||||||
color="#9B5B52"
|
color="#9B5B52"
|
||||||
/>
|
/>
|
||||||
|
</g>
|
||||||
|
</g>
|
||||||
</g>
|
</g>
|
||||||
</g>
|
</svg>
|
||||||
</g>
|
</>
|
||||||
</svg>
|
) : null}
|
||||||
|
|
||||||
|
{rig && v3Active ? (
|
||||||
|
<ClientAvatarV3
|
||||||
|
rig={rig}
|
||||||
|
code={persona.code}
|
||||||
|
state={state}
|
||||||
|
affect={affect}
|
||||||
|
openness={v3Openness}
|
||||||
|
surfaceIntensity={v3SurfaceIntensity}
|
||||||
|
speech={speech}
|
||||||
|
running={animated}
|
||||||
|
reducedMotion={reduced}
|
||||||
|
onLoadStateChange={setV3LoadState}
|
||||||
|
/>
|
||||||
|
) : null}
|
||||||
</div>
|
</div>
|
||||||
|
|
||||||
{/* 페르소나 메타 + 상태 텍스트 */}
|
{/* 페르소나 메타 + 상태 텍스트 */}
|
||||||
|
|
|
||||||
153
apps/web/src/components/avatar/engine/speechDriver.ts
Normal file
153
apps/web/src/components/avatar/engine/speechDriver.ts
Normal file
|
|
@ -0,0 +1,153 @@
|
||||||
|
/* =====================================================================
|
||||||
|
아바타 v3 발화 구동 공용 모듈 — 결정문 §8.5 "세션 연결". AvatarLab의
|
||||||
|
runSpeechTimeline / playSpeech / playSpeechWithAudioFile / endSpeech가
|
||||||
|
하던 엔진 구동 부분(buildPerformance → buildSpeechTimeline[+포락선]
|
||||||
|
→ buildCoSpeechPlan → playPerformance → 매 프레임 표본)을 여기로 옮겨
|
||||||
|
Lab과 v3 래퍼(ClientAvatarV3)가 함께 쓴다. 화면 표시(ref 갱신)는 onFrame
|
||||||
|
콜백으로 호출자에 남긴다. 발화 종료 시 상태 복귀(Lab은 listening, 래퍼는
|
||||||
|
prop state)도 호출자가 onEnd에서 정한다 — 이 모듈은 setSpeechShape(null)·
|
||||||
|
setSpeechMotion(null)과 rAF 해제까지만 한다.
|
||||||
|
===================================================================== */
|
||||||
|
|
||||||
|
import type { AvatarExpression } from "../persona";
|
||||||
|
import type { AvatarEngine } from "./engine";
|
||||||
|
import type { SpeechStyle } from "./demeanorDefaults";
|
||||||
|
import { buildPerformance, type PerformanceCue } from "./performance";
|
||||||
|
import {
|
||||||
|
buildSpeechTimeline,
|
||||||
|
currentViseme,
|
||||||
|
sampleSpeech,
|
||||||
|
type PhraseKind,
|
||||||
|
type SpeechShape,
|
||||||
|
type SpeechTimeline,
|
||||||
|
type VisemeId,
|
||||||
|
} from "./lipsync";
|
||||||
|
import { buildCoSpeechPlan, isStressPulseActive, sampleCoSpeech } from "./coSpeech";
|
||||||
|
import { computeEnvelope, type SpeechEnvelope } from "./speechEnvelope";
|
||||||
|
|
||||||
|
export interface SpeechDriverAudio {
|
||||||
|
buffer: AudioBuffer;
|
||||||
|
context: BaseAudioContext;
|
||||||
|
/** source.start(when)의 when(초, context 시계). */
|
||||||
|
startAt: number;
|
||||||
|
}
|
||||||
|
|
||||||
|
export interface SpeechFrameSample {
|
||||||
|
viseme: VisemeId;
|
||||||
|
shape: SpeechShape;
|
||||||
|
phraseKind: PhraseKind | null;
|
||||||
|
stressed: boolean;
|
||||||
|
}
|
||||||
|
|
||||||
|
export interface StartSpeechParams {
|
||||||
|
engine: AvatarEngine;
|
||||||
|
/** 괄호 지문을 포함한 발화 원문. */
|
||||||
|
text: string;
|
||||||
|
expression: AvatarExpression;
|
||||||
|
intensity: number;
|
||||||
|
openness: number;
|
||||||
|
seed: number;
|
||||||
|
speech: SpeechStyle;
|
||||||
|
/** 엔진 시계(engine.evaluate에 넘기는 시계와 같은 시계). */
|
||||||
|
now: () => number;
|
||||||
|
/** 있으면 오디오 선분석 정렬 경로. 없으면 텍스트 타이밍(speech.syllablesPerSec). */
|
||||||
|
audio?: SpeechDriverAudio;
|
||||||
|
/** 매 프레임 표본 결과(화면 표시용). React state 갱신은 호출자 책임(ref로 받는다). */
|
||||||
|
onFrame?: (sample: SpeechFrameSample) => void;
|
||||||
|
/** 타임라인이 끝까지 재생됐을 때(stop() 호출로 끝난 경우에는 부르지 않는다). */
|
||||||
|
onEnd?: () => void;
|
||||||
|
}
|
||||||
|
|
||||||
|
export interface SpeechHandle {
|
||||||
|
stop(): void;
|
||||||
|
/** 표시·디버그용 — buildPerformance가 파싱한 cue와 미대응 지문. */
|
||||||
|
cues: PerformanceCue[];
|
||||||
|
unmatched: string[];
|
||||||
|
}
|
||||||
|
|
||||||
|
/** tMs 시점까지 시작한 가장 최근 구의 종류. 구 사이 휴지 중에는 그 직전 구를 보인다. */
|
||||||
|
function currentPhraseKind(timeline: SpeechTimeline, localMs: number): PhraseKind | null {
|
||||||
|
let kind: PhraseKind | null = null;
|
||||||
|
for (const ph of timeline.phrases) {
|
||||||
|
if (ph.startMs > localMs) break;
|
||||||
|
kind = ph.kind;
|
||||||
|
}
|
||||||
|
return kind;
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* 발화 하나를 시작한다. 오디오가 있으면 포락선을 선분석해 오디오 시계로 표본하고,
|
||||||
|
* 없으면 텍스트 타이밍으로 엔진 시계에서 300ms 뒤 시작한다(결정문 §5.4·§8.5).
|
||||||
|
* 지문은 playPerformance로 함께 재생한다. 반환된 stop()은 호출자가 발화를 바꾸거나
|
||||||
|
* 끝낼 때 부른다(speechShape·speechMotion을 null로 되돌리고 rAF를 해제, 상태 복귀는 않는다).
|
||||||
|
*/
|
||||||
|
export function startSpeech(params: StartSpeechParams): SpeechHandle {
|
||||||
|
const { engine, text, expression, intensity, openness, seed, speech, now, audio, onFrame, onEnd } = params;
|
||||||
|
const nowMs = now();
|
||||||
|
const { performance: perf, unmatched } = buildPerformance({ text, expression, intensity, openness, seed });
|
||||||
|
|
||||||
|
let envelope: SpeechEnvelope | undefined;
|
||||||
|
let timeline: SpeechTimeline;
|
||||||
|
let speechStartMs: number;
|
||||||
|
let sampleClock: () => number;
|
||||||
|
|
||||||
|
if (audio) {
|
||||||
|
const channels: Float32Array[] = [];
|
||||||
|
for (let c = 0; c < audio.buffer.numberOfChannels; c++) channels.push(audio.buffer.getChannelData(c));
|
||||||
|
envelope = computeEnvelope(channels, audio.buffer.sampleRate);
|
||||||
|
timeline = buildSpeechTimeline({ text, syllablesPerSec: speech.syllablesPerSec, envelope });
|
||||||
|
speechStartMs = nowMs + (audio.startAt - audio.context.currentTime) * 1000;
|
||||||
|
/* 오디오 시계 기준(초 → ms, §8.5 "매 프레임 표본 시각은 (context.currentTime − startAt)·1000"). */
|
||||||
|
sampleClock = () => (audio.context.currentTime - audio.startAt) * 1000 + speechStartMs;
|
||||||
|
} else {
|
||||||
|
timeline = buildSpeechTimeline({ text, syllablesPerSec: speech.syllablesPerSec });
|
||||||
|
speechStartMs = nowMs + 300;
|
||||||
|
sampleClock = now;
|
||||||
|
}
|
||||||
|
|
||||||
|
const coSpeechPlan = buildCoSpeechPlan(timeline, speech, seed);
|
||||||
|
engine.playPerformance(perf, { speechStartMs, speechDurationMs: timeline.totalDurationMs }, nowMs);
|
||||||
|
|
||||||
|
let rafId = 0;
|
||||||
|
let prevLocalMs = -Infinity;
|
||||||
|
const endMs = speechStartMs + timeline.totalDurationMs;
|
||||||
|
|
||||||
|
function stop(): void {
|
||||||
|
cancelAnimationFrame(rafId);
|
||||||
|
engine.setSpeechShape(null);
|
||||||
|
engine.setSpeechMotion(null);
|
||||||
|
}
|
||||||
|
|
||||||
|
const step = () => {
|
||||||
|
const t = sampleClock();
|
||||||
|
if (t < speechStartMs) {
|
||||||
|
rafId = requestAnimationFrame(step);
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
if (t >= endMs) {
|
||||||
|
stop();
|
||||||
|
onEnd?.();
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
const localMs = t - speechStartMs;
|
||||||
|
const shape = sampleSpeech(timeline, localMs, speech.articulation);
|
||||||
|
engine.setSpeechShape(shape);
|
||||||
|
|
||||||
|
const sample = sampleCoSpeech(coSpeechPlan, localMs, prevLocalMs, envelope);
|
||||||
|
engine.setSpeechMotion(sample.delta);
|
||||||
|
if (sample.blinkNow) engine.requestSpeechBlink(now());
|
||||||
|
|
||||||
|
onFrame?.({
|
||||||
|
viseme: currentViseme(timeline, localMs),
|
||||||
|
shape,
|
||||||
|
phraseKind: currentPhraseKind(timeline, localMs),
|
||||||
|
stressed: isStressPulseActive(coSpeechPlan, localMs),
|
||||||
|
});
|
||||||
|
|
||||||
|
prevLocalMs = localMs;
|
||||||
|
rafId = requestAnimationFrame(step);
|
||||||
|
};
|
||||||
|
rafId = requestAnimationFrame(step);
|
||||||
|
|
||||||
|
return { stop, cues: perf.cues, unmatched };
|
||||||
|
}
|
||||||
210
apps/web/src/components/avatar/v3/ClientAvatarV3.tsx
Normal file
210
apps/web/src/components/avatar/v3/ClientAvatarV3.tsx
Normal file
|
|
@ -0,0 +1,210 @@
|
||||||
|
/* =====================================================================
|
||||||
|
ClientAvatarV3 — 리노컷 리그가 있는 페르소나를 위한 v3 그림 영역 래퍼.
|
||||||
|
결정문 §8.5 "세션 연결". ClientAvatar가 그림 영역만 이 컴포넌트로 바꿔
|
||||||
|
그린다(루트 data-*·캡션·메타는 ClientAvatar가 그대로 유지한다).
|
||||||
|
|
||||||
|
엔진은 이 컴포넌트가 소유한다(persona.code로 demeanor·시드를 만든다).
|
||||||
|
시계는 모듈 수준 상수 함수(engineNowMs)를 LinocutAvatar에 안정 참조로
|
||||||
|
넘긴다 — 인라인 화살표를 넘기면 리렌더마다 rAF가 다시 시작된다.
|
||||||
|
크롭(bust/face)은 이 래퍼의 실측 폭(ResizeObserver)으로 정한다 —
|
||||||
|
session.css의 !important 규칙이 실제 렌더 크기를 덮어쓰기 때문에 size
|
||||||
|
prop은 쓰지 않는다(§8.5).
|
||||||
|
===================================================================== */
|
||||||
|
|
||||||
|
import { useEffect, useMemo, useRef, useState } from "react";
|
||||||
|
import type { AvatarExpression, AvatarState } from "../persona";
|
||||||
|
import { createAvatarEngine } from "../engine/engine";
|
||||||
|
import { demeanorFor } from "../engine/demeanorDefaults";
|
||||||
|
import { hashString } from "../engine/rng";
|
||||||
|
import { startSpeech, type SpeechDriverAudio } from "../engine/speechDriver";
|
||||||
|
import type { VisemeId } from "../engine/lipsync";
|
||||||
|
import LinocutAvatar from "./LinocutAvatar";
|
||||||
|
import { observeWidth } from "./observeWidth";
|
||||||
|
import type { LinocutRig, RigCrop } from "./linocutRig";
|
||||||
|
import "./client-avatar-v3.css";
|
||||||
|
|
||||||
|
/** ClientAvatar의 speech? prop 타입 — ClientAvatar 모듈에서 재노출한다. */
|
||||||
|
export interface AvatarSpeech {
|
||||||
|
id: string;
|
||||||
|
/** 괄호 지문을 포함한 답변 원문. */
|
||||||
|
text: string;
|
||||||
|
/** 있으면 오디오 선분석 경로. startAt은 source.start(when)의 when(초, context 시계). */
|
||||||
|
audio?: SpeechDriverAudio;
|
||||||
|
}
|
||||||
|
|
||||||
|
export interface ClientAvatarV3Props {
|
||||||
|
rig: LinocutRig;
|
||||||
|
code: string | null | undefined;
|
||||||
|
state: AvatarState;
|
||||||
|
affect: AvatarExpression;
|
||||||
|
/** 호출부가 openness ?? rapport ?? 0.5 로 보정해 넘긴다. */
|
||||||
|
openness: number;
|
||||||
|
/** 호출부가 surfaceIntensity ?? 0.5 로 보정해 넘긴다. */
|
||||||
|
surfaceIntensity: number;
|
||||||
|
speech?: AvatarSpeech | null;
|
||||||
|
/** false면 rAF를 멈추고 정지한 한 프레임만 그린다(시작 전 화면·썸네일). */
|
||||||
|
running: boolean;
|
||||||
|
reducedMotion: boolean;
|
||||||
|
className?: string;
|
||||||
|
onLoadStateChange?: (state: "loading" | "ready" | "error") => void;
|
||||||
|
}
|
||||||
|
|
||||||
|
const CROP_BUST_MIN_WIDTH = 120;
|
||||||
|
|
||||||
|
/** 엔진 시계는 performance.now()다. 안정 참조로 LinocutAvatar의 nowMs에 넘긴다
|
||||||
|
(인라인 화살표를 넘기면 리렌더마다 rAF가 다시 시작된다, §8.5). */
|
||||||
|
function engineNowMs(): number {
|
||||||
|
return performance.now();
|
||||||
|
}
|
||||||
|
|
||||||
|
export default function ClientAvatarV3({
|
||||||
|
rig,
|
||||||
|
code,
|
||||||
|
state,
|
||||||
|
affect,
|
||||||
|
openness,
|
||||||
|
surfaceIntensity,
|
||||||
|
speech = null,
|
||||||
|
running,
|
||||||
|
reducedMotion,
|
||||||
|
className,
|
||||||
|
onLoadStateChange,
|
||||||
|
}: ClientAvatarV3Props) {
|
||||||
|
const demeanor = useMemo(() => demeanorFor(code), [code]);
|
||||||
|
const seed = useMemo(() => hashString(code ?? ""), [code]);
|
||||||
|
const engine = useMemo(
|
||||||
|
() => createAvatarEngine({ demeanor, seed, reducedMotion }),
|
||||||
|
[demeanor, seed, reducedMotion],
|
||||||
|
);
|
||||||
|
|
||||||
|
const stateRef = useRef(state);
|
||||||
|
const speakingRef = useRef(false);
|
||||||
|
|
||||||
|
/* 엔진이 바뀌면(페르소나·reduced motion 전환) 지금 값으로 한 번에 다시 맞춘다
|
||||||
|
(AvatarLab.tsx의 같은 패턴). */
|
||||||
|
useEffect(() => {
|
||||||
|
engine.setState(speakingRef.current ? "speaking" : state, engineNowMs());
|
||||||
|
engine.setOpenness(openness);
|
||||||
|
engine.setSurface(affect, surfaceIntensity);
|
||||||
|
// eslint-disable-next-line react-hooks/exhaustive-deps
|
||||||
|
}, [engine]);
|
||||||
|
|
||||||
|
useEffect(() => {
|
||||||
|
stateRef.current = state;
|
||||||
|
if (!speakingRef.current) engine.setState(state, engineNowMs());
|
||||||
|
}, [engine, state]);
|
||||||
|
|
||||||
|
useEffect(() => {
|
||||||
|
engine.setOpenness(openness);
|
||||||
|
}, [engine, openness]);
|
||||||
|
|
||||||
|
useEffect(() => {
|
||||||
|
engine.setSurface(affect, surfaceIntensity);
|
||||||
|
}, [engine, affect, surfaceIntensity]);
|
||||||
|
|
||||||
|
/* 래퍼 박스 루트 — 크롭 판정(ResizeObserver)과 발화 data-* 속성 쓰기가 함께 쓴다. */
|
||||||
|
const rootRef = useRef<HTMLDivElement | null>(null);
|
||||||
|
|
||||||
|
/* 크롭(§8.5): 래퍼 박스의 실측 폭. 120px 이상이면 bust, 미만이면 face다. */
|
||||||
|
const [crop, setCrop] = useState<RigCrop>("face");
|
||||||
|
useEffect(() => {
|
||||||
|
const el = rootRef.current;
|
||||||
|
if (!el) return;
|
||||||
|
return observeWidth(el, (widthPx) => setCrop(widthPx >= CROP_BUST_MIN_WIDTH ? "bust" : "face"));
|
||||||
|
}, []);
|
||||||
|
|
||||||
|
const [loadState, setLoadState] = useState<"loading" | "ready" | "error">("loading");
|
||||||
|
function handleLoadStateChange(next: "loading" | "ready" | "error"): void {
|
||||||
|
setLoadState(next);
|
||||||
|
onLoadStateChange?.(next);
|
||||||
|
}
|
||||||
|
|
||||||
|
/* 발화(§8.5·§5.4·§5.5): id가 바뀌면 이전 발화를 멈추고 새 발화를 시작한다. 발화가
|
||||||
|
진행 중이면 엔진 상태를 "speaking"으로 유지하고, 끝나면 그때의 prop state로
|
||||||
|
되돌린다. 프레임마다 바뀌는 data-viseme은 React state가 아니라 ref로 쓴다. */
|
||||||
|
const speechHandleRef = useRef<ReturnType<typeof startSpeech> | null>(null);
|
||||||
|
const activeSpeechIdRef = useRef<string | null>(null);
|
||||||
|
|
||||||
|
function writeSpeechAttrs(source: "audio" | "text" | null, viseme: VisemeId): void {
|
||||||
|
const el = rootRef.current;
|
||||||
|
if (!el) return;
|
||||||
|
if (source) {
|
||||||
|
el.setAttribute("data-speech-source", source);
|
||||||
|
el.setAttribute("data-viseme", viseme);
|
||||||
|
} else {
|
||||||
|
el.removeAttribute("data-speech-source");
|
||||||
|
el.setAttribute("data-viseme", "X");
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
useEffect(() => {
|
||||||
|
if (!speech) {
|
||||||
|
if (activeSpeechIdRef.current !== null) {
|
||||||
|
speechHandleRef.current?.stop();
|
||||||
|
speechHandleRef.current = null;
|
||||||
|
activeSpeechIdRef.current = null;
|
||||||
|
speakingRef.current = false;
|
||||||
|
engine.setState(stateRef.current, engineNowMs());
|
||||||
|
writeSpeechAttrs(null, "X");
|
||||||
|
}
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
if (speech.id === activeSpeechIdRef.current) return;
|
||||||
|
|
||||||
|
speechHandleRef.current?.stop();
|
||||||
|
activeSpeechIdRef.current = speech.id;
|
||||||
|
speakingRef.current = true;
|
||||||
|
engine.setState("speaking", engineNowMs());
|
||||||
|
const source: "audio" | "text" = speech.audio ? "audio" : "text";
|
||||||
|
writeSpeechAttrs(source, "X");
|
||||||
|
|
||||||
|
speechHandleRef.current = startSpeech({
|
||||||
|
engine,
|
||||||
|
text: speech.text,
|
||||||
|
expression: affect,
|
||||||
|
intensity: surfaceIntensity,
|
||||||
|
openness,
|
||||||
|
seed,
|
||||||
|
speech: demeanor.speech,
|
||||||
|
now: engineNowMs,
|
||||||
|
audio: speech.audio,
|
||||||
|
onFrame: ({ viseme }) => writeSpeechAttrs(source, viseme),
|
||||||
|
onEnd: () => {
|
||||||
|
speakingRef.current = false;
|
||||||
|
speechHandleRef.current = null;
|
||||||
|
activeSpeechIdRef.current = null;
|
||||||
|
engine.setState(stateRef.current, engineNowMs());
|
||||||
|
writeSpeechAttrs(null, "X");
|
||||||
|
},
|
||||||
|
});
|
||||||
|
// eslint-disable-next-line react-hooks/exhaustive-deps
|
||||||
|
}, [speech]);
|
||||||
|
|
||||||
|
/* 언마운트 시 발화 rAF를 해제한다(페르소나 전환 등으로 이 래퍼 자체가 사라질 때). */
|
||||||
|
useEffect(() => {
|
||||||
|
return () => speechHandleRef.current?.stop();
|
||||||
|
}, []);
|
||||||
|
|
||||||
|
const transitionMs = reducedMotion ? 0 : 300;
|
||||||
|
|
||||||
|
return (
|
||||||
|
<div
|
||||||
|
ref={rootRef}
|
||||||
|
className={"vg-avatar-v3" + (className ? " " + className : "")}
|
||||||
|
data-avatar-renderer="linocut"
|
||||||
|
data-load-state={loadState}
|
||||||
|
data-viseme="X"
|
||||||
|
style={{ opacity: loadState === "ready" ? 1 : 0, transitionDuration: `${transitionMs}ms` }}
|
||||||
|
>
|
||||||
|
<LinocutAvatar
|
||||||
|
rig={rig}
|
||||||
|
engine={engine}
|
||||||
|
running={running}
|
||||||
|
nowMs={engineNowMs}
|
||||||
|
crop={crop}
|
||||||
|
backdropExpression={affect}
|
||||||
|
onLoadStateChange={handleLoadStateChange}
|
||||||
|
/>
|
||||||
|
</div>
|
||||||
|
);
|
||||||
|
}
|
||||||
|
|
@ -23,6 +23,7 @@ import {
|
||||||
type LinocutFrame,
|
type LinocutFrame,
|
||||||
type JawWarpLayer,
|
type JawWarpLayer,
|
||||||
} from "./linocutGeometry";
|
} from "./linocutGeometry";
|
||||||
|
import { observeWidth } from "./observeWidth";
|
||||||
import "./linocut-avatar.css";
|
import "./linocut-avatar.css";
|
||||||
|
|
||||||
export interface LinocutAvatarProps {
|
export interface LinocutAvatarProps {
|
||||||
|
|
@ -442,25 +443,23 @@ export default function LinocutAvatar({
|
||||||
const el = svgRef.current;
|
const el = svgRef.current;
|
||||||
if (!el) return;
|
if (!el) return;
|
||||||
const viewBoxWidth = rig.crops[crop][2];
|
const viewBoxWidth = rig.crops[crop][2];
|
||||||
const update = (widthPx: number) => {
|
return observeWidth(el, (widthPx) => {
|
||||||
if (widthPx <= 0) return;
|
|
||||||
const k = widthPx / viewBoxWidth;
|
const k = widthPx / viewBoxWidth;
|
||||||
sizeScaleRef.current = Math.min(2, Math.max(1, 0.3 / k));
|
sizeScaleRef.current = Math.min(2, Math.max(1, 0.3 / k));
|
||||||
};
|
|
||||||
update(el.getBoundingClientRect().width);
|
|
||||||
const observer = new ResizeObserver((entries) => {
|
|
||||||
for (const entry of entries) {
|
|
||||||
const w = entry.contentBoxSize?.[0]?.inlineSize ?? entry.contentRect.width;
|
|
||||||
update(w);
|
|
||||||
}
|
|
||||||
});
|
});
|
||||||
observer.observe(el);
|
|
||||||
return () => observer.disconnect();
|
|
||||||
}, [rig, crop]);
|
}, [rig, crop]);
|
||||||
|
|
||||||
|
/* running=false(§8.5 세션 연결 — 시작 전 화면·썸네일)여도 정지한 한 프레임은 계산해
|
||||||
|
그린다. 그렇지 않으면 벡터 부위(눈·눈썹·입)가 비어 보인다. 마운트할 때, loadState가
|
||||||
|
ready가 될 때, engine·rig·crop이 바뀔 때 다시 계산한다(crop·loadState를 deps에
|
||||||
|
넣는다). running=true면 그대로 rAF 루프를 돈다. */
|
||||||
useEffect(() => {
|
useEffect(() => {
|
||||||
if (!running) return;
|
|
||||||
const clock = nowMs ?? (() => performance.now());
|
const clock = nowMs ?? (() => performance.now());
|
||||||
|
if (!running) {
|
||||||
|
const v = engine.evaluate(clock());
|
||||||
|
applyFrame(refs.current, computeLinocutFrame(rig, v, sizeScaleRef.current));
|
||||||
|
return;
|
||||||
|
}
|
||||||
const loop = () => {
|
const loop = () => {
|
||||||
const v = engine.evaluate(clock());
|
const v = engine.evaluate(clock());
|
||||||
const frame = computeLinocutFrame(rig, v, sizeScaleRef.current);
|
const frame = computeLinocutFrame(rig, v, sizeScaleRef.current);
|
||||||
|
|
@ -469,7 +468,7 @@ export default function LinocutAvatar({
|
||||||
};
|
};
|
||||||
rafRef.current = requestAnimationFrame(loop);
|
rafRef.current = requestAnimationFrame(loop);
|
||||||
return () => cancelAnimationFrame(rafRef.current);
|
return () => cancelAnimationFrame(rafRef.current);
|
||||||
}, [engine, running, nowMs, rig]);
|
}, [engine, running, nowMs, rig, crop, loadState]);
|
||||||
|
|
||||||
const viewBox = rig.crops[crop].join(" ");
|
const viewBox = rig.crops[crop].join(" ");
|
||||||
const faceClipId = `${uid}-face-clip`;
|
const faceClipId = `${uid}-face-clip`;
|
||||||
|
|
|
||||||
20
apps/web/src/components/avatar/v3/client-avatar-v3.css
Normal file
20
apps/web/src/components/avatar/v3/client-avatar-v3.css
Normal file
|
|
@ -0,0 +1,20 @@
|
||||||
|
/* =====================================================================
|
||||||
|
client-avatar-v3.css — ClientAvatarV3 래퍼 스타일. 결정문 §8.5 "화면".
|
||||||
|
|
||||||
|
원형 오브 안(.vg-avatar__stage, position:relative)에 꽉 차게 올라가
|
||||||
|
기존 SVG와 정확히 같은 자리를 차지한다(inset:0) — session.css의
|
||||||
|
!important 크기 규칙이 스테이지 쪽에 이미 걸려 있어 이 규칙만으로
|
||||||
|
모든 브레이크포인트에서 기존 SVG와 같은 크기가 된다. 원형이 아닌 곳에
|
||||||
|
이 컴포넌트를 쓰면 border-radius는 그대로 두되 바깥에서 inset을
|
||||||
|
덮어써야 한다(지금은 호출부가 모두 원형 오브다).
|
||||||
|
===================================================================== */
|
||||||
|
|
||||||
|
.vg-avatar-v3 {
|
||||||
|
position: absolute;
|
||||||
|
inset: 0;
|
||||||
|
z-index: 1;
|
||||||
|
border-radius: 50%;
|
||||||
|
overflow: hidden;
|
||||||
|
transition-property: opacity;
|
||||||
|
transition-timing-function: ease;
|
||||||
|
}
|
||||||
15
apps/web/src/components/avatar/v3/observeWidth.ts
Normal file
15
apps/web/src/components/avatar/v3/observeWidth.ts
Normal file
|
|
@ -0,0 +1,15 @@
|
||||||
|
/** 요소의 렌더 폭(px)을 지금 한 번, 이후 크기가 바뀔 때마다 onWidth로 알린다(프레임마다 아님).
|
||||||
|
폭이 0 이하인 측정(아직 배치 전)은 건너뛴다. 반환 함수로 관찰을 끝낸다. */
|
||||||
|
export function observeWidth(el: Element, onWidth: (widthPx: number) => void): () => void {
|
||||||
|
const report = (widthPx: number) => {
|
||||||
|
if (widthPx > 0) onWidth(widthPx);
|
||||||
|
};
|
||||||
|
report(el.getBoundingClientRect().width);
|
||||||
|
const observer = new ResizeObserver((entries) => {
|
||||||
|
for (const entry of entries) {
|
||||||
|
report(entry.contentBoxSize?.[0]?.inlineSize ?? entry.contentRect.width);
|
||||||
|
}
|
||||||
|
});
|
||||||
|
observer.observe(el);
|
||||||
|
return () => observer.disconnect();
|
||||||
|
}
|
||||||
16
apps/web/src/components/avatar/v3/rigs/index.ts
Normal file
16
apps/web/src/components/avatar/v3/rigs/index.ts
Normal file
|
|
@ -0,0 +1,16 @@
|
||||||
|
/* =====================================================================
|
||||||
|
아바타 v3 리그 레지스트리 — 결정문 §8.5 "분기". 페르소나 코드 → 리노컷
|
||||||
|
리그. 지금은 P1만 있고 나머지는 null(기존 SVG로 남는다). 빌드 플래그
|
||||||
|
VITE_AVATAR_V3=0 이면 항상 null(기본은 켜짐).
|
||||||
|
===================================================================== */
|
||||||
|
|
||||||
|
import type { LinocutRig } from "../linocutRig";
|
||||||
|
import { P1_LINOCUT_RIG } from "./p1Rig";
|
||||||
|
|
||||||
|
const LINOCUT_RIGS: Partial<Record<string, LinocutRig>> = { P1: P1_LINOCUT_RIG };
|
||||||
|
|
||||||
|
export function linocutRigFor(code: string | null | undefined): LinocutRig | null {
|
||||||
|
if (import.meta.env.VITE_AVATAR_V3 === "0") return null;
|
||||||
|
if (!code) return null;
|
||||||
|
return LINOCUT_RIGS[code.toUpperCase()] ?? null;
|
||||||
|
}
|
||||||
|
|
@ -14,18 +14,9 @@ import { CHANNEL_IDS } from "../components/avatar/engine/channels";
|
||||||
import { createAvatarEngine, type AvatarEngine, type DebugSnapshot } from "../components/avatar/engine/engine";
|
import { createAvatarEngine, type AvatarEngine, type DebugSnapshot } from "../components/avatar/engine/engine";
|
||||||
import { REACTION_CLIPS, REACTION_CLIP_IDS, type ReactionClipId } from "../components/avatar/engine/clipCatalog";
|
import { REACTION_CLIPS, REACTION_CLIP_IDS, type ReactionClipId } from "../components/avatar/engine/clipCatalog";
|
||||||
import { demeanorFor } from "../components/avatar/engine/demeanorDefaults";
|
import { demeanorFor } from "../components/avatar/engine/demeanorDefaults";
|
||||||
import { buildPerformance, type Performance, type PerformanceCue } from "../components/avatar/engine/performance";
|
import type { Performance, PerformanceCue } from "../components/avatar/engine/performance";
|
||||||
import {
|
import type { PhraseKind, SpeechShape, VisemeId } from "../components/avatar/engine/lipsync";
|
||||||
buildSpeechTimeline,
|
import { startSpeech, type SpeechHandle } from "../components/avatar/engine/speechDriver";
|
||||||
currentViseme,
|
|
||||||
sampleSpeech,
|
|
||||||
type PhraseKind,
|
|
||||||
type SpeechShape,
|
|
||||||
type SpeechTimeline,
|
|
||||||
type VisemeId,
|
|
||||||
} from "../components/avatar/engine/lipsync";
|
|
||||||
import { buildCoSpeechPlan, isStressPulseActive, sampleCoSpeech, type CoSpeechPlan } from "../components/avatar/engine/coSpeech";
|
|
||||||
import { computeEnvelope, type SpeechEnvelope } from "../components/avatar/engine/speechEnvelope";
|
|
||||||
import { AVATAR_EXPRESSION_LIBRARY, type AvatarExpression, type AvatarState } from "../components/avatar/persona";
|
import { AVATAR_EXPRESSION_LIBRARY, type AvatarExpression, type AvatarState } from "../components/avatar/persona";
|
||||||
import "./avatar-lab.css";
|
import "./avatar-lab.css";
|
||||||
|
|
||||||
|
|
@ -83,16 +74,6 @@ function createClock(): Clock {
|
||||||
};
|
};
|
||||||
}
|
}
|
||||||
|
|
||||||
/** 표시용: localMs 시점까지 시작한 가장 최근 구의 종류. 구 사이 휴지 중에는 그 직전 구를 보인다. */
|
|
||||||
function currentPhraseKind(timeline: SpeechTimeline, localMs: number): PhraseKind | null {
|
|
||||||
let kind: PhraseKind | null = null;
|
|
||||||
for (const ph of timeline.phrases) {
|
|
||||||
if (ph.startMs > localMs) break;
|
|
||||||
kind = ph.kind;
|
|
||||||
}
|
|
||||||
return kind;
|
|
||||||
}
|
|
||||||
|
|
||||||
function seedFromQuery(): number {
|
function seedFromQuery(): number {
|
||||||
const raw = new URLSearchParams(window.location.search).get("seed");
|
const raw = new URLSearchParams(window.location.search).get("seed");
|
||||||
const parsed = raw === null ? NaN : Number(raw);
|
const parsed = raw === null ? NaN : Number(raw);
|
||||||
|
|
@ -132,7 +113,7 @@ export default function AvatarLab() {
|
||||||
const [unmatched, setUnmatched] = useState<string[]>([]);
|
const [unmatched, setUnmatched] = useState<string[]>([]);
|
||||||
const [snapshot, setSnapshot] = useState<DebugSnapshot | null>(null);
|
const [snapshot, setSnapshot] = useState<DebugSnapshot | null>(null);
|
||||||
|
|
||||||
const speechRafRef = useRef<number>(0);
|
const speechHandleRef = useRef<SpeechHandle | null>(null);
|
||||||
const meterRefs = useRef<Record<string, HTMLTableCellElement | null>>({});
|
const meterRefs = useRef<Record<string, HTMLTableCellElement | null>>({});
|
||||||
|
|
||||||
/* 발화층(립싱크) — 원시 SpeechShape·비짐 표시용. 채널 미터와 달리 매 프레임 갱신하지
|
/* 발화층(립싱크) — 원시 SpeechShape·비짐 표시용. 채널 미터와 달리 매 프레임 갱신하지
|
||||||
|
|
@ -213,7 +194,7 @@ export default function AvatarLab() {
|
||||||
|
|
||||||
useEffect(() => {
|
useEffect(() => {
|
||||||
return () => {
|
return () => {
|
||||||
cancelAnimationFrame(speechRafRef.current);
|
speechHandleRef.current?.stop();
|
||||||
try {
|
try {
|
||||||
audioSourceRef.current?.stop();
|
audioSourceRef.current?.stop();
|
||||||
} catch {
|
} catch {
|
||||||
|
|
@ -247,76 +228,38 @@ export default function AvatarLab() {
|
||||||
coSpeechDisplayRef.current = { phraseKind: null, stressed: false };
|
coSpeechDisplayRef.current = { phraseKind: null, stressed: false };
|
||||||
}
|
}
|
||||||
|
|
||||||
/** timeline을 nowFn() 시계로 매 프레임 표본해 engine.setSpeechShape·setSpeechMotion에 흘려보낸다.
|
|
||||||
coSpeechPlan이 있으면 발화 동반층(§5.5)도 같은 프레임에 표본한다. */
|
|
||||||
function runSpeechTimeline(
|
|
||||||
timeline: SpeechTimeline,
|
|
||||||
speechStartMs: number,
|
|
||||||
nowFn: () => number,
|
|
||||||
articulation: number,
|
|
||||||
coSpeechPlan: CoSpeechPlan | null,
|
|
||||||
envelope?: SpeechEnvelope,
|
|
||||||
): void {
|
|
||||||
cancelAnimationFrame(speechRafRef.current);
|
|
||||||
const endMs = speechStartMs + timeline.totalDurationMs;
|
|
||||||
let prevLocalMs = -Infinity;
|
|
||||||
const step = () => {
|
|
||||||
const t = nowFn();
|
|
||||||
if (t < speechStartMs) {
|
|
||||||
speechRafRef.current = requestAnimationFrame(step);
|
|
||||||
return;
|
|
||||||
}
|
|
||||||
if (t >= endMs) {
|
|
||||||
endSpeech(clock.now());
|
|
||||||
return;
|
|
||||||
}
|
|
||||||
const localMs = t - speechStartMs;
|
|
||||||
const shape = sampleSpeech(timeline, localMs, articulation);
|
|
||||||
engine.setSpeechShape(shape);
|
|
||||||
speechDisplayRef.current = { viseme: currentViseme(timeline, localMs), shape };
|
|
||||||
|
|
||||||
if (coSpeechPlan) {
|
|
||||||
const sample = sampleCoSpeech(coSpeechPlan, localMs, prevLocalMs, envelope);
|
|
||||||
engine.setSpeechMotion(sample.delta);
|
|
||||||
if (sample.blinkNow) engine.requestSpeechBlink(clock.now());
|
|
||||||
coSpeechDisplayRef.current = {
|
|
||||||
phraseKind: currentPhraseKind(timeline, localMs),
|
|
||||||
stressed: isStressPulseActive(coSpeechPlan, localMs),
|
|
||||||
};
|
|
||||||
}
|
|
||||||
prevLocalMs = localMs;
|
|
||||||
|
|
||||||
speechRafRef.current = requestAnimationFrame(step);
|
|
||||||
};
|
|
||||||
speechRafRef.current = requestAnimationFrame(step);
|
|
||||||
}
|
|
||||||
|
|
||||||
function playSpeech(): void {
|
function playSpeech(): void {
|
||||||
|
speechHandleRef.current?.stop();
|
||||||
const nowMs = clock.now();
|
const nowMs = clock.now();
|
||||||
const { performance: perf, unmatched: um } = buildPerformance({
|
const speech = demeanorFor(personaCode).speech;
|
||||||
|
setAvatarState("speaking");
|
||||||
|
engine.setState("speaking", nowMs);
|
||||||
|
const handle = startSpeech({
|
||||||
|
engine,
|
||||||
text: speechText,
|
text: speechText,
|
||||||
expression: selectedExpression,
|
expression: selectedExpression,
|
||||||
intensity,
|
intensity,
|
||||||
openness,
|
openness,
|
||||||
seed,
|
seed,
|
||||||
|
speech,
|
||||||
|
now: clock.now,
|
||||||
|
onFrame: ({ viseme, shape, phraseKind, stressed }) => {
|
||||||
|
speechDisplayRef.current = { viseme, shape };
|
||||||
|
coSpeechDisplayRef.current = { phraseKind, stressed };
|
||||||
|
},
|
||||||
|
onEnd: () => endSpeech(clock.now()),
|
||||||
});
|
});
|
||||||
const speech = demeanorFor(personaCode).speech;
|
speechHandleRef.current = handle;
|
||||||
const timeline = buildSpeechTimeline({ text: speechText, syllablesPerSec: speech.syllablesPerSec });
|
setParsedCues(handle.cues);
|
||||||
const coSpeechPlan = buildCoSpeechPlan(timeline, speech, seed);
|
setUnmatched(handle.unmatched);
|
||||||
const speechStartMs = nowMs + 300;
|
|
||||||
setAvatarState("speaking");
|
|
||||||
engine.setState("speaking", nowMs);
|
|
||||||
engine.playPerformance(perf, { speechStartMs, speechDurationMs: timeline.totalDurationMs }, nowMs);
|
|
||||||
setParsedCues(perf.cues);
|
|
||||||
setUnmatched(um);
|
|
||||||
runSpeechTimeline(timeline, speechStartMs, clock.now, speech.articulation, coSpeechPlan);
|
|
||||||
}
|
}
|
||||||
|
|
||||||
function handleAudioFileChange(e: ChangeEvent<HTMLInputElement>): void {
|
function handleAudioFileChange(e: ChangeEvent<HTMLInputElement>): void {
|
||||||
setAudioFileName(e.target.files?.[0]?.name ?? null);
|
setAudioFileName(e.target.files?.[0]?.name ?? null);
|
||||||
}
|
}
|
||||||
|
|
||||||
/** "오디오 파일로 말하기" — 로컬 오디오를 디코드해 포락선을 만들고, 오디오 시계로 표본한다. */
|
/** "오디오 파일로 말하기" — 로컬 오디오를 디코드해 speechDriver에 넘긴다(포락선 선분석·
|
||||||
|
오디오 시계 표본은 startSpeech 안에서 한다). */
|
||||||
async function playSpeechWithAudioFile(): Promise<void> {
|
async function playSpeechWithAudioFile(): Promise<void> {
|
||||||
const file = audioFileInputRef.current?.files?.[0];
|
const file = audioFileInputRef.current?.files?.[0];
|
||||||
if (!file) return;
|
if (!file) return;
|
||||||
|
|
@ -330,21 +273,6 @@ export default function AvatarLab() {
|
||||||
|
|
||||||
const arrayBuffer = await file.arrayBuffer();
|
const arrayBuffer = await file.arrayBuffer();
|
||||||
const audioBuffer = await ctx.decodeAudioData(arrayBuffer.slice(0));
|
const audioBuffer = await ctx.decodeAudioData(arrayBuffer.slice(0));
|
||||||
const channels: Float32Array[] = [];
|
|
||||||
for (let c = 0; c < audioBuffer.numberOfChannels; c++) channels.push(audioBuffer.getChannelData(c));
|
|
||||||
const envelope = computeEnvelope(channels, audioBuffer.sampleRate);
|
|
||||||
|
|
||||||
const nowMs = clock.now();
|
|
||||||
const { performance: perf, unmatched: um } = buildPerformance({
|
|
||||||
text: speechText,
|
|
||||||
expression: selectedExpression,
|
|
||||||
intensity,
|
|
||||||
openness,
|
|
||||||
seed,
|
|
||||||
});
|
|
||||||
const speech = demeanorFor(personaCode).speech;
|
|
||||||
const timeline = buildSpeechTimeline({ text: speechText, syllablesPerSec: speech.syllablesPerSec, envelope });
|
|
||||||
const coSpeechPlan = buildCoSpeechPlan(timeline, speech, seed);
|
|
||||||
|
|
||||||
try {
|
try {
|
||||||
audioSourceRef.current?.stop();
|
audioSourceRef.current?.stop();
|
||||||
|
|
@ -356,19 +284,33 @@ export default function AvatarLab() {
|
||||||
source.connect(ctx.destination);
|
source.connect(ctx.destination);
|
||||||
audioSourceRef.current = source;
|
audioSourceRef.current = source;
|
||||||
|
|
||||||
const speechStartMs = nowMs + 300;
|
|
||||||
setAvatarState("speaking");
|
|
||||||
engine.setState("speaking", nowMs);
|
|
||||||
engine.playPerformance(perf, { speechStartMs, speechDurationMs: timeline.totalDurationMs }, nowMs);
|
|
||||||
setParsedCues(perf.cues);
|
|
||||||
setUnmatched(um);
|
|
||||||
|
|
||||||
const startAtCtx = ctx.currentTime + 0.3;
|
const startAtCtx = ctx.currentTime + 0.3;
|
||||||
source.start(startAtCtx);
|
source.start(startAtCtx);
|
||||||
/* 오디오 시계 기준(초 → ms). Lab의 clock(performance.now() 기반)과는 별개 시계다 —
|
|
||||||
재생 시작을 같은 300ms로 맞췄지만 독립 시계라 아주 긴 발화에서는 드리프트가 있을 수 있다. */
|
speechHandleRef.current?.stop();
|
||||||
const nowFn = () => (ctx.currentTime - startAtCtx) * 1000 + speechStartMs;
|
const nowMs = clock.now();
|
||||||
runSpeechTimeline(timeline, speechStartMs, nowFn, speech.articulation, coSpeechPlan, envelope);
|
const speech = demeanorFor(personaCode).speech;
|
||||||
|
setAvatarState("speaking");
|
||||||
|
engine.setState("speaking", nowMs);
|
||||||
|
const handle = startSpeech({
|
||||||
|
engine,
|
||||||
|
text: speechText,
|
||||||
|
expression: selectedExpression,
|
||||||
|
intensity,
|
||||||
|
openness,
|
||||||
|
seed,
|
||||||
|
speech,
|
||||||
|
now: clock.now,
|
||||||
|
audio: { buffer: audioBuffer, context: ctx, startAt: startAtCtx },
|
||||||
|
onFrame: ({ viseme, shape, phraseKind, stressed }) => {
|
||||||
|
speechDisplayRef.current = { viseme, shape };
|
||||||
|
coSpeechDisplayRef.current = { phraseKind, stressed };
|
||||||
|
},
|
||||||
|
onEnd: () => endSpeech(clock.now()),
|
||||||
|
});
|
||||||
|
speechHandleRef.current = handle;
|
||||||
|
setParsedCues(handle.cues);
|
||||||
|
setUnmatched(handle.unmatched);
|
||||||
}
|
}
|
||||||
|
|
||||||
function playLeakTest(): void {
|
function playLeakTest(): void {
|
||||||
|
|
|
||||||
|
|
@ -23,7 +23,8 @@ import {
|
||||||
ClientAvatar,
|
ClientAvatar,
|
||||||
expressionLabelFor,
|
expressionLabelFor,
|
||||||
} from "../components/avatar/ClientAvatar";
|
} from "../components/avatar/ClientAvatar";
|
||||||
import type { AvatarState, AvatarAffect } from "../components/avatar/ClientAvatar";
|
import type { AvatarAffect, AvatarSpeech, AvatarState } from "../components/avatar/ClientAvatar";
|
||||||
|
import { surfaceIntensityFor } from "../components/avatar/engine/performance";
|
||||||
import { Kicker, Button, Icon, surfaceClassName } from "../components/ui";
|
import { Kicker, Button, Icon, surfaceClassName } from "../components/ui";
|
||||||
import { InnerReactionCard } from "../components/inner-reaction/InnerReactionCard";
|
import { InnerReactionCard } from "../components/inner-reaction/InnerReactionCard";
|
||||||
import {
|
import {
|
||||||
|
|
@ -268,6 +269,9 @@ export default function Session() {
|
||||||
const [voiceConsentError, setVoiceConsentError] = useState<string | null>(null);
|
const [voiceConsentError, setVoiceConsentError] = useState<string | null>(null);
|
||||||
const [resumedSessionLoaded, setResumedSessionLoaded] = useState(false);
|
const [resumedSessionLoaded, setResumedSessionLoaded] = useState(false);
|
||||||
const [voiceAnalyser, setVoiceAnalyser] = useState<AnalyserNode | null>(null);
|
const [voiceAnalyser, setVoiceAnalyser] = useState<AnalyserNode | null>(null);
|
||||||
|
// v3(리노컷) 아바타 발화 구동(결정문 §8.5). TTS 재생 시작 시점에 세팅하고, 정지·실패·
|
||||||
|
// 종료 경로에서 null로 되돌린다.
|
||||||
|
const [avatarSpeech, setAvatarSpeech] = useState<AvatarSpeech | null>(null);
|
||||||
const [endDialogOpen, setEndDialogOpen] = useState(false);
|
const [endDialogOpen, setEndDialogOpen] = useState(false);
|
||||||
const [ending, setEnding] = useState(false);
|
const [ending, setEnding] = useState(false);
|
||||||
const endCancelRef = useRef<HTMLButtonElement>(null);
|
const endCancelRef = useRef<HTMLButtonElement>(null);
|
||||||
|
|
@ -346,6 +350,9 @@ export default function Session() {
|
||||||
const ttsPlaybackActiveRef = useRef(false);
|
const ttsPlaybackActiveRef = useRef(false);
|
||||||
const ttsRequestAbortRef = useRef<AbortController | null>(null);
|
const ttsRequestAbortRef = useRef<AbortController | null>(null);
|
||||||
const playTtsAudioRef = useRef<(() => Promise<void>) | null>(null);
|
const playTtsAudioRef = useRef<(() => Promise<void>) | null>(null);
|
||||||
|
// v3 아바타 발화 텍스트(결정문 §8.5 2g): 텍스트 모드는 speakTextClientTurn 직전,
|
||||||
|
// 음성 모드는 reply 이벤트에서 담는다. 발화를 한 번 내보내면 비운다.
|
||||||
|
const pendingSpeechTextRef = useRef<string | null>(null);
|
||||||
const pendingVoiceLearnerIdRef = useRef<number | null>(null);
|
const pendingVoiceLearnerIdRef = useRef<number | null>(null);
|
||||||
const pendingVoiceLearnerTextRef = useRef<string>("");
|
const pendingVoiceLearnerTextRef = useRef<string>("");
|
||||||
const coachEvidenceCloseRef = useRef<HTMLButtonElement>(null);
|
const coachEvidenceCloseRef = useRef<HTMLButtonElement>(null);
|
||||||
|
|
@ -419,6 +426,7 @@ export default function Session() {
|
||||||
ttsPlaybackCleanupRef.current = null;
|
ttsPlaybackCleanupRef.current = null;
|
||||||
if (cleanup) cleanup();
|
if (cleanup) cleanup();
|
||||||
setVoiceAnalyser(null);
|
setVoiceAnalyser(null);
|
||||||
|
setAvatarSpeech(null);
|
||||||
}, []);
|
}, []);
|
||||||
|
|
||||||
useEffect(() => {
|
useEffect(() => {
|
||||||
|
|
@ -1284,6 +1292,15 @@ export default function Session() {
|
||||||
}
|
}
|
||||||
}, [acceptConsent, consentChecked, pushSignal]);
|
}, [acceptConsent, consentChecked, pushSignal]);
|
||||||
|
|
||||||
|
// TTS를 받지 못했거나 재생하지 못한 경로(결정문 §8.5 2g): 발화 텍스트 ref에 남은
|
||||||
|
// 텍스트가 있으면 텍스트 타이밍 발화로 v3 아바타를 말하게 한다. 내보내면 ref를 비운다.
|
||||||
|
const speakPendingTextFallback = useCallback(() => {
|
||||||
|
const text = pendingSpeechTextRef.current;
|
||||||
|
pendingSpeechTextRef.current = null;
|
||||||
|
if (!text) return;
|
||||||
|
setAvatarSpeech({ id: randomUuid(), text });
|
||||||
|
}, []);
|
||||||
|
|
||||||
const speakTextClientTurn = useCallback(
|
const speakTextClientTurn = useCallback(
|
||||||
async (sessionId: string, turnSeq: number) => {
|
async (sessionId: string, turnSeq: number) => {
|
||||||
const requestId = ttsPlaybackRequestRef.current + 1;
|
const requestId = ttsPlaybackRequestRef.current + 1;
|
||||||
|
|
@ -1311,13 +1328,14 @@ export default function Session() {
|
||||||
setVoiceStatus("degraded");
|
setVoiceStatus("degraded");
|
||||||
setVoiceDetail("내담자 음성을 재생하지 못했습니다. 자막 응답은 화면에 남겼습니다.");
|
setVoiceDetail("내담자 음성을 재생하지 못했습니다. 자막 응답은 화면에 남겼습니다.");
|
||||||
pushSignal("warn", "AI 음성 재생 실패");
|
pushSignal("warn", "AI 음성 재생 실패");
|
||||||
|
speakPendingTextFallback();
|
||||||
} finally {
|
} finally {
|
||||||
if (ttsRequestAbortRef.current === abortController) {
|
if (ttsRequestAbortRef.current === abortController) {
|
||||||
ttsRequestAbortRef.current = null;
|
ttsRequestAbortRef.current = null;
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
},
|
},
|
||||||
[pushSignal],
|
[pushSignal, speakPendingTextFallback],
|
||||||
);
|
);
|
||||||
|
|
||||||
const appendServerClientReply = useCallback(
|
const appendServerClientReply = useCallback(
|
||||||
|
|
@ -1477,6 +1495,7 @@ export default function Session() {
|
||||||
setAvatarState("listening");
|
setAvatarState("listening");
|
||||||
if (!conversationStopped && !qualityRetryable) {
|
if (!conversationStopped && !qualityRetryable) {
|
||||||
if (clientReply && typeof done.turn_seq === "number") {
|
if (clientReply && typeof done.turn_seq === "number") {
|
||||||
|
pendingSpeechTextRef.current = clientReply;
|
||||||
void speakTextClientTurn(liveSessionId, done.turn_seq);
|
void speakTextClientTurn(liveSessionId, done.turn_seq);
|
||||||
}
|
}
|
||||||
void requestLiveCoach({
|
void requestLiveCoach({
|
||||||
|
|
@ -1613,12 +1632,15 @@ export default function Session() {
|
||||||
ttsPlaybackActiveRef.current = false;
|
ttsPlaybackActiveRef.current = false;
|
||||||
ttsPlaybackCleanupRef.current = null;
|
ttsPlaybackCleanupRef.current = null;
|
||||||
setVoiceAnalyser(null);
|
setVoiceAnalyser(null);
|
||||||
|
setAvatarSpeech(null);
|
||||||
setAvatarState("listening");
|
setAvatarState("listening");
|
||||||
setVoiceStatus("idle");
|
setVoiceStatus("idle");
|
||||||
setVoiceDetail("응답이 끝났습니다. 마이크를 다시 켜 발화하세요.");
|
setVoiceDetail("응답이 끝났습니다. 마이크를 다시 켜 발화하세요.");
|
||||||
closeVoiceSocket();
|
closeVoiceSocket();
|
||||||
};
|
};
|
||||||
|
|
||||||
|
// TTS를 받았지만 재생하지 못한 경로(결정문 §8.5 2g): 발화 텍스트가 있으면
|
||||||
|
// 텍스트 타이밍 발화로 v3 아바타를 대신 말하게 한다.
|
||||||
const failPlayback = (detail: string, status: VoiceStatus = "error") => {
|
const failPlayback = (detail: string, status: VoiceStatus = "error") => {
|
||||||
ttsPlaybackActiveRef.current = false;
|
ttsPlaybackActiveRef.current = false;
|
||||||
ttsPlaybackCleanupRef.current = null;
|
ttsPlaybackCleanupRef.current = null;
|
||||||
|
|
@ -1627,6 +1649,7 @@ export default function Session() {
|
||||||
setVoiceStatus(status);
|
setVoiceStatus(status);
|
||||||
setVoiceDetail(detail);
|
setVoiceDetail(detail);
|
||||||
closeVoiceSocket();
|
closeVoiceSocket();
|
||||||
|
speakPendingTextFallback();
|
||||||
};
|
};
|
||||||
|
|
||||||
const ctx = ensureVoiceAudioContext();
|
const ctx = ensureVoiceAudioContext();
|
||||||
|
|
@ -1677,7 +1700,13 @@ export default function Session() {
|
||||||
setAvatarState("speaking");
|
setAvatarState("speaking");
|
||||||
setVoiceStatus("speaking");
|
setVoiceStatus("speaking");
|
||||||
setVoiceDetail(`${clientName} 음성을 재생 중입니다.`);
|
setVoiceDetail(`${clientName} 음성을 재생 중입니다.`);
|
||||||
source.start();
|
const startAt = ctx.currentTime + 0.08;
|
||||||
|
source.start(startAt);
|
||||||
|
const speechText = pendingSpeechTextRef.current;
|
||||||
|
pendingSpeechTextRef.current = null;
|
||||||
|
if (speechText) {
|
||||||
|
setAvatarSpeech({ id: randomUuid(), text: speechText, audio: { buffer: decoded, context: ctx, startAt } });
|
||||||
|
}
|
||||||
return;
|
return;
|
||||||
} catch {
|
} catch {
|
||||||
setVoiceAnalyser(null);
|
setVoiceAnalyser(null);
|
||||||
|
|
@ -1711,11 +1740,14 @@ export default function Session() {
|
||||||
|
|
||||||
try {
|
try {
|
||||||
await audio.play();
|
await audio.play();
|
||||||
|
const speechText = pendingSpeechTextRef.current;
|
||||||
|
pendingSpeechTextRef.current = null;
|
||||||
|
if (speechText) setAvatarSpeech({ id: randomUuid(), text: speechText });
|
||||||
} catch {
|
} catch {
|
||||||
cleanupElement();
|
cleanupElement();
|
||||||
failPlayback("브라우저가 자동 재생을 막았습니다. 자막 응답은 화면에 남겼습니다.", "degraded");
|
failPlayback("브라우저가 자동 재생을 막았습니다. 자막 응답은 화면에 남겼습니다.", "degraded");
|
||||||
}
|
}
|
||||||
}, [clientName, closeVoiceSocket, ensureVoiceAudioContext, stopTtsPlayback]);
|
}, [clientName, closeVoiceSocket, ensureVoiceAudioContext, speakPendingTextFallback, stopTtsPlayback]);
|
||||||
|
|
||||||
useEffect(() => {
|
useEffect(() => {
|
||||||
playTtsAudioRef.current = playTtsAudio;
|
playTtsAudioRef.current = playTtsAudio;
|
||||||
|
|
@ -1947,6 +1979,8 @@ export default function Session() {
|
||||||
|
|
||||||
if (payload.type === "reply") {
|
if (payload.type === "reply") {
|
||||||
replyReceived = true;
|
replyReceived = true;
|
||||||
|
// 음성 모드 발화 텍스트(결정문 §8.5 2g): reply 이벤트의 payload.text를 담는다.
|
||||||
|
pendingSpeechTextRef.current = payload.text ?? null;
|
||||||
setClientReplyPending(false);
|
setClientReplyPending(false);
|
||||||
if (payload.stage) setStage(payload.stage);
|
if (payload.stage) setStage(payload.stage);
|
||||||
if (typeof payload.effective_openness === "number") {
|
if (typeof payload.effective_openness === "number") {
|
||||||
|
|
@ -3562,6 +3596,14 @@ export default function Session() {
|
||||||
affect={avatarAffect}
|
affect={avatarAffect}
|
||||||
analyser={voiceAnalyser}
|
analyser={voiceAnalyser}
|
||||||
rapport={meters.rapport}
|
rapport={meters.rapport}
|
||||||
|
openness={openness}
|
||||||
|
surfaceIntensity={surfaceIntensityFor({
|
||||||
|
expression: avatarAffect,
|
||||||
|
openness,
|
||||||
|
safety: Boolean(safety),
|
||||||
|
paused,
|
||||||
|
})}
|
||||||
|
speech={avatarSpeech}
|
||||||
size={220}
|
size={220}
|
||||||
/>
|
/>
|
||||||
</div>
|
</div>
|
||||||
|
|
|
||||||
2
apps/web/src/vite-env.d.ts
vendored
2
apps/web/src/vite-env.d.ts
vendored
|
|
@ -3,6 +3,8 @@
|
||||||
interface ImportMetaEnv {
|
interface ImportMetaEnv {
|
||||||
/** API 베이스 URL. 기본 "/api" (vite proxy / nginx 가 백엔드로 라우팅). */
|
/** API 베이스 URL. 기본 "/api" (vite proxy / nginx 가 백엔드로 라우팅). */
|
||||||
readonly VITE_API_BASE?: string;
|
readonly VITE_API_BASE?: string;
|
||||||
|
/** "0"이면 아바타 v3(리노컷)를 끄고 항상 기존 SVG를 쓴다. 기본(미설정)은 켜짐. */
|
||||||
|
readonly VITE_AVATAR_V3?: string;
|
||||||
}
|
}
|
||||||
|
|
||||||
interface ImportMeta {
|
interface ImportMeta {
|
||||||
|
|
|
||||||
Loading…
Add table
Add a link
Reference in a new issue