diff --git a/docs/HANDOFF.md b/docs/HANDOFF.md index 32a05f1..3daf8bf 100644 --- a/docs/HANDOFF.md +++ b/docs/HANDOFF.md @@ -456,7 +456,13 @@ G7 내부 소스 계약은 완료됐고 `scripts/check-g7-external-proof.py`는 - human voice-gain pack은 category와 categorical κ를 필수로 포함하고, preregistration이 held-out 공개보다 앞서야 하며, 양 조건이 모두 관측된 50회기/150축만 paired 표본으로 집계한다. 한쪽 결측이면 양쪽 모두 최대오류 ITT로 처리한다. `scripts/check-g7-human-voice-gain.py --input `을 먼저 통과하지 못하면 production runner는 마이크를 - 열기 전에 exit 2로 끝난다. 스키마는 `--print-schema`로 출력한다. 관련 API+script 계약은 97/97 통과했다. + 열기 전에 exit 2로 끝난다. 스키마는 `--print-schema`로 출력한다. +- 실제 평가팀 intake는 `scripts/prepare-g7-human-voice-gain-intake.py --create-template `로 시작한다. 생성물은 + header-only이고 `template_only=true`라 증거가 아니다. provenance와 비식별 participant/labeler/observation CSV를 + 채운 뒤 `--compile --out `을 실행하면 외부 report의 ICC·κ와 행 재계산값의 일치를 확인하고 + production gate가 전부 통과할 때만 기존 파일을 덮어쓰지 않고 최종 pack을 만든다. 콘솔은 경로·키·라벨·잘못된 + 셀 값을 반사하지 않는다. + evaluator/checker 97건과 intake 6건을 합쳐 관련 계약 103/103을 통과했다. 실제 사람 데이터는 아직 없다. - 무마이크 rehearsal 산출물은 `D:\workspace\vignette-runtime-evidence\g7-rehearsal-b34623f3db05-20260809T134105Z`에 있다. voice `61af2e98…009c`, runtime `c6e8670e…4454`, topology `89f1cb86…a251`이며 세 leg 모두 passed, diff --git a/docs/TODO.md b/docs/TODO.md index 899838e..17832d0 100644 --- a/docs/TODO.md +++ b/docs/TODO.md @@ -113,6 +113,10 @@ runner는 이 pack의 category·κ·preregistration·paired completeness·결측 ITT를 3,120초 캡처 전에 검증하며, 부적합하면 마이크를 열지 않고 exit 2로 끝낸다. 독립 검수는 `python -X utf8 -B scripts/check-g7-human-voice-gain.py --input `이고 `--print-schema`로 계약을 확인한다. + 독립 평가팀용 `prepare-g7-human-voice-gain-intake.py`는 빈 header-only 템플릿을 만들고, 비식별 CSV를 + production floor·ICC·κ·gain·CI로 재계산해 전부 통과한 경우에만 최종 pack을 원자 생성한다. 템플릿은 + `template_only=true`라 증거로 컴파일할 수 없고 이름·이메일·원음·축어록·자유서술 열을 받지 않는다. + 도구 준비만 끝난 상태이며 실제 독립 평가자 데이터는 아직 없으므로 G7 external GATE를 유지한다. **공개 선행 조건 완료(2026-08-09):** detached-clean `b34623f3…`·tree `32cc85c7…`에서 OpenAPI 126, `/admin/voice-runtime`, exact local voice provider/model과 queue 4를 제공한다. fresh launcher receipt `9f8d1941…a21bf`가 passed이고 두 Scheduled Task도 동일 source에 pin돼 result 0이다. Cloudflare의 Python @@ -123,7 +127,7 @@ `≥3000s`를 checker가 강제한다. Windows artifact는 detached-clean HEAD/tree, runner/collector/checker SHA, exact `psutil==6.1.1`을 매 sample 전후 pin한다. fresh launcher는 legacy API와 exact-config cloudflared를 bounded 교체하고 새 PID/start/exe/command SHA/실제 cwd의 raw command-line 없는 receipt를 만든다. G7 runner/checker/topology - human pack 계약 97/97과 기존 runtime sampler 4/4를 통과했다. `--rehearse`는 인증 WSS ready/ping/close 1000, + human pack evaluator/checker+intake 계약 103/103과 기존 runtime sampler 4/4를 통과했다. `--rehearse`는 인증 WSS ready/ping/close 1000, runtime 7 samples, Windows topology 7 samples를 모두 통과했고 마이크 capture false·UUID/email literal 0이다. fail-closed 경계(동의 없음·pack 없음·3,120초 미만·host/Origin/scheme 불일치)는 CLI로 실증했다. 상세: `ops/outcome-os-g7-external-proof-readiness-2026-08-07.md`. diff --git a/docs/dev_dashboard.html b/docs/dev_dashboard.html index e19517e..52cc004 100644 --- a/docs/dev_dashboard.html +++ b/docs/dev_dashboard.html @@ -674,7 +674,7 @@

195차 적용(2026-08-10): clean HEAD 13d1c39a…·tree beb93801…의 4차 수동 full run은 앞선 SSE hang을 제거해 6.3분에 종료했고, 실제 동일 학습자 SSE·review ready·G4 첫 POST와 원장 count 0→1까지 통과했다. 이후 full page reload에서 이미 저장된 현재 회기 episode를 카드가 hydrate하지 않아 원장 반영됨 대신 학습자만 실행으로 되돌아가는 제품 정직성 결함을 fail-closed했다. receipt periodic-learner-e2e-20260810-003150.json은 공개·NAS·active DB 접촉 0과 container/volume/network/listener/temp 잔여 0을 증명한다. UI는 현재 session+prescription의 persisted episode를 read model에서 복원해 현재 원장 상태와 재확인 행동을 표시하도록 수정했고, 실제 reload를 포함한 자기주도 route fixture desktop/mobile 6/6·typecheck를 통과했다. full runner는 재시도하지 않았으므로 첫 GREEN receipt와 자동 주기 승격은 여전히 미완료다.

196차 적용(2026-08-10): clean HEAD 9a7dd98c…·tree 1dce20c5…의 5차 단일 full run은 실제 동일 학습자 브라우저 단계에서 SSE·review ready·G4 0→1→1·G5 0→1→1을 78초에 모두 통과해 핵심 폐루프 첫 GREEN을 만들었다. 다만 같은 fixture를 이어서 사용한 returned-practice desktop/mobile 회귀가 이미 적재된 G4/G5를 초기 0으로 오인해 2/4 실패했고 전체 receipt periodic-learner-e2e-20260810-010038.json은 정직하게 FAILED다. 공개·NAS·active DB 접촉 0, container/volume/network/listener/temp 잔여 0은 유지했다. runner는 핵심 폐루프와 후속 회귀 전에 각각 새 학습자·회기 fixture를 준비하도록 분리했고 unit 12/12·Ruff·compile을 통과했다. 실패 화면에서 리뷰 요약 literal [NAME]/[ORG]익명 내담자가고 조사 결함도 발견해 표시 전용 자연어 치환과 복합 조사 우선순위를 보강했고 session-review desktop/mobile 2/2·typecheck·build를 통과했다. 당시에는 전체 GREEN receipt가 없어 자동 full runner를 등록하지 않았다.

197차 적용(2026-08-10): clean HEAD 4cd35488…·tree 001a7c74…의 6차 full run은 핵심 폐루프를 두 번째로 통과했고, 별도 returned-practice fixture도 초기 G4/G5 0을 증명했다. desktop G4/G5 2건은 GREEN이었지만 desktop이 만든 성공 상태를 mobile G4가 다시 읽을 때 테스트가 반영 상태 다시 확인 대신 최초 버튼을 찾아 5분 timeout, mobile G5는 serial 중단됐다. receipt periodic-learner-e2e-20260810-012909.json은 FAILED이며 public/NAS/active DB 접촉 0·exact cleanup GREEN이다. mobile 성공-state locator를 교정했고 실패 화면에서 추가로 확인한 강점·개선점·다음 문장·내담자 반응·워크시트 근거의 literal [NAME]/[ORG]도 저장/API는 유지한 채 표시층에서 자연어화했다. focused session-review desktop/mobile 2/2·typecheck·live spec 4건 collection을 통과했지만 full runner는 재실행하지 않았다.

-

198차 적용(2026-08-10): clean HEAD aa81af29…·tree 3901d71c…의 7차 full run이 499.8초에 첫 전체 GREEN receipt periodic-learner-e2e-20260810-020117.json·SHA-256 bae963de…f32e1a를 만들었다. 실제 브라우저 same-learner session 생성·SSE·review ready·G4/G5 0→1→1, 별도 zero-state returned-practice desktop/mobile 4/4, route mock 0을 통과했다. public 8001·active DB 55432·engine 9099·NAS 접촉은 0이며 전용 container 4·volume 3·network 1·engine PID와 6578-6581 listener는 cleanup 뒤 모두 0이다. 기존 vignette-e2e hourly heartbeat는 중복 생성 없이 갱신해 최신 GREEN이 6시간 이상 오래됐거나 material milestone이 바뀐 때만 preflight 뒤 full runner를 최대 1회 실행하고, 실패 시 같은 heartbeat에서 재시도하지 않는다. G7 human pack validator·runner/checker current 계약도 97/97로 재검증했으며 물리 마이크는 열지 않았다.

+

198차 적용(2026-08-10): clean HEAD aa81af29…·tree 3901d71c…의 7차 full run이 499.8초에 첫 전체 GREEN receipt periodic-learner-e2e-20260810-020117.json·SHA-256 bae963de…f32e1a를 만들었다. 실제 브라우저 same-learner session 생성·SSE·review ready·G4/G5 0→1→1, 별도 zero-state returned-practice desktop/mobile 4/4, route mock 0을 통과했다. public 8001·active DB 55432·engine 9099·NAS 접촉은 0이며 전용 container 4·volume 3·network 1·engine PID와 6578-6581 listener는 cleanup 뒤 모두 0이다. 기존 vignette-e2e hourly heartbeat는 중복 생성 없이 갱신해 최신 GREEN이 6시간 이상 오래됐거나 material milestone이 바뀐 때만 preflight 뒤 full runner를 최대 1회 실행하고, 실패 시 같은 heartbeat에서 재시도하지 않는다. G7 human pack evaluator·runner/checker와 후속 intake compiler 계약도 103/103으로 재검증했으며 물리 마이크는 열지 않았다.

래스터만 사용이미지 생성 도구 산출물은 PNG 기반 시안이다. SVG·벡터·와이어프레임·로고 시트로 해석하지 않는다.
기능 우선메인 라우트의 실제 액션과 정보 구조를 먼저 반영한다. 장식은 기능을 가리지 않는 수준에서만 쓴다.
@@ -929,7 +929,7 @@
-
내부 구현·공개 선행조건 DONE

consent→단일 clock→독립 text/voice/fusion→철회·tombstone, streaming interim/final·word timestamp·provider event, HMAC word pseudonym, 열린 stream의 1초 동의 재검사·abort, 텍스트 보존·음성 재연결 UX를 연결했다. 운영 기본 provider는 노트북 상주 local_whisper/melotts이고 외부 Deepgram/OpenAI adapter는 fallback으로 보존한다. 격리 synthetic PCM soak는 59/59를 통과했다. Windows host topology는 detached-clean commit/tree·도구 SHA·psutil version과 PID/start/exe·command SHA/cwd·process/TCP high-water를 fail-closed로 수집하고, 3,120초 capture의 공통 3,000초를 canonical exit와 결속한다. 공개 API/cloudflared는 clean b34623f3…·tree 32cc85c7…로 fresh 승격돼 OpenAPI 126·/admin/voice-runtime·local_whisper/small·melotts/melotts-korean·WS queue 4를 제공한다. receipt 9f8d1941…a21bf가 passed이고 두 source-pinned task도 result 0이다. human pack은 category·κ·preregistration·paired completeness·결측 ITT를 캡처 전에 검증하며 관련 계약 97/97·runtime sampler 4/4를 통과했다.

외부 종료 GATE

30초 무마이크 rehearsal은 authenticated public WSS ready/ping/close1000, runtime 7 samples, Windows topology 7 samples를 모두 통과했고 physical_capture=false·UUID/email literal 0이다. 남은 것은 사용자가 선택한 장치와 실행 직전 명시 동의를 받은 물리 마이크 3,120초 양방향 soak, 같은 public host·공통 3,000초 시간창의 worker/Uvicorn queue와 process/TCP high-water, 최소 31명·held-out 50회기·150축·blind evaluator 2인의 독립 human voice-gain pack이다. pack은 check-g7-human-voice-gain.py를 먼저 통과해야 하며 부적합하면 마이크를 열기 전에 exit 2다. 네 artifact가 scripts/check-g7-external-proof.py exit 0을 만들기 전에는 메인 상태를 DONE으로 바꾸지 않는다.

+
내부 구현·공개 선행조건 DONE

consent→단일 clock→독립 text/voice/fusion→철회·tombstone, streaming interim/final·word timestamp·provider event, HMAC word pseudonym, 열린 stream의 1초 동의 재검사·abort, 텍스트 보존·음성 재연결 UX를 연결했다. 운영 기본 provider는 노트북 상주 local_whisper/melotts이고 외부 Deepgram/OpenAI adapter는 fallback으로 보존한다. 격리 synthetic PCM soak는 59/59를 통과했다. Windows host topology는 detached-clean commit/tree·도구 SHA·psutil version과 PID/start/exe·command SHA/cwd·process/TCP high-water를 fail-closed로 수집하고, 3,120초 capture의 공통 3,000초를 canonical exit와 결속한다. 공개 API/cloudflared는 clean b34623f3…·tree 32cc85c7…로 fresh 승격돼 OpenAPI 126·/admin/voice-runtime·local_whisper/small·melotts/melotts-korean·WS queue 4를 제공한다. receipt 9f8d1941…a21bf가 passed이고 두 source-pinned task도 result 0이다. human pack은 category·κ·preregistration·paired completeness·결측 ITT를 캡처 전에 검증한다. 새 intake compiler는 non-evidence 빈 틀에서 시작해 비식별 CSV의 production floor·ICC·κ·gain·CI를 재계산하고 모두 통과한 경우에만 pack을 쓰며 관련 계약은 103/103·runtime sampler 4/4다.

외부 종료 GATE

30초 무마이크 rehearsal은 authenticated public WSS ready/ping/close1000, runtime 7 samples, Windows topology 7 samples를 모두 통과했고 physical_capture=false·UUID/email literal 0이다. 남은 것은 사용자가 선택한 장치와 실행 직전 명시 동의를 받은 물리 마이크 3,120초 양방향 soak, 같은 public host·공통 3,000초 시간창의 worker/Uvicorn queue와 process/TCP high-water, 최소 31명·held-out 50회기·150축·blind evaluator 2인의 실제 독립 human voice-gain pack이다. 작성 도구만 준비됐고 사람 데이터는 아직 없다. pack은 check-g7-human-voice-gain.py를 먼저 통과해야 하며 부적합하면 마이크를 열기 전에 exit 2다. 네 artifact가 scripts/check-g7-external-proof.py exit 0을 만들기 전에는 메인 상태를 DONE으로 바꾸지 않는다.

diff --git a/docs/guides/testing.md b/docs/guides/testing.md index 779c252..56c74c1 100644 --- a/docs/guides/testing.md +++ b/docs/guides/testing.md @@ -448,7 +448,9 @@ hourly heartbeat는 최신 GREEN이 6시간 이상 오래됐거나 material mile - 2026-08-09 G7 human pack preflight: production runner는 캡처 전에 category·categorical κ≥0.70, preregistration 선행, 50회기/150축 complete pairing, 양 조건 결측 ITT를 검증한다. 독립 검수는 `python -X utf8 -B scripts/check-g7-human-voice-gain.py --input `이며 `--print-schema`도 지원한다. - 부적합 pack은 마이크를 열기 전에 exit 2다. API evaluator 11/11 + G7 script 86/86, 합계 97/97 통과했다. + 부적합 pack은 마이크를 열기 전에 exit 2다. `prepare-g7-human-voice-gain-intake.py`는 non-evidence 빈 템플릿과 + strict CSV compiler를 제공하며 production gate 통과 전에는 pack을 쓰지 않는다. API evaluator 11/11 + G7 script + 86/86 + intake compiler 6/6, 합계 103/103을 통과했다. 실제 사람 데이터가 없으므로 external GATE는 열려 있다. - 2026-08-07 G7/G8 시각 QA: 데스크톱·Pixel 5에서 focused **16/16**, typecheck·build·cosmetic-filter를 통과했다. G8 승인 차단 이유를 `title` 의존에서 상시 문구·`aria-describedby`·semantic form/Enter 제출로 바꾸고, G7 이벤트 최소 24×24px와 시간축 키보드 포커스를 수치 회귀로 고정했다. 상세는 diff --git a/docs/ops/backlog-2026-06-26.md b/docs/ops/backlog-2026-06-26.md index 92a4de9..ea49028 100644 --- a/docs/ops/backlog-2026-06-26.md +++ b/docs/ops/backlog-2026-06-26.md @@ -71,6 +71,9 @@ 참가자를 포함한 총 최소 31명/held-out 50회기/150축 독립 blind human-held-out voice-gain evaluator와 이를 같은 public host·동시 50분 시간창으로 묶는 canonical checker까지 code/internal 준비를 마쳤다. + 독립 평가 입력은 `prepare-g7-human-voice-gain-intake.py --create-template`의 명시적 non-evidence 빈 틀로 시작하고, + `--compile ... --out ...`가 비식별 CSV의 production floor·ICC·κ·gain·CI를 재계산해 모두 통과할 때만 pack을 + 생성한다. 이 작성 도구는 준비됐지만 실제 참가자·독립 평가자 데이터는 아직 수집되지 않았다. 운영 TTS는 MeloTTS Korean(MIT) loopback으로 전환했고 외부 OpenAI adapter와 회기 전·중 AI 생성 음성 고지는 보존했다. clean public runtime·`local_whisper`/`melotts` exact ready·authenticated WSS의 30초 무마이크 rehearsal은 voice/runtime/topology 각 7 samples로 통과했다. 남은 범위: 실행 직전 장치 선택·명시 동의 물리 마이크와 diff --git a/docs/ops/outcome-os-g7-external-proof-readiness-2026-08-07.md b/docs/ops/outcome-os-g7-external-proof-readiness-2026-08-07.md index 3c93b1f..18a487f 100644 --- a/docs/ops/outcome-os-g7-external-proof-readiness-2026-08-07.md +++ b/docs/ops/outcome-os-g7-external-proof-readiness-2026-08-07.md @@ -79,6 +79,19 @@ runtime/topology interval 기본값은 child collector 상한과 같은 60초이 사람 데이터 운영자는 52분 실행과 독립적으로 pack을 먼저 검증할 수 있다. `--print-schema`는 authoritative JSON Schema를 출력하고, 오류는 JSON pointer와 유형만 반환해 participant/labeler key·라벨 값·PII를 반사하지 않는다. +복잡한 nested JSON을 사람이 직접 조립하지 않도록 intake 도구를 제공한다. `--create-template`는 header-only CSV와 +`template_only=true` manifest를 만들어 **증거로 사용할 수 없다**. 실제 외부 protocol·consent·split·labeling·analysis +artifact SHA와 비식별 participant/labeler/observation 수치를 채운 뒤 `--compile`하면 ICC·categorical κ와 production +floor·gain·CI를 행에서 다시 계산한다. 모든 gate가 통과한 경우에만 기존 파일을 덮어쓰지 않고 최종 pack을 원자 생성한다. +이름·이메일·원음·축어록·자유서술 열은 받지 않으며 콘솔에도 입력 경로·키·라벨·잘못된 셀 값을 출력하지 않는다. + +```powershell +& $py -X utf8 -B scripts/prepare-g7-human-voice-gain-intake.py --create-template +# 실제 독립 평가팀이 manifest와 CSV를 채운 뒤에만 실행 +& $py -X utf8 -B scripts/prepare-g7-human-voice-gain-intake.py ` + --compile --out +``` + ```powershell & $py -X utf8 -B scripts/check-g7-human-voice-gain.py --print-schema & $py -X utf8 -B scripts/check-g7-human-voice-gain.py --input diff --git a/scripts/prepare-g7-human-voice-gain-intake.py b/scripts/prepare-g7-human-voice-gain-intake.py new file mode 100644 index 0000000..b6bee4b --- /dev/null +++ b/scripts/prepare-g7-human-voice-gain-intake.py @@ -0,0 +1,578 @@ +#!/usr/bin/env python3 +"""Prepare or compile a deidentified G7 human voice-gain intake. + +Template mode creates empty, explicitly non-evidence input files. Compile mode +accepts only deidentified metrics, recomputes reliability and production gates, +and writes a final pack only when every production requirement passes. Console +output never includes input paths, participant/labeler keys, labels, or values +from invalid cells. +""" + +from __future__ import annotations + +import argparse +import csv +import json +import os +import re +import sys +import tempfile +from dataclasses import dataclass +from pathlib import Path +from typing import Any, Iterable + +from pydantic import ValidationError + + +REPO_ROOT = Path(__file__).resolve().parents[1] +API_ROOT = REPO_ROOT / "apps/api" +if str(API_ROOT) not in sys.path: + sys.path.insert(0, str(API_ROOT)) + +from app.contracts.g7_external_evidence import ( # noqa: E402 + G7HumanVoiceGainEvidencePack, +) +from app.services.g7_voice_gain_evidence import ( # noqa: E402 + evaluate_human_voice_gain, +) + + +MANIFEST_NAME = "manifest.json" +PARTICIPANTS_NAME = "participants.csv" +LABELERS_NAME = "labelers.csv" +OBSERVATIONS_NAME = "observations.csv" +README_NAME = "README.md" +INTAKE_SCHEMA_VERSION = "g7_human_voice_gain_intake_v1" + +MANIFEST_KEYS = { + "intake_schema_version", + "template_only", + "provenance", + "text_only_model", + "voice_enabled_model", + "power_plan", + "reported_icc", + "reported_categorical_kappa", + "reliability_report_sha256", +} +PARTICIPANT_FIELDS = ( + "participant_key", + "split", + "consent_receipt_sha256", +) +LABELER_FIELDS = ( + "labeler_key", + "blinded_to_model_condition", + "blinded_to_other_labelers", + "labeled_independently", + "attestation_sha256", +) +OBSERVATION_FIELDS = ( + "observation_id", + "participant_key", + "session_key", + "axis", + "text_only_status", + "text_only_score", + "voice_enabled_status", + "voice_enabled_score", + "labeler_key", + "label_score", + "label_category", +) + +DEIDENTIFIED_PARTICIPANT_RE = re.compile( + r"^(?:calibration|held)-[A-Za-z0-9][A-Za-z0-9._:-]{2,95}$" +) +DEIDENTIFIED_LABELER_RE = re.compile(r"^labeler-[A-Za-z0-9][A-Za-z0-9._:-]{2,95}$") +DEIDENTIFIED_SESSION_RE = re.compile(r"^session-[A-Za-z0-9][A-Za-z0-9._:-]{2,95}$") + + +@dataclass(frozen=True, slots=True) +class IntakeError(Exception): + code: str + source: str = "intake" + row: int | None = None + field: str | None = None + + +def _base_report() -> dict[str, Any]: + return { + "schema_version": "vignette.g7-human-voice-gain-intake.v1", + "passed": False, + "pack_written": False, + "clinical_claim_allowed": False, + "privacy_boundary": { + "input_paths_logged": False, + "participant_keys_logged": False, + "labeler_keys_logged": False, + "labels_logged": False, + "raw_cell_values_logged": False, + "raw_audio_or_transcript_accepted": False, + }, + "errors": [], + "result": {}, + } + + +def _safe_error(error: IntakeError) -> dict[str, Any]: + result: dict[str, Any] = {"code": error.code, "source": error.source} + if error.row is not None: + result["row"] = error.row + if error.field is not None: + result["field"] = error.field + return result + + +def _json_pointer(location: tuple[int | str, ...]) -> str: + if not location: + return "/" + return "/" + "/".join( + str(item).replace("~", "~0").replace("/", "~1") for item in location + ) + + +def _validation_errors(error: ValidationError) -> list[dict[str, Any]]: + return [ + { + "code": f"contract:{item['type']}", + "source": "compiled_pack", + "pointer": _json_pointer(tuple(item["loc"])), + } + for item in error.errors( + include_url=False, + include_context=False, + include_input=False, + ) + ] + + +def _require_exact_keys(payload: dict[str, Any]) -> None: + if set(payload) != MANIFEST_KEYS: + raise IntakeError("manifest_fields_invalid", source=MANIFEST_NAME) + + +def _read_manifest(directory: Path) -> dict[str, Any]: + try: + value = json.loads((directory / MANIFEST_NAME).read_text(encoding="utf-8")) + except (OSError, UnicodeError, json.JSONDecodeError) as exc: + raise IntakeError( + f"manifest_read:{type(exc).__name__}", source=MANIFEST_NAME + ) from None + if not isinstance(value, dict): + raise IntakeError("manifest_object_required", source=MANIFEST_NAME) + _require_exact_keys(value) + if value.get("intake_schema_version") != INTAKE_SCHEMA_VERSION: + raise IntakeError("intake_schema_version_invalid", source=MANIFEST_NAME) + if value.get("template_only") is not False: + raise IntakeError("template_cannot_compile", source=MANIFEST_NAME) + return value + + +def _read_csv( + directory: Path, filename: str, fields: tuple[str, ...] +) -> list[dict[str, str]]: + try: + with (directory / filename).open( + "r", encoding="utf-8-sig", newline="" + ) as handle: + reader = csv.DictReader(handle) + if reader.fieldnames != list(fields): + raise IntakeError("csv_headers_invalid", source=filename) + rows = [] + for row_number, row in enumerate(reader, start=2): + if None in row or any(value is None for value in row.values()): + raise IntakeError( + "csv_shape_invalid", source=filename, row=row_number + ) + if any(len(value) > 256 for value in row.values()): + raise IntakeError( + "csv_cell_too_long", source=filename, row=row_number + ) + rows.append(dict(row)) + except IntakeError: + raise + except (OSError, UnicodeError, csv.Error) as exc: + raise IntakeError(f"csv_read:{type(exc).__name__}", source=filename) from None + if not rows: + raise IntakeError("csv_rows_required", source=filename) + return rows + + +def _parse_true(value: str, *, source: str, row: int, field: str) -> bool: + if value != "true": + raise IntakeError("literal_true_required", source, row, field) + return True + + +def _parse_score( + value: str, + *, + nullable: bool, + source: str, + row: int, + field: str, +) -> float | None: + if value == "" and nullable: + return None + try: + result = float(value) + except ValueError: + raise IntakeError("score_invalid", source, row, field) from None + if not 0.0 <= result <= 1.0: + raise IntakeError("score_out_of_range", source, row, field) + return result + + +def _parse_participants(rows: list[dict[str, str]]) -> list[dict[str, Any]]: + participants = [] + for row_number, row in enumerate(rows, start=2): + key = row["participant_key"] + split = row["split"] + if DEIDENTIFIED_PARTICIPANT_RE.fullmatch(key) is None: + raise IntakeError( + "deidentified_participant_key_required", + PARTICIPANTS_NAME, + row_number, + "participant_key", + ) + expected_prefix = "calibration-" if split == "calibration" else "held-" + if split not in {"calibration", "held_out"} or not key.startswith( + expected_prefix + ): + raise IntakeError( + "participant_split_key_mismatch", + PARTICIPANTS_NAME, + row_number, + "split", + ) + participants.append(row) + return participants + + +def _parse_labelers(rows: list[dict[str, str]]) -> list[dict[str, Any]]: + labelers = [] + for row_number, row in enumerate(rows, start=2): + key = row["labeler_key"] + if DEIDENTIFIED_LABELER_RE.fullmatch(key) is None: + raise IntakeError( + "deidentified_labeler_key_required", + LABELERS_NAME, + row_number, + "labeler_key", + ) + labelers.append( + { + "labeler_key": key, + "blinded_to_model_condition": _parse_true( + row["blinded_to_model_condition"], + source=LABELERS_NAME, + row=row_number, + field="blinded_to_model_condition", + ), + "blinded_to_other_labelers": _parse_true( + row["blinded_to_other_labelers"], + source=LABELERS_NAME, + row=row_number, + field="blinded_to_other_labelers", + ), + "labeled_independently": _parse_true( + row["labeled_independently"], + source=LABELERS_NAME, + row=row_number, + field="labeled_independently", + ), + "attestation_sha256": row["attestation_sha256"], + } + ) + return labelers + + +def _parse_observations(rows: list[dict[str, str]]) -> list[dict[str, Any]]: + grouped: dict[tuple[str, str, str, str], dict[str, Any]] = {} + order: list[tuple[str, str, str, str]] = [] + for row_number, row in enumerate(rows, start=2): + participant_key = row["participant_key"] + session_key = row["session_key"] + labeler_key = row["labeler_key"] + if DEIDENTIFIED_PARTICIPANT_RE.fullmatch(participant_key) is None: + raise IntakeError( + "deidentified_participant_key_required", + OBSERVATIONS_NAME, + row_number, + "participant_key", + ) + if DEIDENTIFIED_SESSION_RE.fullmatch(session_key) is None: + raise IntakeError( + "deidentified_session_key_required", + OBSERVATIONS_NAME, + row_number, + "session_key", + ) + if DEIDENTIFIED_LABELER_RE.fullmatch(labeler_key) is None: + raise IntakeError( + "deidentified_labeler_key_required", + OBSERVATIONS_NAME, + row_number, + "labeler_key", + ) + key = ( + row["observation_id"], + participant_key, + session_key, + row["axis"], + ) + core = { + "observation_id": row["observation_id"], + "participant_key": participant_key, + "session_key": session_key, + "axis": row["axis"], + "text_only_status": row["text_only_status"], + "text_only_score": _parse_score( + row["text_only_score"], + nullable=True, + source=OBSERVATIONS_NAME, + row=row_number, + field="text_only_score", + ), + "voice_enabled_status": row["voice_enabled_status"], + "voice_enabled_score": _parse_score( + row["voice_enabled_score"], + nullable=True, + source=OBSERVATIONS_NAME, + row=row_number, + field="voice_enabled_score", + ), + "raw_audio_included": False, + "transcript_included": False, + } + if key not in grouped: + grouped[key] = {**core, "labels": []} + order.append(key) + elif any(grouped[key][name] != value for name, value in core.items()): + raise IntakeError( + "observation_core_mismatch", OBSERVATIONS_NAME, row_number + ) + grouped[key]["labels"].append( + { + "labeler_key": labeler_key, + "score": _parse_score( + row["label_score"], + nullable=False, + source=OBSERVATIONS_NAME, + row=row_number, + field="label_score", + ), + "category": row["label_category"], + } + ) + return [grouped[key] for key in order] + + +def _draft_pack(directory: Path) -> dict[str, Any]: + manifest = _read_manifest(directory) + participants = _parse_participants( + _read_csv(directory, PARTICIPANTS_NAME, PARTICIPANT_FIELDS) + ) + labelers = _parse_labelers(_read_csv(directory, LABELERS_NAME, LABELER_FIELDS)) + observations = _parse_observations( + _read_csv(directory, OBSERVATIONS_NAME, OBSERVATION_FIELDS) + ) + labeler_keys = [item["labeler_key"] for item in labelers] + return { + "provenance": manifest["provenance"], + "text_only_model": manifest["text_only_model"], + "voice_enabled_model": manifest["voice_enabled_model"], + "power_plan": manifest["power_plan"], + "participants": participants, + "labeler_attestations": labelers, + "reliability": { + "labeler_keys": labeler_keys, + "reported_icc": manifest["reported_icc"], + "reported_categorical_kappa": manifest["reported_categorical_kappa"], + "report_sha256": manifest["reliability_report_sha256"], + }, + "observations": observations, + } + + +def compile_intake(directory: Path) -> tuple[dict[str, Any], dict[str, Any] | None]: + report = _base_report() + try: + payload = _draft_pack(directory) + pack = G7HumanVoiceGainEvidencePack.model_validate(payload) + result = evaluate_human_voice_gain(pack) + except IntakeError as exc: + report["errors"] = [_safe_error(exc)] + return report, None + except ValidationError as exc: + report["errors"] = _validation_errors(exc) + return report, None + except Exception as exc: + report["errors"] = [ + {"code": f"evaluation:{type(exc).__name__}", "source": "compiled_pack"} + ] + return report, None + + report["passed"] = result.passed + report["result"] = result.model_dump(mode="json") + if not result.passed: + return report, None + return report, pack.model_dump(mode="json") + + +def _atomic_create_json(path: Path, payload: dict[str, Any]) -> None: + if path.exists(): + raise IntakeError("output_exists", source="output") + path.parent.mkdir(parents=True, exist_ok=True) + temporary: Path | None = None + try: + with tempfile.NamedTemporaryFile( + "w", + encoding="utf-8", + newline="\n", + dir=path.parent, + prefix=f".{path.name}.", + suffix=".tmp", + delete=False, + ) as handle: + temporary = Path(handle.name) + json.dump(payload, handle, ensure_ascii=False, indent=2, sort_keys=True) + handle.write("\n") + handle.flush() + os.fsync(handle.fileno()) + try: + os.link(temporary, path) + except FileExistsError: + raise IntakeError("output_exists", source="output") from None + except OSError as exc: + raise IntakeError( + f"output_link:{type(exc).__name__}", source="output" + ) from None + try: + temporary.unlink() + except OSError as exc: + path.unlink(missing_ok=True) + raise IntakeError( + f"output_cleanup:{type(exc).__name__}", source="output" + ) from None + temporary = None + finally: + if temporary is not None: + temporary.unlink(missing_ok=True) + + +def create_template(directory: Path) -> None: + if directory.exists(): + raise IntakeError("template_directory_exists", source="template") + directory.mkdir(parents=True) + manifest = { + "intake_schema_version": INTAKE_SCHEMA_VERSION, + "template_only": True, + "provenance": { + "protocol_sha256": "", + "consent_protocol_sha256": "", + "dataset_manifest_sha256": "", + "split_manifest_sha256": "", + "labeling_protocol_sha256": "", + "analysis_plan_sha256": "", + "registered_at": "", + "held_out_labels_opened_at": "", + }, + "text_only_model": { + "role": "text_only_baseline", + "provider": "", + "model_id": "", + "model_version": "", + "artifact_sha256": "", + "configuration_sha256": "", + }, + "voice_enabled_model": { + "role": "voice_enabled_candidate", + "provider": "", + "model_id": "", + "model_version": "", + "artifact_sha256": "", + "configuration_sha256": "", + }, + "power_plan": { + "primary_metric": "paired_one_minus_mae_gain", + "clustering_unit": "participant", + "required_held_out_participants": 30, + "required_held_out_sessions": 50, + "required_paired_axis_observations": 150, + "alpha": 0.05, + "target_power": 0.8, + "minimally_detectable_gain": 0.01, + "planned_bootstrap_samples": 10000, + }, + "reported_icc": None, + "reported_categorical_kappa": None, + "reliability_report_sha256": "", + } + (directory / MANIFEST_NAME).write_text( + json.dumps(manifest, ensure_ascii=False, indent=2, sort_keys=True) + "\n", + encoding="utf-8", + newline="\n", + ) + for filename, fields in ( + (PARTICIPANTS_NAME, PARTICIPANT_FIELDS), + (LABELERS_NAME, LABELER_FIELDS), + (OBSERVATIONS_NAME, OBSERVATION_FIELDS), + ): + with (directory / filename).open("w", encoding="utf-8", newline="") as handle: + csv.writer(handle, lineterminator="\n").writerow(fields) + (directory / README_NAME).write_text( + "# G7 human voice-gain intake\n\n" + "이 디렉터리는 빈 입력 틀이지 증거가 아니다. `manifest.json`의 모든 provenance를 " + "실제 외부 연구 산출물 SHA-256으로 채우고 `template_only`를 `false`로 바꿔라. " + "CSV에는 비식별 키와 수치만 넣고 이름, 이메일, 원음, 축어록, 자유서술을 넣지 마라. " + "각 held-out 회기는 goal/task/bond 3축과 동일한 blind labeler panel을 가져야 한다.\n", + encoding="utf-8", + newline="\n", + ) + + +def parser() -> argparse.ArgumentParser: + result = argparse.ArgumentParser(description=__doc__) + mode = result.add_mutually_exclusive_group(required=True) + mode.add_argument("--create-template", type=Path, metavar="DIR") + mode.add_argument("--compile", type=Path, metavar="DIR") + result.add_argument("--out", type=Path, help="final pack JSON; compile only") + return result + + +def main(argv: Iterable[str] | None = None) -> int: + args = parser().parse_args(list(argv) if argv is not None else None) + report = _base_report() + if args.create_template is not None: + if args.out is not None: + parser().error("--out is compile-only") + try: + create_template(args.create_template) + except IntakeError as exc: + report["errors"] = [_safe_error(exc)] + print(json.dumps(report, ensure_ascii=False, indent=2, sort_keys=True)) + return 1 + report["template_created"] = True + report["template_is_evidence"] = False + print(json.dumps(report, ensure_ascii=False, indent=2, sort_keys=True)) + return 0 + + if args.out is None: + parser().error("--compile requires --out") + report, pack_payload = compile_intake(args.compile) + if pack_payload is not None: + try: + _atomic_create_json(args.out, pack_payload) + except IntakeError as exc: + report["passed"] = False + report["errors"] = [_safe_error(exc)] + else: + report["pack_written"] = True + print(json.dumps(report, ensure_ascii=False, indent=2, sort_keys=True)) + return 0 if report["passed"] is True and report["pack_written"] else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/scripts/test_prepare_g7_human_voice_gain_intake.py b/scripts/test_prepare_g7_human_voice_gain_intake.py new file mode 100644 index 0000000..acada79 --- /dev/null +++ b/scripts/test_prepare_g7_human_voice_gain_intake.py @@ -0,0 +1,270 @@ +from __future__ import annotations + +import contextlib +import csv +import importlib.util +import io +import json +import sys +import tempfile +import unittest +from pathlib import Path + + +REPO_ROOT = Path(__file__).resolve().parents[1] +API_ROOT = REPO_ROOT / "apps/api" +if str(API_ROOT) not in sys.path: + sys.path.insert(0, str(API_ROOT)) + +from app.contracts.g7_external_evidence import ( # noqa: E402 + G7HumanVoiceGainEvidencePack, +) +from app.test_g7_voice_gain_evidence import _valid_payload # noqa: E402 +from scripts.test_g7_external_proof import human_pack # noqa: E402 + + +SCRIPT_PATH = Path(__file__).with_name("prepare-g7-human-voice-gain-intake.py") +SPEC = importlib.util.spec_from_file_location( + "prepare_g7_human_voice_gain_intake", SCRIPT_PATH +) +assert SPEC is not None and SPEC.loader is not None +MODULE = importlib.util.module_from_spec(SPEC) +sys.modules[SPEC.name] = MODULE +SPEC.loader.exec_module(MODULE) + + +def _write_csv( + path: Path, fields: tuple[str, ...], rows: list[dict[str, object]] +) -> None: + with path.open("w", encoding="utf-8", newline="") as handle: + writer = csv.DictWriter(handle, fieldnames=fields, lineterminator="\n") + writer.writeheader() + writer.writerows(rows) + + +def _write_intake(directory: Path, payload: dict[str, object]) -> None: + reliability = payload["reliability"] + assert isinstance(reliability, dict) + manifest = { + "intake_schema_version": MODULE.INTAKE_SCHEMA_VERSION, + "template_only": False, + "provenance": payload["provenance"], + "text_only_model": payload["text_only_model"], + "voice_enabled_model": payload["voice_enabled_model"], + "power_plan": payload["power_plan"], + "reported_icc": reliability["reported_icc"], + "reported_categorical_kappa": reliability["reported_categorical_kappa"], + "reliability_report_sha256": reliability["report_sha256"], + } + (directory / MODULE.MANIFEST_NAME).write_text( + json.dumps(manifest), encoding="utf-8" + ) + + participants = payload["participants"] + assert isinstance(participants, list) + _write_csv( + directory / MODULE.PARTICIPANTS_NAME, + MODULE.PARTICIPANT_FIELDS, + participants, + ) + + attestations = payload["labeler_attestations"] + assert isinstance(attestations, list) + labeler_rows = [] + for item in attestations: + assert isinstance(item, dict) + labeler_rows.append( + { + "labeler_key": item["labeler_key"], + "blinded_to_model_condition": "true", + "blinded_to_other_labelers": "true", + "labeled_independently": "true", + "attestation_sha256": item["attestation_sha256"], + } + ) + _write_csv( + directory / MODULE.LABELERS_NAME, + MODULE.LABELER_FIELDS, + labeler_rows, + ) + + observations = payload["observations"] + assert isinstance(observations, list) + observation_rows = [] + for observation in observations: + assert isinstance(observation, dict) + labels = observation["labels"] + assert isinstance(labels, list) + for label in labels: + assert isinstance(label, dict) + observation_rows.append( + { + "observation_id": observation["observation_id"], + "participant_key": observation["participant_key"], + "session_key": observation["session_key"], + "axis": observation["axis"], + "text_only_status": observation["text_only_status"], + "text_only_score": ( + "" + if observation.get("text_only_score") is None + else observation["text_only_score"] + ), + "voice_enabled_status": observation["voice_enabled_status"], + "voice_enabled_score": ( + "" + if observation.get("voice_enabled_score") is None + else observation["voice_enabled_score"] + ), + "labeler_key": label["labeler_key"], + "label_score": label["score"], + "label_category": label["category"], + } + ) + _write_csv( + directory / MODULE.OBSERVATIONS_NAME, + MODULE.OBSERVATION_FIELDS, + observation_rows, + ) + + +class G7HumanVoiceGainIntakeTests(unittest.TestCase): + def test_empty_template_is_explicitly_not_evidence(self) -> None: + with tempfile.TemporaryDirectory() as root: + target = Path(root) / "intake" + output = io.StringIO() + with contextlib.redirect_stdout(output): + exit_code = MODULE.main(["--create-template", str(target)]) + + self.assertEqual(0, exit_code) + self.assertEqual( + { + MODULE.MANIFEST_NAME, + MODULE.PARTICIPANTS_NAME, + MODULE.LABELERS_NAME, + MODULE.OBSERVATIONS_NAME, + MODULE.README_NAME, + }, + {item.name for item in target.iterdir()}, + ) + manifest = json.loads( + (target / MODULE.MANIFEST_NAME).read_text(encoding="utf-8") + ) + self.assertIs(True, manifest["template_only"]) + self.assertFalse(json.loads(output.getvalue())["template_is_evidence"]) + + report, pack = MODULE.compile_intake(target) + self.assertIsNone(pack) + self.assertEqual("template_cannot_compile", report["errors"][0]["code"]) + + def test_production_intake_compiles_and_checks_reported_reliability(self) -> None: + with tempfile.TemporaryDirectory() as root: + directory = Path(root) / "private-human-intake" + directory.mkdir() + payload = human_pack() + _write_intake(directory, payload) + output_path = Path(root) / "compiled.json" + stdout = io.StringIO() + with contextlib.redirect_stdout(stdout): + exit_code = MODULE.main( + ["--compile", str(directory), "--out", str(output_path)] + ) + + self.assertEqual(0, exit_code) + report = json.loads(stdout.getvalue()) + self.assertTrue(report["passed"]) + self.assertTrue(report["pack_written"]) + self.assertEqual(30, report["result"]["held_out_participants"]) + serialized_report = json.dumps(report, ensure_ascii=False) + self.assertNotIn("private-human-intake", serialized_report) + self.assertNotIn("held-000", serialized_report) + self.assertNotIn("labeler-001", serialized_report) + + compiled = json.loads(output_path.read_text(encoding="utf-8")) + pack = G7HumanVoiceGainEvidencePack.model_validate(compiled) + self.assertEqual(1.0, pack.reliability.reported_icc) + self.assertEqual(1.0, pack.reliability.reported_categorical_kappa) + self.assertTrue( + all(not item.raw_audio_included for item in pack.observations) + ) + self.assertTrue( + all(not item.transcript_included for item in pack.observations) + ) + + def test_underpowered_intake_never_writes_a_pack(self) -> None: + with tempfile.TemporaryDirectory() as root: + directory = Path(root) / "intake" + directory.mkdir() + _write_intake(directory, _valid_payload()) + output_path = Path(root) / "compiled.json" + + report, pack = MODULE.compile_intake(directory) + + self.assertIsNone(pack) + self.assertFalse(report["passed"]) + self.assertIn( + "production_participant_floor", + report["result"]["failure_reasons"], + ) + self.assertFalse(output_path.exists()) + + def test_reported_reliability_must_match_rows(self) -> None: + with tempfile.TemporaryDirectory() as root: + directory = Path(root) / "intake" + directory.mkdir() + payload = human_pack() + reliability = payload["reliability"] + assert isinstance(reliability, dict) + reliability["reported_icc"] = 0.8 + _write_intake(directory, payload) + + report, pack = MODULE.compile_intake(directory) + + self.assertIsNone(pack) + self.assertFalse(report["passed"]) + self.assertIn( + "reported_icc_matches_rows", + report["result"]["failure_reasons"], + ) + + def test_unexpected_raw_material_column_is_rejected_without_echo(self) -> None: + with tempfile.TemporaryDirectory() as root: + directory = Path(root) / "person-at-example.test" + directory.mkdir() + _write_intake(directory, human_pack()) + path = directory / MODULE.OBSERVATIONS_NAME + rows = list(csv.reader(path.read_text(encoding="utf-8").splitlines())) + rows[0].append("raw_transcript") + rows[1].append("private words") + with path.open("w", encoding="utf-8", newline="") as handle: + csv.writer(handle, lineterminator="\n").writerows(rows) + + report, pack = MODULE.compile_intake(directory) + + self.assertIsNone(pack) + self.assertEqual("csv_headers_invalid", report["errors"][0]["code"]) + serialized = json.dumps(report, ensure_ascii=False) + self.assertNotIn("private words", serialized) + self.assertNotIn("person-at-example.test", serialized) + + def test_existing_output_is_not_overwritten(self) -> None: + with tempfile.TemporaryDirectory() as root: + directory = Path(root) / "intake" + directory.mkdir() + _write_intake(directory, human_pack()) + output_path = Path(root) / "compiled.json" + output_path.write_text("preserve", encoding="utf-8") + stdout = io.StringIO() + with contextlib.redirect_stdout(stdout): + exit_code = MODULE.main( + ["--compile", str(directory), "--out", str(output_path)] + ) + + self.assertEqual(1, exit_code) + self.assertEqual("preserve", output_path.read_text(encoding="utf-8")) + report = json.loads(stdout.getvalue()) + self.assertEqual("output_exists", report["errors"][0]["code"]) + self.assertFalse(report["pack_written"]) + + +if __name__ == "__main__": + unittest.main()