d3ro-voice/server/supabase/migrations/20260929030000_knowledge_chunks_vector_index.sql
Yun Chan ba9ef9741e fix: red-team round 3 hardening across desktop, mobile, core and server
Batch of red-team r3 fixes that were in the working tree before the
2026-09-28 design overhaul, committed as one unit with their tests.

- desktop main: STT timeouts and sidecar, voice recording store, sync
  (credentials, audio, knowledge reindex, push gates), runtime
  provisioner, update policy, AltGr keybindings, voice-command policy,
  dictionary file codec/limits, meeting transcript condensing and a
  local recording ledger so interrupted-session recovery only closes
  meetings this device recorded (a phone's live meeting is left alone).
- mobile: login CSRF via implicit token callbacks rejected, account
  deletion/retention, durable queue retention, knowledge realtime
  without unfiltered DELETE, meeting re-record failure paths, cloud STT
  client, preferences store/resync.
- core: text chunking splits long unbroken transcripts to fit, template
  field policy, dictionary limits, meeting markdown inline handling.
- server: payple webhook policy and cancellation order scope, meeting
  document generation quota, team RPC null-role guard, unified LLM
  quota in-flight accounting, knowledge chunk vector index, meeting
  re-record failure paths (migrations 20260929*).
- ci: portable/runtime feed gates, update-policy schema, Forgejo file
  delete and alias planning.

Four older tests are updated to the new contracts rather than the old
behavior: token-pair auth callbacks are rejected, knowledge realtime no
longer subscribes to DELETE, long transcript lines are split, and
meeting recovery requires the local recording ledger for empty rows.
2026-09-28 20:45:52 +09:00

102 lines
4.1 KiB
PL/PgSQL

-- ============================================================================
-- Knowledge vector search: rank exactly within the caller's visible chunks
-- ============================================================================
--
-- Problem (20260410000004_pgvector_knowledge.sql)
-- * idx_knowledge_chunks_embedding was an IVFFlat index (lists = 100) built
-- in the same migration that added the embedding column, so it was
-- trained on zero vectors and its centroids are meaningless. Nothing ever
-- ran the REINDEX that the comment asked for.
-- * match_knowledge_chunks ended with
-- ORDER BY kc.embedding <=> query_embedding LIMIT match_count
-- which lets the planner walk that ANN index across ALL tenants with the
-- default ivfflat.probes = 1, and only then apply the user/team filter,
-- the similarity threshold and RLS. A user whose chunks are a small share
-- of the table usually got 0-1 rows although exact matches existed, so
-- search-knowledge returned `results: []` and answers were generated
-- without the user's documents.
--
-- Fix
-- 1) Drop the broken IVFFlat index. It is not replaced by another ANN index
-- (HNSW etc.): match_knowledge_chunks below never orders the whole table
-- by distance, so an ANN index could not serve it and would only add
-- write cost to embed-chunks. Per-caller candidate sets (own documents +
-- teams the caller belongs to) are small and are reached through
-- idx_knowledge_chunks_document.
-- 2) match_knowledge_chunks first materializes the ids of the documents the
-- caller can see, computes the exact cosine distance for their chunks
-- only, and ranks inside that set. The MATERIALIZED CTEs are an
-- optimization fence: the planner cannot turn the ORDER BY back into a
-- global approximate index scan, so recall is exact regardless of what
-- indexes exist on knowledge_chunks.
-- 3) match_count is clamped to [0, 50] (NULL -> default 5). LIMIT NULL used
-- to mean "no limit". The search-knowledge edge function already caps
-- requests at 20, so it is unaffected.
--
-- Signature, return columns, SECURITY INVOKER (RLS still applies) and the
-- existing EXECUTE grants are unchanged (CREATE OR REPLACE keeps the ACL).
-- ============================================================================
DROP INDEX IF EXISTS public.idx_knowledge_chunks_embedding;
CREATE OR REPLACE FUNCTION public.match_knowledge_chunks(
query_embedding vector(1536),
match_count integer DEFAULT 5,
similarity_threshold double precision DEFAULT 0.5
)
RETURNS TABLE (
id uuid,
document_id uuid,
chunk_index integer,
content text,
similarity double precision
)
LANGUAGE plpgsql
STABLE
AS $$
DECLARE
-- Upper bound on rows one call may return. Keep >= search-knowledge's
-- MAX_MATCH_COUNT (20).
max_match_count CONSTANT integer := 50;
effective_count integer := LEAST(GREATEST(COALESCE(match_count, 5), 0), max_match_count);
BEGIN
IF effective_count = 0 OR query_embedding IS NULL THEN
RETURN;
END IF;
RETURN QUERY
WITH visible_documents AS MATERIALIZED (
SELECT kd.id AS visible_document_id
FROM public.knowledge_documents kd
WHERE
kd.user_id = auth.uid()
OR (
kd.team_id IS NOT NULL
AND kd.team_id IN (
SELECT tm.team_id FROM public.team_members tm WHERE tm.user_id = auth.uid()
)
)
),
candidates AS MATERIALIZED (
SELECT
kc.id AS chunk_id,
kc.document_id AS chunk_document_id,
kc.chunk_index AS chunk_position,
kc.content AS chunk_content,
(kc.embedding <=> query_embedding)::double precision AS distance
FROM visible_documents vd
INNER JOIN public.knowledge_chunks kc ON kc.document_id = vd.visible_document_id
WHERE kc.embedding IS NOT NULL
)
SELECT
c.chunk_id,
c.chunk_document_id,
c.chunk_position,
c.chunk_content,
(1 - c.distance)::double precision AS similarity
FROM candidates c
WHERE (1 - c.distance) > similarity_threshold
ORDER BY c.distance, c.chunk_id
LIMIT effective_count;
END;
$$;