Vector DB 하이브리드 검색
LHG 2026-06-07
처음 보는 사람도 이해되게 먼저 풀어둔다
키워드와 의미 검색은 서로 다른 형태로 실패한다. RRF가 그 실패를 메운다
getUserById
findUserById
점수 스케일이 다른 두 검색을 정규화 없이 합치는 방법이다
GIN, HNSW, RRF를 SQL 함수 하나에 담은 레퍼런스 구현이다
create table documents ( id bigint primary key generated always as identity, content text, fts tsvector generated always as (to_tsvector('english', content)) stored, embedding extensions.vector(512) ); create index on documents using gin(fts); create index on documents using hnsw (embedding vector_ip_ops);
연산자와 인덱스 방식이 짝이 안 맞으면 인덱스가 조용히 무시된다.
with full_text as ( select id, row_number() over( order by ts_rank_cd(fts, websearch_to_tsquery(query_text)) desc) as rank_ix from documents where fts @@ websearch_to_tsquery(query_text) limit least(match_count, 30) * 2 ), semantic as ( select id, row_number() over(order by embedding <#> query_embedding) as rank_ix from documents limit least(match_count, 30) * 2 ) select documents.* from full_text full outer join semantic on full_text.id = semantic.id join documents on coalesce(full_text.id, semantic.id) = documents.id order by coalesce(1.0/(rrf_k+full_text.rank_ix),0)*full_text_weight + coalesce(1.0/(rrf_k+semantic.rank_ix),0)*semantic_weight desc limit least(match_count, 30);
ts_rank_cd
FULL OUTER JOIN
coalesce
full_text_weight
semantic_weight
rrf_k
match_count
hnsw.ef_search
네이티브로 지원하는 DB와 직접 코드로 합쳐야 하는 DB로 나뉜다
client.query_points( collection_name="docs", prefetch=[ models.Prefetch(query=sparse_vec, using="sparse", limit=20), models.Prefetch(query=dense_vec, using="dense", limit=20), ], query=models.RrfQuery(rrf=models.Rrf(k=60)), limit=10, )
Qdrant는 최상위 순위를 0부터 세므로 다른 시스템과 점수를 직접 비교하면 안 된다.
response = docs.query.hybrid( query="food", alpha=0.5, # 0이면 키워드, 1이면 의미 검색 fusion_type=HybridFusion.RELATIVE_SCORE, # 또는 RANKED, 즉 RRF limit=10, )
alpha 방식은 직관적이지만 두 점수의 분포가 다르면 예상과 다르게 움직일 수 있다.
results = client.hybrid_search( collection_name="docs", reqs=[dense_req, sparse_req], ranker=RRFRanker(k=60), limit=10, )
Milvus는 2.5부터 BM25 함수를 컬렉션 스키마에 직접 정의할 수 있다.
recall, 속도, 메모리, 빌드 시간의 다차원 트레이드오프다
m
ef_construction
ef_search
DB가 죽어도 데이터가 살아남는지는 운영에서 가장 먼저 확인해야 한다
이 데이터셋이 RAM에 들어갈지는 가장 자주 틀리는 계산이다
수치는 변동이 크므로 의사결정의 유일한 근거로 쓰지 않는다
임베딩에는 원본 텍스트와 메타데이터가 함께 들어간다
alter table documents enable row level security; create policy tenant_isolation on documents using (tenant_id = current_setting('app.current_tenant')::uuid); set app.current_tenant = '...uuid...';
이 패턴 하나로 SQL 인젝션과 테넌트 간 데이터 누수를 함께 막는다.
측정 없이 유명한 것을 고르는 게 가장 큰 함정이다