pipeline/ — 최종 코드 + 데이터 + 결과 + 보고서
tests/matching/ 의 320 개 파일 중 최종본만 복사한 정리 폴더.
tests/matching/ 는 원본 그대로 보존 (이 폴더는 복사본).
상위 문서: ../PIPELINE.md — V1~V4 시스템 통합 정리
폴더 구조
code/ (30 개) — 매칭 시스템 코드
메인 pipeline (V1 ~ V4)
| 파일 |
단계 |
역할 |
pipeline_06_2_mdx_matching.py |
V1 |
키워드 매칭 — 0.414 핵심 + 0.320 세트 + 0.265 연관 |
pipeline_08_v2_semantic_rerank.py |
V2 |
의미 매칭 — ko-sroberta cosine |
pipeline_08_v3_r5_structure_rerank.py |
V3 (최종) |
구조 매칭 — 0.40 layout + 0.35 aff + 0.25 intent |
pipeline_17_v4_full32.py |
V4 (최종) |
종합 판정 — 32 frame 전체 평가, 5축 가중합 |
사전 작업 (입력 + 키워드 정리)
| 파일 |
역할 |
pipeline_01_extract_nodes.py |
Figma + BEPS + MDX 텍스트 추출 (MDX_SECTIONS, MDX_DIR 정의) |
pipeline_02_tokenize.py |
Kiwi 형태소 분석 |
pipeline_02_5_special_forms.py |
특수 형식 (DX/BIM 등) 보존 |
pipeline_04_normalize.py |
표기 통합 + 동의어 처리 |
pipeline_05_anchor_candidates.py |
핵심 키워드 후보 추출 |
pipeline_06_anchor_sets.py |
키워드 세트 생성 |
pipeline_07_auto_anchor_candidates.py |
자동 anchor 후보 |
Frame 구조 라벨링 (templates_v2)
| 파일 |
역할 |
pipeline_12_generate_templates_v2.py |
r1 — 초기 ontology 생성 (AFFINITY_KEYWORDS, INTENT_KEYWORDS 정의) |
pipeline_12_r2_generate_templates_v2.py |
r2 |
pipeline_12_r3_generate_templates_v2.py |
r3 (최종 키워드 사전) |
pipeline_12_finalize_v2.py |
final ontology 마무리 |
V3 의존 라운드 (r5 가 r4, r3, r2 의존)
| 파일 |
역할 |
pipeline_08_v3_r2_structure_rerank.py |
r2 — MDX_LAYOUT_STRUCTURE 정의 |
pipeline_08_v3_r3_structure_rerank.py |
r3 — STRONG 패턴 검증 도입 |
pipeline_08_v3_r4_structure_rerank.py |
r4 — detect_mdx_v2_profile_r4, v3_r4_score 정의 |
검증 / Ablation
| 파일 |
역할 |
pipeline_15_logistic_regression.py |
V1 가중치 학습 + LOOCV 검증 (4/4) |
pipeline_15_bm25_with_sets.py |
BM25 비교 (현 방식 우위 확인) |
pipeline_18_slot_axis_ablation.py |
V4 slot 축 ablation (Top-1 7/7 동일 확인) |
보고서 생성
| 파일 |
역할 |
pipeline_16_deck_4pages.py |
DECK 1~7 HTML 생성 |
의존 모듈
| 파일 |
역할 |
detect_mdx.py |
MDX 본문 분석 (detect_mdx_analysis, slots_from_table 등) |
template_fit.py |
V4 confidence 계산 (compute_template_fit, route, 5축 가중치) |
structure_v3.py |
V3 score 함수 (structural_match_v3) |
phase_common.py |
공통 유틸 (detect_mdx_layout_v2, _COMPAT, 표기 통합) |
methods.py |
Kiwi 캐싱, content token 추출, layout 감지 |
embeddings.py |
ko-sroberta 임베딩 + cosine |
common.py |
단순 토크나이저, char_ngrams |
keyword_normalizer.py |
키워드 정규화 |
data/ (14 개) — 입력 데이터 + Frame DB + 키워드 사전
Frame 구조 DB
| 파일 |
내용 |
structure_ontology_v2_final_r2.yaml |
32 Frame 구조 DB (layout, family, cardinality, content_affinity, structure_intent, slots, alternative_patterns) |
키워드 사전
| 파일 |
내용 |
auto_anchor_candidates.yaml |
키워드 sets (frame 별 source_text_sets, terms, frame_stats) |
anchor_sets_input.yaml |
anchor sets 입력 |
keyword_base.yaml |
기본 키워드 사전 |
domain_terms.yaml |
도메인 용어 |
synonyms_production.yaml |
동의어 (운영용) |
synonyms.yaml |
동의어 후보 |
정답 / 검증 라벨
| 파일 |
내용 |
ground_truth.yaml |
정답 매칭 라벨 |
holdout_labels.yaml |
Holdout 평가 라벨 |
중간 데이터 (전처리 결과)
| 파일 |
내용 |
text_canonical.yaml |
표준화된 텍스트 |
normalized_text_tokens.yaml |
정규화된 토큰 |
actual_text_nodes.yaml |
추출된 텍스트 노드 |
actual_text_tokens.yaml |
추출된 토큰 |
special_forms_inventory.yaml |
특수 형식 인벤토리 |
results/ (4 개) — V1~V4 최종 결과
| 파일 |
단계 |
핵심 정보 |
mdx_matching_result.yaml |
V1 |
per_frame_detail (standalone/keyword_group/related 점수, hit/total), rank_by_matching_score |
v2_semantic_rerank_result.yaml |
V2 |
v2_rerank (semantic_score, V2 순위) |
v3_structure_rerank_r5_result.yaml |
V3 |
v3_r5_rerank (total, breakdown — layout_compat/content_affinity/structure_intent), mdx_profile |
v4_full32_result.yaml |
V4 |
judgments_full32 (32 frame 전체 confidence + label), axes (5축 점수) |
reports/ (14 개) — 임원 보고용 + 보조 보고서
임원 보고용 7 페이지
| 파일 |
내용 |
DECK_01_TARGET_01_2.html |
TARGET MDX 01-2 × V1~V4 Top-3 매칭 결과 (이미지 + 점수) |
DECK_02_HOLDOUT_02_2_1.html |
Holdout MDX 02-2.1 × V1~V4 Top-3 |
DECK_03_METHODS.html |
V1~V4 방법 설명 |
DECK_04_DB_STRUCTURE.html |
MDX 1 vs Top-3 frame 5 컬럼 비교표 — 키워드/구조/종합 매칭 |
DECK_05_STRUCTURE_KEYWORD.html |
구조 매칭 4 항목 상세 (콘텐츠 성격/시각 의도/편집 슬롯/항목수·관계) + 키워드 사전 |
DECK_06_FIGMA_TO_DB.html |
Frame 18 의 DB 저장 형태 샘플 (YAML) |
DECK_07_KEYWORD_PIPELINE.html |
키워드집 정리 과정 + Frame 18 샘플 |
※ DECK 1, 3, 4, 5, 7 은 사용자가 직접 편집한 (최종) 버전 복사. matching/ 의 DECK_XX(최종).html 이 원본.
보조 보고서
| 파일 |
내용 |
KEYWORD_INVENTORY.html / .md |
키워드집 생성 보고서 (1,713 고유 + 13 중요 표기 + 451/1238/24 분포) |
ATTACH_07_WEIGHT_VALIDATION.html |
V1 가중치 Logistic Regression + LOOCV 검증 |
V4_SLOT_ABLATION.md |
V4 slot 축 ablation (Top-1 7/7 동일 — 영향 없음) |
FIGMA_KEYWORDS_REPORT.html / .md |
Frame 별 키워드 인벤토리 |
FRAME_KEYWORD_REVIEW.html / .md |
Frame 별 키워드 검토 |
사용 시나리오
새 작업자 (또는 새 Claude) 가 시스템을 처음 보는 경우
- ../PIPELINE.md 읽기 — V1~V4 통합 정리 (산식 / 가중치 / 도출 과정 / 결과 / 약점)
reports/DECK_03_METHODS.html 보기 — V1~V4 방법 시각화
reports/DECK_04_DB_STRUCTURE.html 보기 — 실제 매칭 비교 사례
code/pipeline_17_v4_full32.py 읽기 — V4 최종 매칭 로직
매칭 결과 분석
results/mdx_matching_result.yaml (V1)
results/v3_structure_rerank_r5_result.yaml (V3)
results/v4_full32_result.yaml (V4)
Frame DB 확인
data/structure_ontology_v2_final_r2.yaml
data/auto_anchor_candidates.yaml
검증 / Ablation 결과
reports/ATTACH_07_WEIGHT_VALIDATION.html (V1 가중치)
reports/V4_SLOT_ABLATION.md (slot 축)
주의
- 이 폴더는 복사본.
tests/matching/ 가 원본.
- 코드 수정 시 두 곳 동기화 필요 (또는 matching/ 만 수정 후 다시 복사).
- 코드 실행 시 import 가 같은 디렉토리 가정 (
from detect_mdx import ...) — code/ 디렉토리 안에서 실행 필요.