## `EvidenceChunker`
>`evidence_chunker.EvidenceChunker`
PDF → Evidence Unit 파이프라인의 메인 진입점.
```python
EvidenceChunker(
parser: PdfParser | None = None,
artifacts_path: str | None = None,
bbox_threshold: float = 300.0,
sim_threshold: float = 0.0,
)
```
| 파라미터 | 타입 | 설명 |
| --- | --- | --- |
| `parser` | `PdfParser \| None` | `parse(path) -> ParsedDoc` 구현체.
`None`이면 기본 `DoclingParser` 사용. 주입 시 `chunk()`만 교체 가능(`build_corpus()`는 미지원) |
| `artifacts_path` | `str \| None` | Docling 로컬 모델 경로. (`parser` 직접 주입 시 무시) |
| `bbox_threshold` | `float` | 표 위/아래 단락 수집 범위(pt). 기본 300.0 |
| `sim_threshold` | `float` | 문맥 단락 채택 코사인 유사도 임계값. 기본 0.0 |
### 메서드
| 메서드 | 반환 타입 | 설명 |
| --- | --- | --- |
| `chunk(pdf_path, doc_id=None)` | `list[EvidenceUnit]` | PDF에서 표(Evidence Unit)만 추출. 512토큰 초과 EU는 행 단위로 분할된 상태로 반환 |
| `build_corpus(pdf_path, doc_id=None)` | `list[RetrievalChunk]` | 표(EU) + 일반 본문(`TextChunk`)을 합친 검색 코퍼스.
커스텀 `parser` 주입 시 `NotImplementedError` |
---
## `EvidenceUnit`
> `evidence_chunker.EvidenceUnit`
표 하나를 표현하는 검색 단위 dataclass.
```mermaid
flowchart TD
L1["Case 1: 일반"] ~~~ P1([p.1]) --> EU1["eu_id: p1-1
page_span={1}"]
L2["Case 2: 인접 페이지
문맥 흡수"] ~~~ P2([p.2]) --> EU2["eu_id: p2-1
page_span={1,2}"]
L3["Case 3: 512토큰 초과
→ 행 분할"] ~~~ P4([p.4])
P4 --> EU4a["p4-1-s1
행 1~N"]
P4 --> EU4b["p4-1-s2
행 N+1~"]
style P1 fill:#EBCB8B,stroke:#D08770,stroke-width:1.5px,color:#2E3440
style P2 fill:#EBCB8B,stroke:#D08770,stroke-width:1.5px,color:#2E3440
style P4 fill:#EBCB8B,stroke:#D08770,stroke-width:1.5px,color:#2E3440
style EU1 fill:#88C0D0,stroke:#5E81AC,stroke-width:2px,color:#2E3440
style EU2 fill:#88C0D0,stroke:#5E81AC,stroke-width:2px,color:#2E3440
style EU4a fill:#A3BE8C,stroke:#4C566A,stroke-width:1.5px,color:#2E3440
style EU4b fill:#A3BE8C,stroke:#4C566A,stroke-width:1.5px,color:#2E3440
style L1 fill:none,stroke:none,color:#4C566A
style L2 fill:none,stroke:none,color:#4C566A
style L3 fill:none,stroke:none,color:#4C566A
```
### 주요 필드
| 필드 | 타입 | 설명 |
| --- | --- | --- |
| `eu_id` | `str` | `"{doc_id}-p{page}-{idx}"` 형식 식별자. 분할 조각은 `f"{eu_id}-s{n}"` |
| `page_no` | `int` | 표 자신의 페이지 |
| `page_span` | `set[int]` | 표+채택된 문맥 단락이 걸친 모든 페이지 |
| `caption_text` | `str \| None` | 연결된 캡션 원문 |
| `caption_confidence` | `"direct" \| "inferred" \| "none"` | 캡션 연결 신뢰도 |
| `table_html` | `str \| None` | 표 HTML (Docling `export_to_html()` 결과) |
| `context_before` / `context_after` | `list[str]` | 표 위/아래 채택된 설명 단락 |
| `flattened_rows` | `list[str]` | 셀 데이터를 "행헤더 \| 열헤더: 값" 문장으로 변환한 목록 |
| `table_abstract` | `str \| None` | 규칙 기반 표 요약(캡션+열헤더+행 수) |
| `bbox` | `tuple[float, float, float, float]` | 0~1 정규화 BOTTOMLEFT 좌표 |
| `is_split` / `split_index` / `total_splits` | `bool` / `int \| None` / `int \| None` | 분할 여부와 조각 정보 |
### Property (자동 계산, 직접 대입 금지)
| Property | 반환 타입 | 설명 |
| --- | --- | --- |
| `text` | `str` | LLM 컨텍스트용 전체 텍스트. `table_html` 포함 |
| `retrieval_text` | `str` | 임베딩/검색용. `table_html` **제외** (대신 `flattened_rows` 사용) |
| `retrieval_units` | `list[str]` | 표 요약 + 문단/행/각주를 개별 단위로 쪼갠 목록 (small-to-big 패턴용) |
| `chunk_id` | `str` | `eu_id`와 동일. `export.RetrievalChunk` 프로토콜 구현 |
| `is_atomic` | `bool` | 항상 `False` (`retrieval_units`로 쪼갤 수 있는 대상) |
| `metadata` | `dict` | LangChain/LlamaIndex 문서 메타데이터
- `chunk_id`, `eu_id`, `page_span`, `caption_text` 등 |
| `safe_caption` | `str \| None` | 캡션이 Fig/Figure/그림으로 시작하면 `None`
- 그림 캡션이 표 캡션으로 오인되는 Case 방어 |
### `retrieval_text` vs `text`
| 상황 | 쓸 것 |
| --- | --- |
| 벡터스토어에 임베딩할 때 | `retrieval_text` (HTML 노이즈 없음, 토큰 절약) |
| 검색된 EU를 LLM 컨텍스트로 넘길 때 | `text` (표 구조를 HTML로 보존) |
`to_langchain()`은 기본으로 `page_content=eu.text`를 사용한다.
검색 정확도를 최대화하려면 벡터스토어 구성 시 `page_content=eu.retrieval_text`로 직접 변경해 사용하는 것을 권장한다.
---
## `evidence_chunker.export.langchain`
| 함수/클래스 | 시그니처 | 설명 |
| --- | --- | --- |
| `to_langchain` | `(chunks: list[RetrievalChunk]) -> list[Document]` | 1 chunk = 1 Document. `page_content=c.text` |
| `to_langchain_units` | `(chunks: list[RetrievalChunk]) -> list[Document]` | small-to-big.
`is_atomic=False`인 chunk만 `retrieval_units` 단위로 쪼갬 |
| `dedupe_by_chunk_id` | `(results, k=None, key="chunk_id") -> list` | 정렬된 검색 결과에서 같은 `chunk_id`끼리 첫 등장(=최고 점수)만 남김 |
| `EvidenceRetriever` | `(vectorstore, k=5, fetch_k=None, dedupe=True)` | max-pool dedupe가 기본 적용된 검색 래퍼.
`get_relevant_documents(query)` / `invoke(query)` 제공 |
`EvidenceRetriever`는 `vectorstore.similarity_search_with_score(query, k=fetch_k)`를 지원하는 LangChain `VectorStore`가 필요하다. `fetch_k` 기본값은 `max(k*4, 20)`.
---
## `evidence_chunker.export.llamaindex`
`export.langchain`과 동일한 함수 구성(`to_llamaindex`, `to_llamaindex_units`, `dedupe_by_chunk_id`, `EvidenceRetriever`)을 LlamaIndex `TextNode`/retriever 기준으로 제공한다.
`EvidenceRetriever(base_retriever, k=5, dedupe=True)`는 `.retrieve(query)`를 제공하는 LlamaIndex retriever를 감싼다.
---
## `evidence_chunker.export`
| 함수/클래스 | 설명 |
| --- | --- |
| `RetrievalChunk` (Protocol) | `chunk_id`/`is_atomic`/`text`/`retrieval_text`/`retrieval_units`/`metadata` 속성 계약.
`EvidenceUnit`과 `TextChunk` 둘 다 만족 |
| `TextChunk` | Docling `HybridChunker`가 만든 일반 본문 청크를 `RetrievalChunk` 프로토콜로 감싼 래퍼. `is_atomic=True` |
| `filter_consumed_paragraphs(chunks, eu_list, min_substring_len=20)` | EU가 이미 흡수한 문단과 겹치는 일반 청크 제거 (카니발라이제이션 방지) |
---
## `evidence_chunker.split`
| 함수 | 시그니처 | 설명 |
| --- | --- | --- |
| `split_eu` | `(eu, limit=SPLIT_LIMIT) -> SplitResult` | EU 하나를 토큰 한도에 맞춰 분할.
- `single`: 한도 이내, 그대로 통과
- `row_split`: 행 단위 분할
- `llm_summary`: 실제 요약은 하지 않고 원본 그대로 통과, 상위 레이어에 요약 필요 신호만 전달 |
| `split_oversized_units` | `(eu_list, stats=None) -> list[EvidenceUnit]` | 한도 초과 EU 전체를 일괄 분할 |
---
## `evidence_chunker.tokens`
| 함수 | 설명 |
| --- | --- |
| `count_tokens(text)` | tiktoken(`cl100k_base`) 기준 토큰 수 |
| `count_eu_tokens(eu)` | `eu.text` 기준 토큰 수 |
| `exceeds_token_limit(eu, limit=DEFAULT_TOKEN_LIMIT)` | 한도 초과 여부 |
| `DEFAULT_TOKEN_LIMIT` | `512` |