Skip to content

teff.rag.chunker

teff.rag.chunker

Text chunking strategies for RAG document splitting.

Classes:

Name Description
Chunker

Split text into chunks for embedding and retrieval.

Chunker dataclass

Split text into chunks for embedding and retrieval.

Supports three strategies:

  • token — Split on whitespace into token windows.
  • sentence — Split on sentence boundaries.
  • fixed — Split by fixed character count.

Attributes:

Name Type Description
strategy str

Chunking strategy name.

chunk_size int

Target chunk size (tokens, sentences, or chars).

overlap int

Overlap between consecutive chunks.

Methods:

Name Description
chunk

Split text into chunks using the configured strategy.

Source code in teff/rag/chunker.py
 6
 7
 8
 9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
@dataclass
class Chunker:
    """Split text into chunks for embedding and retrieval.

    Supports three strategies:

    - ``token`` — Split on whitespace into token windows.
    - ``sentence`` — Split on sentence boundaries.
    - ``fixed`` — Split by fixed character count.

    Attributes:
        strategy: Chunking strategy name.
        chunk_size: Target chunk size (tokens, sentences, or chars).
        overlap: Overlap between consecutive chunks.
    """

    strategy: str = "token"
    chunk_size: int = 500
    overlap: int = 50

    def chunk(self, text: str) -> list[str]:
        """Split *text* into chunks using the configured strategy."""
        if self.strategy == "token":
            return self._chunk_token(text)
        if self.strategy == "sentence":
            return self._chunk_sentence(text)
        if self.strategy == "fixed":
            return self._chunk_fixed(text)
        raise ValueError(f"unknown chunk strategy: {self.strategy}")

    def _chunk_token(self, text: str) -> list[str]:
        tokens = text.split()
        chunks = []
        start = 0
        while start < len(tokens):
            end = start + self.chunk_size
            chunk = " ".join(tokens[start:end])
            chunks.append(chunk)
            start += self.chunk_size - self.overlap
            if self.chunk_size - self.overlap <= 0:
                break
        return chunks

    def _chunk_sentence(self, text: str) -> list[str]:
        import re

        sentences = re.split(r"(?<=[.!?])\s+", text)
        chunks = []
        current = []
        for s in sentences:
            current.append(s)
            if len(current) >= self.chunk_size:
                chunks.append(" ".join(current))
                overlap_start = max(0, len(current) - self.overlap)
                current = current[overlap_start:]
        if current:
            chunks.append(" ".join(current))
        return chunks

    def _chunk_fixed(self, text: str) -> list[str]:
        chunks = []
        start = 0
        while start < len(text):
            end = start + self.chunk_size
            chunks.append(text[start:end])
            start += self.chunk_size - self.overlap
            if self.chunk_size - self.overlap <= 0:
                break
        return chunks

chunk

chunk(text)

Split text into chunks using the configured strategy.

Source code in teff/rag/chunker.py
26
27
28
29
30
31
32
33
34
def chunk(self, text: str) -> list[str]:
    """Split *text* into chunks using the configured strategy."""
    if self.strategy == "token":
        return self._chunk_token(text)
    if self.strategy == "sentence":
        return self._chunk_sentence(text)
    if self.strategy == "fixed":
        return self._chunk_fixed(text)
    raise ValueError(f"unknown chunk strategy: {self.strategy}")