# What TokenTextSplitter does to a 40-sentence document document: 4268 characters, 880 cl100k_base tokens --- new TokenTextSplitter() (defaults: 800 tokens, 350 min chars, 5 min length to embed) --- chunks: 2, sizes in tokens: [792, 88] --- chunk size 100 tokens --- chunks: 10 chunk 0: 88 tokens, 419 chars, starts "Sentence 1 describes t", ends "ry full-time employee." chunk 1: 88 tokens, 420 chars, starts "Sentence 5 describes t", ends "ry full-time employee." chunk 2: 88 tokens, 426 chars, starts "Sentence 9 describes t", ends "ry full-time employee." chunk 3: 88 tokens, 428 chars, starts "Sentence 13 describes ", ends "ry full-time employee." chunk 4: 88 tokens, 427 chars, starts "Sentence 17 describes ", ends "ry full-time employee." chunk 5: 88 tokens, 427 chars, starts "Sentence 21 describes ", ends "ry full-time employee." chunk 6: 88 tokens, 428 chars, starts "Sentence 25 describes ", ends "ry full-time employee." chunk 7: 88 tokens, 428 chars, starts "Sentence 29 describes ", ends "ry full-time employee." chunk 8: 88 tokens, 428 chars, starts "Sentence 33 describes ", ends "ry full-time employee." chunk 9: 88 tokens, 427 chars, starts "Sentence 37 describes ", ends "ry full-time employee." metadata of chunk 1 (parent_document_id, a random UUID, left out): {chunk_index=1, page_number=1, source_file=long.txt, total_chunks=10} --- is there any overlap between neighbouring chunks? --- boundaries where the first 30 characters of a chunk already appear in the chunk before it: 0 of 9 --- minChunkSizeChars: where the cut lands --- minChunkSizeChars=350: 10 chunks, 9 of the first 9 end on a full stop minChunkSizeChars= 50: 10 chunks, 9 of the first 9 end on a full stop