#storage

2 postswith this tag

657,867 documents later: the full corpus in 3.5 GB, 6.7 million chunks, and the bugs that only appear at real scale

We built the production foundations on the 657,867 documents of the Official Journal of the Federation (DOF): the full compressed corpus occupies 3.52 GiB, the chunk index holds 6.73 million 91-byte recipes, and the binary-vector pilot confirms the vector index will fit in less than 1 GiB. Along the way: a GROUP BY that consumed 35 GB of disk, a document that became 21 GB of chunks, and a lesson in result parity.

Joaquín Bravo ContrerasRead more

Storing 31 GB of Text in 3 GB (and Recovering Every Byte): the Storage Proof of Concept

Fourth installment of the benchmark: we built the compressed corpus on 10,000 real documents. sqlite-zstd compresses 10.9x with random access, chunks are stored as 110-byte recipes instead of text, the word-search index ended up 15x smaller than feared, and TurboQuant quantization ties full-vector quality. Everything fits in ~11 GB.

Joaquín Bravo ContrerasRead more