The Library of Babel; Embedding Information
Origionally Written for Math and Philosophy (2024) - Ted Theodosopoulos
The Library of Babel, imagined by writer and poet Jorge Luis Borges in his short story of the same name, is a theoretical, infinitely large library containing all possible information, constructed of symbols both familar and unknown. This library would contain the sum total of all human knowledge, as well as all future works which could be created. However, these peices would be lost in a sea of gibberish.
How could we define this library from a mathematical standpoint? The simlplest case, though we would lose information, is that of a standard alphabet and size. Let’s say we have 32 characters, each letter in the English alphabet, along with some basic punctuation symbols: (, . : ? ; space). We’re skipping quite a lot here, such as all of mathematics. However, with this set, it’s not unreasonable to claim that most information is conveyable.
In the original story, each book has a standard template: 410 pages, 40 lines on each page, and 80 characters in each line. We can imagine a library that contains all possible combinations of characters needed to fill this book. This library would be unimaginably massive, with the number of possible books having over 2 million digits, assuming there are no repetitions. Given, most of these would consist of nothing but random letters. However, it would also contain exact copies of all books written with less than 1.6 million characters, and variations of each with one-character differences. It contains all possible iterations of ASCII art, and therefore copies of every painting made with symbols.
Let’s abstract this a bit, stepping away from the story itself. It would be better not to have a limit on the number of characters we can put, so let’s do just that. We still have constraints, namely the fact that while the books themselves can be any size, they still must be finite. This new library would be truly infinite, but would still be countable, due to our defined number of characters. To create a correspondence between this library and the natural numbers, we could assign an order to our characters, then order our books as such:
1,1,1,1… 1
1,1,1,1… 2
…
1,1,1,1… 32
1,1,1,1… 2, 1
1,1,1,1… 2, 2
…
1,1,1,1… 2, 32
1,1,1,1… 3, 1
…
1,1,1,1… 32, 32
1,1,1,1… 2, 1, 1
…
32, 32, 32… 32
This is similar to counting in base 32. With this, we can establish a correspondence with the natural numbers, meaning the number of books is countably infinite.
So how would we organize the library when accepting books of any alphabet? We could use Paul Dancstep’s presentation of the library. For each combination of a number of characters and a number of spaces we can put them in, we create a collection of possible books.
We can organize these collections in a grid pattern, with the block closest to the origin using an alphabet with 1 character and having 1 space for it to go. One axis is the number of characters, and the other is the number of spaces. This grid is filled with blocks, which are collections of the possible books that can be in them. We can create a clear mapping between the structure of these categories and the x-y plane, and therefore, the number of categories is countable. Additionally, each collection will contain a finite number of books, so we can conclude that this new library is still countable.
One thing which we have overlooked is the order of characters, which is necessary for us to order the books within these categories. We can order each collection with the same process we used to count the books earlier, allowing us to call books directly. We could imagine a distinct vector that allows us to find the exact book we want with three distinct variables: alphabet length, book length, and the index within each collection. This is a useful way to index our library, though it is limited. To perform any analysis of a book, we would have to return to our model and find the structure of the book we are referencing. Also, we would be unable to perform operations between the vectors themselves, as they only act as an index. For the purpose of this essay, we want to try and find something generalizable to linear algebra, in which we have a very established framework for the interaction between vectors. While our current model does index books as vectors, we can’t perform actions such as matrix multiplication or addition. Let’s try another approach which uses linear algebra.
Before we do so, we have to slightly shift our focus. Exploring our infinite library has led us to a deeper task: quantifying langauge. We must break down the books into paragraphs, sentences, words, and even morphemes (pre, anti, etc.). Therefore, when moving forward, we will keep the same structure of the library in mind, but instead of thinking of books in their entirety, we will think of short words with much smaller amounts of characters. Another point: while we wish we could use every language format, there are different meanings for the same symbols across languages. Therefore, it will be extremely difficult to derive meaning unless we decide on a specific language. We can even generalize our format so it works with any language, but once we choose one, we can’t switch interchangeably. For now we may stick to english.
Instead of trying to fit an index of the books into a vector, we can build the system from the basis. We will avoid the generalization of languages for now and focus on defining a vector space for each possible alphabet. This way, we can represent words as vectors, and in this vector space perform interactions between them, allowing us to interpret them mathematically. For example, say we have a collection of books with 5 possible characters. We can establish a basis from these characters as such:
[1, 0, 0, 0, 0] for A
[0, 1, 0, 0 ,0] for B
[0, 0, 1, 0, 0] for C
[0, 0, 0, 1, 0] for D
[0, 0, 0, 0 ,1] for L
We can generalize this format to languages with any number of characters, simply adding more bases. Now, when making a book, sentence, or word, we can put together these rows in a matrix. For example, CABAL would look like such:
[0, 0, 1, 0, 0]
[1, 0, 0, 0, 0]
[0, 1, 0, 0, 0]
[1, 0, 0, 0, 0]
[0, 0, 0, 0, 1]
This is interesting, and now we have strings represented in a matrix, though when putting them together we might need to add a space into our set of characters. However, this is not that practical, especially if we want to represent something as long as a full book. As of now, we have little more than a simple cipher. But this structure is the building block of more complicated mathematical representations of language. Let’s discuss embedding.
One of the most prevalent cases of language-to-math conversion is AI and the study of natural language processing. However, they don’t observe texts using this one-hot vector form. This form allows us to represent each character as a vector, but it is rigid and mostly uninformative, not telling us anything about how one character relates to another. Therefore, an AI encodes words using an embedding matrix. An embedding matrix defines, for each of our characters, a specific number that determines the placement of that character. It might be easier to work with an example:

General form:
[a1, a2, a3, a4, a5]
[b1, b2, b3, b4, b5]
[c1, c2, c3, c4, c5]
[d1, d2, d3, d4, d5]
[l1, l2, l3, l4, l5]
Our word cabal might look like:
\[\begin{bmatrix} c_1 & c_2 & c_3 & c_4 & c_5 \\ a_1 & a_2 & a_3 & a_4 & a_5 \\ b_1 & b_2 & b_3 & b_4 & b_5 \\ a_6 & a_7 & a_8 & a_9 & a_{10} \\ l_1 & l_2 & l_3 & l_4 & l_5 \end{bmatrix}\]Each row is a 5-dimensional embedding for the corresponding character. We can use this to express a word as a matrix of embedded vectors. One important caveat here is that while this form is usually discussed as a matrix, it is more akin to a list of vectors for each letter, which when combined act as a matrix. When referencing a word, AI needs not only the letters comprising it, but also the context surrounding their usage. In this form, the rows determine the letter order, and we still have the other information about weights.
These weights are not random, they are chosen by AI systems to represent the use of those characters in the position which each weight refers to. The weights are generally not readable by humans, or at least we have little to no intuition surrounding them. AI models are trained to create these weights. To understand this further, we need to break down the way AI models work, and how they are trained.
AI is trained by human input in the form of pre-labeled sentences which contain meaning. Sometimes, this meaning is made explicit, as would be the case with early stage development, using small sets of sorted data. However, it could also assume this implicitly, such as when training on huge sets of data such as the internet. Without this training, an AI has no way to distinguish real words from random combinations of letters.
This system is different from the way humans learn, but there are many parallels to be drawn, and without training of our own, humans are similarly incapable of deciphering meaning. If a human were to walk around the Library of Babel, and pick books to read at random, almost every time they would encounter nothing but random strings of letters. However, If a book you pick up happens to be a real story written in English, I, or someone similar, would recognize it as containing meaning. This decision would be based on prior knowledge and understanding of the English language. If there was a book written in a language you had no knowledge of, you would have a very difficult time differentiating it from something meaningless, or random letters. In the same way, a human who somehow had no knowledge of written language in any form would have no idea what anything in the library means, including books written in English.
Before training, AI acts similarly to someone like this, with no knowledge of how sentences are put together. When it is trained with human labeled data, it recalibrates its weights to fit the real examples of meaning. For example, an AI might receive a sentence such as
“The cat slept on the ___.”
It would then attempt to guess what the last word might be. To do this, it would use the weights and determine what it believes is most likely to conclude the sentence, based on past data it has seen. Let’s say the AI decided on “roof”, though the actual answer was “mat”. It would then compare the response it gave, calculating the dot product between its proposed answer and the correct answer. Then, it changes its weights, altering its prediction matrices to incorporate the error. Through this process, an AI develops its embedding matrix, which is used to discern the probabilities of words being used in the context of their surroundings. Advanced models are trained on unimaginable amounts of data, which is what allows it to be so accurate to human speech. Additionally, they can determine when specific kinds of responses are appropriate, for example when prompted to respond casually versus academically, they have different probability matrices for the way they decide their words.
Before we link AI back to our infinite library, let’s finish our explanation of how it works. When an AI is reading a sentence, as much as we would like it to be the case, it doesn’t just generate an embedded sequence of matrices for each word. First, it breaks down the sentence into tokens, which are normally words, but can be meaningful subsections of words as well. It then assigns positional matrices, which tell the AI the order in which the tokens are organized. Then, it sends the tokens through its constructed embedding matrix, and outputs them as embedded vectors. So the example I gave earlier is not entirely representative, most of the time a sentence will be represented as a list of embedded vectors corresponding to each word. The length of an embedded vector is dependent on the sophistication of the language model being used, as it is what the AI uses to encode meaning. This length can range from dozens to thousands. So the full process is as such: tokenization → embedding matrix → positional matrix.
Let’s explore how an AI might be able to help us organize our library of babel. If we are starting with a completely untrained AI, it would be even more lost than us, as even if it encountered sensible text it would not know what it meant. However, let’s say in addition to the AI we had a human interpreter. The AI could bring to this person books it believes to be containing real text, and have them assess its classification. Over time, by recognizing patterns in the books it brings which contain real information, the AI would gradually train itself. After enough time has passed, it will be able to correctly distinguish between sense and nonsense with some level of acceptable accuracy. We could then have this AI read as much of the library as possible, and return with the collection of books it deems to be sensible. This is a possible method of sorting the library, and probably one of the more efficient ones.
When Borges originally wrote his short story pertaining to this library, it was laced with futility, as he believed that the library was unexplorable to a meaningful extent. He begins the story with hope for the utility of the library, but as the characters explore the library, the amount of useless data becomes insurmountable. The number of books is truly infinite, and therefore, the number of illogical books is infinite as well, along with the number of logical ones. However, with the method described using AI, we can observe a very large number of books, arguably any finite number given enough time. Therefore, depending on how the library is organized, we could imagine sorting all books under a certain specific length. This throws the pessimistic view of the library from borges into question, as even if we can’t truly look through each and every book, we can find some extremely interesting things. There’s so many more questions to ask, and we haven’t even begun to delve into many of the interesting philosophical ramifications of having such a library, which was the main focus of the original story. However we have used it as an excuse to mathematically interpret language and explain an interesting technology, which I think is pretty cool. Thanks for reading!