<?xml version="1.0" encoding="utf-8"?><feed xmlns="http://www.w3.org/2005/Atom" ><generator uri="https://jekyllrb.com/" version="4.4.1">Jekyll</generator><link href="https://lukakirigin.com/feed.xml" rel="self" type="application/atom+xml" /><link href="https://lukakirigin.com/" rel="alternate" type="text/html" /><updated>2026-10-11T21:10:18+00:00</updated><id>https://lukakirigin.com/feed.xml</id><entry><title type="html">BRIEF: Lagrange Multipliers</title><link href="https://lukakirigin.com/blog/2026/07/17/Lagrange/" rel="alternate" type="text/html" title="BRIEF: Lagrange Multipliers" /><published>2026-07-17T00:00:00+00:00</published><updated>2026-07-17T00:00:00+00:00</updated><id>https://lukakirigin.com/blog/2026/07/17/Lagrange</id><content type="html" xml:base="https://lukakirigin.com/blog/2026/07/17/Lagrange/"><![CDATA[<p>Find the maximum value of</p>

\[f(x,y,z) = x^2 + y^2 - z\]

<p>subject to the constraint</p>

\[x + y - 2z = 0\]

<p>The function f maps a subset of $\mathbb{R}^3$ to real values in $\mathbb{R}$, and the constraint is a level set of a similar function $g$. What is a level set, you might ask? Great question. A level set of a function is a collection of all input points that give the same output.</p>

<p>Example: Consider the function $f(x,y) = x2 + y2$. Which set of inputs $(x,y)$ make $f(x,y) = 4$?</p>

<p>You have seen $x^2 + y^2 = 4$, and other equations like it. As shown in the graph below, where $f(x,y)$ is blue, when we set our function equal to a given value, we get the set of points that output that value. The actual level set exists on the x-y plane, but for the sake of visualization, I have included both the level set and the level set projected onto $f$’s surface.</p>

<figure style="margin: 1em 0; text-align: center;">
  <img src="/images/lagrange/lg_1.jpg" alt="Desmos function graph" style="width: 350px;" />
</figure>

<p>In this case, our function’s graph is a surface, and the level set is a curve.</p>

<p>Note: The surface is a 2D object in 3D, and the curve is a 1D object in 2D.</p>

<p>Going back to our original problem, $g$ is the function $g(x,y,z) = x + y - 2z$. Our specific constraint is the level set of this function where $g(x,y,z) = 0$. Rather than a curve, this gives us a surface.</p>

<p>We want to find the maximum output of $f(x,y,z)$ subject to this constraint. All points $(x,y,z)$ that reach that maximum output form a level surface of $f$, though we only care about points that also satisfy the constraint.</p>

<p>We solve our problem by finding the level sets of $f$ associated with the maximum value.</p>

<p>There are many possible level sets of $f$. For this problem, changing what the level set is equal to moves the surface up and down, but for other formulas, it can change the surface in different ways.</p>

<p>If we have a point of intersection between our level set and constraint surface that is not a tangent point between the two surfaces, we can move our level set while still having overlap, meaning there are solutions with higher and lower $f$ values. Therefore, this is not an optimized point. Only tangent and boundary points can be maximums or minimums.</p>

<p>We have successfully restructured the goal of our problem: we want to find level sets of $f$ that are tangent to our constraint surface.</p>

<p>How do we do that? We find points where the two gradients are parallel. If you don’t know what a gradient is, think of it as a higher-dimensional derivative. The gradient of a function is perpendicular to its level surface.</p>

<p>Therefore, if a level surface is tangent to the constraint surface, their gradients will be parallel:</p>

\[\nabla f = \lambda \nabla g\]

<p>When the gradients are parallel, we can multiply $\nabla g$ by an unknown scalar ($\lambda$, the Lagrange multiplier), making it equal to $\nabla f$.</p>

<p>To ensure our points lie on our constraint surface, we add the original constraint equation to our system. Every point p already lies on the level surface $f = f(p)$.</p>

<p>The points that satisfy this system of equations are on both surfaces, and we know those surfaces are parallel. Therefore, the surfaces must be tangent at these points.</p>

<p>To solve this problem, we calculate the gradients and set them equal to each other, multiplying $\nabla g$ by a Lagrange multiplier. Together with the constraint equation, this forms a system of equations. The solutions to this system of equations are potential candidates for the maximum or minimum points. We can use these to find true extrema.</p>

<p>That’s pretty boring, so I wrote some code to do it. (<a href="https://github.com/lukadamiank-oss/Lagrange-multipliers">GitHub</a>)</p>

<p>The maximum of our original problem is -1/8,  occurring at (1/4, 1/4, 1/4).</p>

<figure style="margin: 1em 0; text-align: center;">
  <img src="/images/lagrange/lg_2.jpg" alt="MPL function graph" style="width: 400px;" />
</figure>

<p>Let’s look at a different problem:</p>

<p>Find the maximum and minimum values of $f(x,y,z)=xyz$ subject to the constraint $x + 9y^2 + z^2 = 36$. Assume that $x ≥ 0$. (We must assume so, or there will be no absolute extrema)</p>

<p>For this problem:</p>

<p>We reach a maximum of 54 at points (18, -1, -3) and (18, 1, 3).</p>

<p>We reach a minimum of -54 at points (18, -1, 3) and (18, 1, -3).</p>

<p>Having multiple max or min points means that at the optimal level set, there happen to be multiple tangent points. Solution points shown below:</p>

<figure style="margin: 1em 0; text-align: center;">
  <img src="/images/lagrange/lg_3.jpg" alt="MPL function graph" style="width: 600px;" />
</figure>

<p>The two maximum tangent points in Desmos, for context. Cool.</p>

<figure style="margin: 1em 0; text-align: center;">
  <img src="/images/lagrange/lg_4.jpg" alt="Desmos function graph" style="width: 400px;" />
  <figcaption style="font-size: 0.85em; color: #666;">
    Blue: Max objective level surface, Red: constraint surface
  </figcaption>
</figure>

<p>These problems come up fairly often. For example, they are used in the constrained optimization problem behind support vector machines, a fundamental machine learning model: <a href="https://youtu.be/eHsErlPJWUU?si=AxLhSl70jF-Qk_HT&amp;t=1649">Caltech lecture</a>.</p>

<p>You can use Lagrange multipliers to solve optimization problems in 2D, 4D, or any number of dimensions. You can also solve problems with multiple constraints. The visualization only works for 3d, but if you want, you can use the calculator to mess around with higher-dimensional optimization problems. (Again: <a href="https://github.com/lukadamiank-oss/Lagrange-multipliers">GitHub link</a>)</p>]]></content><author><name></name></author><summary type="html"><![CDATA[Find the maximum value of]]></summary></entry><entry><title type="html">Geometric Determinants</title><link href="https://lukakirigin.com/blog/2026/06/05/Determs/" rel="alternate" type="text/html" title="Geometric Determinants" /><published>2026-06-05T00:00:00+00:00</published><updated>2026-06-05T00:00:00+00:00</updated><id>https://lukakirigin.com/blog/2026/06/05/Determs</id><content type="html" xml:base="https://lukakirigin.com/blog/2026/06/05/Determs/"><![CDATA[<p>What is a determinant? You may have used the operation before, but do you know what it means? If you do know what it means, do you know the explanation behind it? This article will discuss the true meaning behind determinants, framed around finding the volume of a parallelepiped constructed by three vectors. This article will assume a baseline level of linear algebra knowledge. However, the following preface will explain the key ideas used, so hopefully those interested will be able to follow.</p>

<p><strong>Preface:</strong></p>

<p>A single dimension can be visualized as a line. Our line has both negative and positive values, which indicate how far we are from the center. Adding a second line, perpendicular to our first, we create a plane. This is the second dimension, typically represented by the x-y plane. If we travel horizontally from our first line, in the direction of our second dimension, our point’s distance from the origin in the first dimension remains unchanged. A point is designated by an ordered pair (typically (x,y)), telling us the point’s position in the first and second dimension, respectively.</p>

<p>If we add a third line, going vertically through the intersection of our first two, we can operate in three dimensions. Our points are now represented by 3-tuples, (x,y,z).</p>

<figure style="margin: 1em 0; text-align: center;">
  <img src="/images/gomd/dimension_viz.jpg" alt="1st, 2nd, 3rd dimension" style="width: 600px;" />
  <figcaption style="font-size: 0.85em; color: #666;">
    Dimension visualization, Mario Estrada, ResearchGate
  </figcaption>
</figure>

<p>This is all fairly comprehensible. However, linear algebra does not limit itself to 3 dimensions. In fact, it allows us to work with objects in any number of dimensions.</p>

<p>The real space formed by n dimensions is referred to as $ℝ^n$.</p>

<p>Visualizing dimensions greater than three can be difficult. Having four or more perpendicular lines intersect at the same point is impossible within our three-dimensional world. However, even as we venture into higher dimensions, the geometry we know and love still holds. We can still calculate the angle between vectors, find the volume of shapes, and create subspaces.</p>

<p>In higher dimensions, the only difference is that our points contain more information. One helpful way to understand 4+ dimensions is to think of our dimensions as variables in a dataset. For example, in a financial model, you might have nine separate variables for each observation. If we want to measure how these variables relate to one another, we can represent the data in nine-dimensional space and use linear algebra to analyze it.</p>

<p>Throughout the course of this article, I will use examples in $ℝ^2$ and $ℝ^3$, which help with intuition, but are hard to visually extrapolate to higher dimensions.</p>

<p>The following objects are fundamental tools of linear algebra.</p>

<p>A vector is an arrow that encodes a direction and magnitude (length). They are typically represented by a point, where the vector is the arrow from the origin (0,0,…,0) to that point.</p>

<p>A matrix, A, represents a linear transformation, T, an operation we can perform on vectors that preserves vector addition and scaling.</p>

<p>For vectors v and u, and scalars a and b:</p>

\[T(a \mathbf v + b \mathbf u) = aT( \mathbf v) + bT(\mathbf u)\]

<p>I will give a brief explanation of what a matrix is, but for people who want a deeper understanding, watch <a href="https://www.youtube.com/watch?v=kYB8IZa5AuE">this</a> incredible video from 3blue1brown. By the way, his video on <a href="https://www.youtube.com/watch?v=MYAxVmbsF0k">determinants</a> serves as a nice introduction to what I will talk about, with far better visuals.</p>

<p>A matrix represents a <em>linear transformation</em>, a function of vectors that outputs exactly one vector for every vector input. They can act between dimensions, corresponding to their length and width (their output and input dimensions, respectively), but in this case, we will focus only on square matrices. A square matrix transforms the coordinate grid into a diagonal grid, preserving lines and linear combinations.</p>

<figure style="margin: 1em 0; text-align: center;">
  <img src="/images/gomd/diagonal-grid.webp" alt="Img of buildings" style="width: 600px;" />
  <figcaption style="font-size: 0.85em; color: #666;">
    Buildings with diagonal grid patterns, Research Gate
  </figcaption>
</figure>

<p>I’m going to assume people understand algebraically how to <a href="https://www.mathsisfun.com/algebra/matrix-multiplying.html">multiply matrices</a>.</p>

<p>Finally, the dot product (also called the inner product) is an operation we can perform on two vectors within the same dimension. It has two formulas: algebraic and geometric. Typically, we calculate the algebraic formula to derive geometric insight. To calculate the dot product algebraically, we sum the products of the vectors’ terms in each dimension. The geometric way is to multiply the magnitude (length) of the vectors together, and then multiply that by the cosine of the angle between them.</p>

\[\mathbf u \cdot \mathbf v = u_1 v_1 + u_2 v_2 + ... + u_1 v_1 = ||\mathbf u || || \mathbf v || \cos (\theta)\]

<p>The two most important facts about the dot product are: we can use it to calculate the projections between vectors, and the dot product of perpendicular vectors is zero.</p>

<p>These two facts are related, and if you don’t have intuition for this, take some time to consider how.</p>

<p><strong>Geometric Determinants:</strong></p>

<p>With our mathematical foundation laid, we begin our exploration of determinants with three vectors in 3d space ($ℝ^3$).</p>

<figure style="margin: 1em 0; text-align: center;">
  <img src="/images/gomd/pp_1.png" alt="Three vectors" style="width: 500px;" />
</figure>

<p>Using these vectors, we form a distorted rectangular prism (a parallelepiped).</p>

<figure style="margin: 1em 0; text-align: center;">
  <img src="/images/gomd/pp_2.png" alt="Parallelepiped" style="width: 700px;" />
  <figcaption style="font-size: 0.85em; color: #666;">
    Pallelepiped, Uregina
  </figcaption>
</figure>

<figure style="margin: 1em 0; text-align: center;">
  <img src="/images/gomd/lewitt.webp" alt="Parallelepiped, in art form" style="width: 300px;" />
  <figcaption style="font-size: 0.85em; color: #666;">
    Sol LeWitt, SF MOMA (with you)
  </figcaption>
</figure>

<p>If we wish to find the volume of this shape, we could calculate it using geometry. Alternatively, we could use something called a determinant.</p>

<p>The determinant is a number associated with a square matrix. Geometrically, it measures how much the matrix scales volume. We can calculate this number using an algebraic process, which I will explain momentarily, but first, let’s discuss its meaning.</p>

<p>The standard basis is a set of vectors that go one unit in each coordinate direction. In  $ℝ^3$, it is ([1,0,0], [0,1,0], [0,0,1]). People refer to these vectors as e1, e2, and e3. For higher dimensions, the standard basis is longer.</p>

<p>In $ℝ^3$, these vectors form a one-by-one-by-one cube, called the unit cube. In $ℝ^2$, they form what is called the unit square, and for higher dimensions, they form what is called the unit hypercube. These objects always have a volume of one. (area/hypervolume in other dimensions)</p>

<p>Consider a 3 by 3 matrix. (Also written as 3 x 3),</p>

\[\begin{bmatrix}
a &amp; b &amp; c \\
d &amp; e &amp; f \\
g &amp; h &amp; i
\end{bmatrix}\]

<p>If we apply this matrix to the first unit vector, we get the first column:</p>

\[\begin{bmatrix}
a &amp; b &amp; c \\
d &amp; e &amp; f \\
g &amp; h &amp; i
\end{bmatrix} \begin{bmatrix}
1 \\ 
0 \\
0
\end{bmatrix} = \begin{bmatrix}
a \\
d \\ 
g \end{bmatrix}\]

<p>If we apply it to the second unit vector, we get the second column:</p>

\[\begin{bmatrix}
a &amp; b &amp; c \\
d &amp; e &amp; f \\
g &amp; h &amp; i
\end{bmatrix} \begin{bmatrix}
0 \\ 
1 \\
0
\end{bmatrix} = \begin{bmatrix}
b \\
e \\ 
h \end{bmatrix}\]

<p>In fact, we can rewrite a matrix based on where it sends our standard basis. (Remember that $e_i$ refers to the ith standard basis vector)</p>

<p>If A is an n x n matrix representing the linear transformation T, then:</p>

\[A = \begin{bmatrix}
\vert &amp; \vert &amp;  &amp; \vert \\
T(e_1) &amp; T(e_2) &amp; ... &amp; T(e_n) \\
\vert &amp; \vert &amp; &amp; \vert 
\end{bmatrix}\]

<p>We can represent any vector as a linear combination of the standard basis. We can technically do the same with any basis, but the standard basis is simple.</p>

\[\begin{bmatrix}
3 \\
4
\end{bmatrix} = 3 \begin{bmatrix}
1 \\ 
0
\end{bmatrix}
+ 4\begin{bmatrix}
0\\
1
\end{bmatrix} = 3 e_1
+ 4 e_2\]

<p>Therefore, if we know how a matrix affects the standard basis, we can determine how it affects any vector.</p>

\[T(\begin{bmatrix}
3 \\
4
\end{bmatrix}) = 3 T(e_1) + 4T(e_2)\]

<p>Now, imagine that we want to measure the effect of a matrix on the unit cube itself. To do so, we could form a parallelepiped from the vectors that the elements of our standard basis were mapped to. (ie, our column vectors)</p>

<p><strong>The signed volume of this parallelepiped is the determinant of our matrix.</strong></p>

<p>So how does one calculate a determinant, and why does it give you this fascinating result?</p>

<p>The determinant of a 2x2 matrix A is calculated as follows: (vertical bars on a matrix mean that you are taking the determinant)</p>

\[A = \begin{bmatrix}
    a &amp; b \\
    c &amp; d
\end{bmatrix} \hspace{1cm}
det(A) = \begin{vmatrix}
    a &amp; b \\
    c &amp; d
\end{vmatrix} = (ad - bc)\]

<p>For larger matrices, we start by choosing any row or column. For each value in this set, we multiply it by the determinant of our matrix excluding both the row and column of our value. Finally, we sum these terms together with alternating signs based on the pattern below, which extends correspondingly to larger matrices.</p>

<figure style="margin: 1em 0; text-align: center;">
  <img src="/images/gomd/plusminus.webp" alt="aternating plus-minus grid" style="width: 200px;" />
</figure>

<p>Determinant of a 3x3 matrix row expansion calculation example:</p>

\[\{\mathbf v_1, \mathbf v_2, \mathbf v_3\} = \{\begin{bmatrix}
1\\
2\\
0
\end{bmatrix}, \begin{bmatrix}
1\\
1\\
1
\end{bmatrix}, \begin{bmatrix}
0 \\
1 \\
2
\end{bmatrix}
\}
\hspace{1cm}
A = \begin{bmatrix}
1 &amp; 1 &amp; 0 \\
2 &amp; 1 &amp; 1 \\
0 &amp; 1 &amp; 2
\end{bmatrix}\]

\[\det(A) = \begin{vmatrix}
1 &amp; 1 &amp; 0 \\
2 &amp; 1 &amp; 1 \\
0 &amp; 1 &amp; 2
\end{vmatrix} = 1 \begin{vmatrix}
    1 &amp; 1 \\ 
    1 &amp; 2
\end{vmatrix} - 1 \begin{vmatrix}
    2 &amp; 1 \\ 
    0 &amp; 2
\end{vmatrix} + 0 \begin{vmatrix}
    2 &amp; 1 \\ 
    0 &amp; 1
\end{vmatrix}\]

\[=1(1*2 - 1*1) - 1(2*2 - 1*0) + 0= (2-1) - (4) = -3\]

<p>This tells us that the volume of the parallelepiped formed by [1,2,0], [1,1,1], and [0,1,2] is -3. The volume being negative means that at some point, our vectors switched order. This is not insignificant, but for finding the volume, we take the absolute value.</p>

<p>One important aspect of the determinant is that it is multiplicative. That is:</p>

<p>det(AB) = det(A)det(B).</p>

<p>Knowing that the determinant measures scaled volume helps with the intuition here.</p>

<p>Now that we have introduced the determinant, we can move on to the real question. Why does it give you the volume of the transformed unit cube?</p>

<p>For two dimensions, you could verify that the formula matches fairly easily. For three dimensions, you could do this as well, albeit with more calculation. However, doing this for higher dimensions would be tedious and unsatisfying.</p>

<p>For a cleaner proof, we start with the determinant of an n x n matrix.</p>

<p>Let us discuss column addition, which is when you add scaled columns to each other. This action does not change the determinant.</p>

<p>Example:</p>

\[\begin{vmatrix}
3 &amp; 1 &amp; 0 \\ 
4 &amp; 2 &amp; 0 \\
2 &amp; 0 &amp; 1 
\end{vmatrix} = \begin{vmatrix}
(3 - 2)&amp; 1 &amp; 0 \\ 
(4-4) &amp; 2 &amp; 0 \\
(2-0) &amp; 0 &amp; 1 
\end{vmatrix} = \begin{vmatrix}
1 &amp; 1 &amp; 0 \\ 
0 &amp; 2 &amp; 0 \\
2 &amp; 0 &amp; 1 
\end{vmatrix}\]

<p>Here, we subtracted two times our second column from our first column.</p>

<p>To understand why the determinant remains unchanged here, we should think about it in terms of what the determinant is geometrically calculating. That is, the area of the parallelotope formed by our column vectors. As a mathematician, it pains me to use the goal of a proof in the explanation. However, this part is more geometric intuition than a strict proof, and it’s the easiest way to understand it.</p>

<p>(‘Parallelotope’ is the n-dimensional name for the shape of the same family that parallelograms and parallelepipeds belong to.)</p>

<p>Column addition doesn’t change the volume of our parallelotope because it is a <em>shear</em>.</p>

<figure style="margin: 1em 0; text-align: center;">
  <img src="/images/gomd/shear.webp" alt="shear visualization" style="width: 600px;" />

  <figcaption style="font-size: 0.85em; color: #666;">
    Shear, MathWorks
  </figcaption>
</figure>

<p>A shear is an affine transformation that moves all points in a fixed direction based on their distance to a specified line parallel to that direction.</p>

<p>Think of a deck of cards that you nudge at an angle, moving the top ones more than the bottom ones. The important idea: shears do not change volume.</p>

<p>Take vectors [3,0] and [1,2], which form the parallelogram below.</p>

<figure style="margin: 1em 0; text-align: center;">
  <img src="/images/gomd/pp_3.webp" alt="aternating plus-minus grid" style="width: 500px;" />
</figure>

<p>If we subtract from [1,2] a scaled version of [3,0], the top of our parallelogram would remain on the line y = 2. All possible resulting shapes have equal volume.</p>

<video controls="" autoplay="" loop="" muted="" playsinline="" preload="metadata" width="100%">
    <source src="/videos/pp_4_vid.mp4" type="video/mp4" />
</video>

<p>You could imagine that at some point, our vectors will be perpendicular to each other. At that point, if we want to calculate the area, instead of needing to measure the height of our parallelogram, we can simply multiply the lengths of our vectors together.</p>

<p>With this example, it is easier to adjust [1,2] to be perpendicular to [3,0], as [3,0] lies on an axis. However, we could also go the other way, subtracting scaled versions of [1,2] from [3,0]. With this process, we could also make the vectors perpendicular.</p>

<p>The next step involves generalizing this process to higher dimensions. Thinking about three dimensions, we start with a parallelepiped. As explained, this parallelepiped is formed with three vectors. From these three vectors, we choose any two to act as our base. These two vectors form a parallelogram. Referring to the parallelepiped diagram, you can see that we have three choices for this base.</p>

<p>The goal of this process is to make our vectors perpendicular to each other, without sacrificing volume.</p>

<p>To do so, we start by performing the parallelogram example we just looked at on the two vectors. We make the vectors in the base perpendicular to each other with a shear, subtracting a scaled version of one of them from the other.</p>

<p>Now, we have two perpendicular vectors forming our base, and a final third vector. If we subtract scaled versions of our two base vectors from this third vector, the top moves, while the volume remains constant.</p>

<p>You can imagine that, at some combination of our subtraction, we end up with a third vector perpendicular to both of the vectors that form our base. This set of three perpendicular vectors forms a rectangular prism, with the same volume as our original parallelepiped. This process is visualized below. (<a href="https://www.desmos.com/3d/afc0uudkdn">Link</a> to Desmos graph)</p>

<video controls="" autoplay="" loop="" muted="" playsinline="" preload="metadata" width="100%">
    <source src="/videos/parallelepiped-viz.mp4" type="video/mp4" />
</video>

<p>This process of making our vectors perpendicular can be generalized to n dimensions, and has a name you may have heard before: the Gram-Schmidt process.</p>

<p>Before we generalize this process, it’s important to understand the method of determining how much of each vector to shear when making them perpendicular. We subtract from each vector the amount it travels in the direction of the other vectors. This is a projection, which we can calculate using the dot product. If you want a full breakdown of the Gram-Schmidt process, watch <a href="https://www.youtube.com/watch?v=enWNFOfudYw">this</a> video from James Hamblin.</p>

<p>For n dimensions, we start by picking a single vector. We then pick a second vector and make it perpendicular to our first. The base at this point is a line, and together the vectors form a rectangle. Next, we pick a third and make it perpendicular to our first two vectors. The base at this point is a rectangle, and together our vectors form a rectangular prism.</p>

<p>For the fourth vector, we do the same thing. We subtract its projection onto our three base vectors, leaving it perpendicular to our base, a parallelepiped. We know this to be true, as if we take the dot product between it and any of our base vectors, we get zero. However, it is essentially impossible to visualize this, as what I’m saying is that we have a vector perpendicular to a three-dimensional object. With that being said, the algebra works; a vector in $ℝ^4$ can be perpendicular to an entire 3-dimensional subspace. One must accept that higher dimensions don’t comply with our understanding of the universe, because we only exist in three.</p>

<p>Moving past that enigmatic fact, we can continue to take new vectors and perpendicularize them until we run out of vectors. This leaves us with an orthogonal set of vectors that form a hyperrectangle with hypervolume equal to our original parallelotope.</p>

<p>Lots of big words. Basically, we make the vectors perpendicular, and now we have a rectangle instead of a parallelogram.</p>

<p>Once we have this new set of vectors, finding the (hyper)volume of our shape is trivial. Because of its rectangular structure, we can determine the volume by multiplying the side lengths of our shape together. In conclusion, we make our vectors perpendicular by subtracting scaled versions of each other, then multiply their magnitudes together to get the volume of our original parallelotope.</p>

<p>Going all the way back to the determinant, we embarked on this whole discussion to prove that we can perform column addition without affecting our determinant’s value. We did so, but also described how we can use column addition to get perpendicular column vectors.</p>

<p>Let’s set up an arbitrary n by n matrix, which we will work with for the remainder of this article. The format below is a way of writing matrices, with $v_1$, $v_2$, etc, as the column vectors.</p>

\[A = 
\begin{bmatrix}
    \vert &amp; \vert &amp;  &amp; \vert \\
    \mathbf v_1   &amp; \mathbf v_2   &amp; ... &amp; \mathbf v_n\\
    \vert &amp; \vert &amp; &amp; \vert
\end{bmatrix}\]

<p>Perform the Gram-Schmidt process on this matrix’s column vectors, and form another matrix:</p>

\[U = 
\begin{bmatrix}
    \vert &amp; \vert &amp;  &amp; \vert \\
    \mathbf u_1   &amp; \mathbf u_2   &amp; ... &amp; \mathbf u_n\\
    \vert &amp; \vert &amp; &amp; \vert
\end{bmatrix}\]

<p>As I mentioned, this process uses only column addition, and therefore, the determinants of these two matrices will be the same.</p>

<p>What I will now prove is that the determinant of our matrix U is equal to the magnitudes of its column vectors multiplied together. As we have shown, this product is equal to the volume of the parallelepiped formed by the column vectors of matrix A.</p>

<p>First, we form a new matrix G equal to U transpose times U. The transpose of a matrix is simply the same matrix flipped along its diagonal. This matrix is called a Gram matrix. When we create this matrix, we get the following structure:</p>

\[G = 
\begin{bmatrix}
    \text{---} &amp; \mathbf u_1 &amp; \text{---} \\
    \text{---} &amp; \mathbf u_2 &amp; \text{---} \\ 
 &amp; ... &amp; \\
\text{---} &amp; \mathbf u_n &amp; \text{---}
\end{bmatrix}
\begin{bmatrix}
    \vert &amp; \vert &amp;  &amp; \vert \\
    \mathbf u_1   &amp; \mathbf u_2   &amp; ... &amp; \mathbf u_n\\
    \vert &amp; \vert &amp; &amp; \vert
\end{bmatrix}\]

\[G = \begin{bmatrix}
    \mathbf u_1\cdot \mathbf u_1 &amp; \mathbf u_1 \cdot \mathbf u_2 &amp; ... &amp; \mathbf u_1 \cdot \mathbf u_n \\
    \mathbf u_2 \cdot \mathbf u_1 &amp; \mathbf u_2 \cdot \mathbf u_2 &amp; ... &amp; \mathbf u_2 \cdot \mathbf u_n \\
    ... &amp; ... &amp; ... &amp; ... \\ 
    \mathbf u_n \cdot \mathbf u_1 &amp; \mathbf u_n \cdot \mathbf u_2 &amp; ... &amp; \mathbf u_n \cdot \mathbf u_n
\end{bmatrix}\]

<p>If you are confused here, try doing the matrix multiplication yourself.</p>

<p>Because the column vectors of our U matrix are perpendicular to each other, when you take the dot product of two distinct U vectors, you get zero.</p>

\[G = \begin{bmatrix}
    \mathbf u_1 \cdot \mathbf u_1 &amp; 0 &amp; ... &amp; 0 \\
    0 &amp; \mathbf u_2 \cdot \mathbf u_2 &amp; ... &amp; 0 \\
    ... &amp; ... &amp; ... &amp; ... \\ 
    0 &amp; 0 &amp; ... &amp; \mathbf u_n \cdot \mathbf u_n
\end{bmatrix}\]

<p>The magnitude of a vector is its Euclidean length, calculated using the following formula:</p>

\[\mathbf u = \begin{bmatrix}
u_1 \\
u_2 \\
... \\ 
u_n
\end{bmatrix}
\hspace{1cm}
|| \mathbf u || = 
\sqrt{u_1 ^ 2 + u_2  ^ 2 + ... + u_n^2}\]

<p>When we take the dot product of a vector with itself, we get its magnitude squared.</p>

\[\mathbf u \cdot \mathbf u = u_1^2 + u_2^2 + ... + u_n^2 = ||\mathbf u || ^2\]

<p>We can rewrite our Gram matrix with this in mind.</p>

\[G = \begin{bmatrix}
    ||\mathbf u_1||^2 &amp; 0 &amp; ... &amp; 0 \\
    0 &amp; ||\mathbf u_2||^2 &amp; ... &amp; 0 \\
    ... &amp; ... &amp; ... &amp; ... \\ 
    0 &amp; 0 &amp; ... &amp; ||\mathbf u_n||^2
\end{bmatrix}\]

<p>Keep in mind that the magnitudes of these vectors are the side lengths we use to calculate the volume of our parallelepiped.</p>

<p>Finally, we take the determinant of G. Because it is diagonal, we get the product of its elements.</p>

\[\det(G) = ||\mathbf u_1 || ^2  ||\mathbf u_2||^2 ...  ||\mathbf u_n||^2 = (||\mathbf u_1 ||  ||\mathbf u_2|| ...  ||\mathbf u_n||)(||\mathbf u_1 || ||\mathbf u_2||...  ||\mathbf u_n||)\]

<p>Now, rewrite the determinant of G in terms of U. This step uses two rules about the determinant: that it is multiplicative, and the determinant of a matrix’s transpose is equal to the determinant of the matrix.</p>

\[\det(G) = \det(U^T U) = \det(U^T)\det(U) = \det(U)^2\]

<p>Next, we set our two equations equal to each other.</p>

\[\det(U)^2 = (||\mathbf u_1 ||  ||\mathbf u_2|| ...  ||\mathbf u_n||)(||\mathbf u_1 ||  ||\mathbf u_2|| ...  ||\mathbf u_n||)\]

<p>Take the square root of both sides:</p>

\[\lvert \det(U) \rvert =  ||\mathbf u_1 ||  ||\mathbf u_2|| ...  ||\mathbf u_n||\]

<p>Remember that the determinant of U is equal to the determinant of our original matrix, A.</p>

<p>And with that, we are done! The absolute value of the determinant of a square matrix is equal to the product of the magnitudes of the orthogonal vectors produced by applying Gram-Schmidt to its column vectors.</p>

<p>I hope this made sense. I think it’s very interesting how these ideas, which are taught in linear algebra, are really all connected, but we never learn about them that way. Thank you for reading, and I will catch you in the next edition of this unnamed math blog.</p>]]></content><author><name></name></author><summary type="html"><![CDATA[What is a determinant? You may have used the operation before, but do you know what it means? If you do know what it means, do you know the explanation behind it? This article will discuss the true meaning behind determinants, framed around finding the volume of a parallelepiped constructed by three vectors. This article will assume a baseline level of linear algebra knowledge. However, the following preface will explain the key ideas used, so hopefully those interested will be able to follow.]]></summary></entry><entry><title type="html">Arrow’s Theorem of Impossibility: A Dialogical Analysis</title><link href="https://lukakirigin.com/blog/2026/01/02/Arrow/" rel="alternate" type="text/html" title="Arrow’s Theorem of Impossibility: A Dialogical Analysis" /><published>2026-01-02T00:00:00+00:00</published><updated>2026-01-02T00:00:00+00:00</updated><id>https://lukakirigin.com/blog/2026/01/02/Arrow</id><content type="html" xml:base="https://lukakirigin.com/blog/2026/01/02/Arrow/"><![CDATA[<p>Jean Carlé: Kenneth Arrow’s Theorem of Impossibility tells us that if we wish to rank a set of outcomes based on a population’s preferences, ranked preference voting will not allow us to determine an optimal solution. In this lecture, I will explain his proof for this theorem and how we may uncover the paradox it outlines. Then I will discuss its context, specifically the implications of how we uncover our paradox, and why it causes many people to cite its conclusion misguidedly. Where shall we start?</p>

<p>Tay Staviski: We must start by establishing our assumptions. I would formalize our function, as well as the conditions Arrow requires for an ‘optimal’ ranking. Then we can create our optimization algorithm.</p>

<p>Jean Carlé: We won’t be creating the social welfare function, just assuming one exists, though we do need to define its structure. Our inputs are ranked lists of potential outcomes. We take an arbitrary number of outcomes, defined by a group X, and an arbitrary number of voters, defined by a group N. Each member of group N submits an ordinal ranking of all elements in group X. Our social welfare function takes these rankings and outputs an optimal ranking of all outcomes. This ranking most accurately matches the aggregate preferences of group N. Arrow’s Theorem walks through a proof showing that an algorithm of this sort cannot violate the conditions he sets up. To understand how it works properly, we must have a complete understanding of these conditions. May you summarize them for us?</p>

<p>Tay Staviski: The social welfare function must give an output for any set of preferences. Second, this output must reflect all unanimous preferences of our population. Third, the ranking of two variables cannot be influenced by the presence of a third variable. Finally, the decision cannot be determined by the preferences of a single voter.</p>

<p>Jean Carlé: These conditions are respectively called: unrestricted domain, Pareto efficiency, the independence of irrelevant alternatives, and non-dictatorship. Besides these, we assume anonymity, meaning each input is considered equally. Going one by one, I will explain each condition.</p>

<p>Non-dictatorship: This condition states that a sole ‘dictator’ cannot determine the output. A voter acts as a dictator if, for every pair of outcomes, the output reflects the voter’s ranking, no matter what other voters prefer. This condition is assumed with the anonymity of our voters, although Arrow explicitly states non-dictatorship as a factor in his original theorem.</p>

<p>Unrestricted domain: Voters must be able to submit any permutation of outcomes. This ensures our social welfare function actually meets the requirements to be classified as a function. Every possible set of finite voter preferences will be mapped to an output.</p>

<p>Pareto efficiency: Our function’s output must be Pareto efficient. This means that every possible change we could make to our final output will, even if it increases the well-being of one person, decrease the well-being of someone else. In other words, our final output cannot be altered to benefit everyone more. This also means that if our population unanimously prefers an outcome X over Y, that preference must be reflected in our output.</p>

<p>Independence of irrelevant alternatives (IIA): The ranked order between two outcomes in our final output cannot be influenced by the addition or subtraction of another outcome. Theoretically, a third option should not influence a voter’s decision between two outcomes. This condition is necessary for our proof to work.</p>

<p>These conditions are necessary for a fair, rational-choice voting system.</p>

<p>Tay Staviski: If the benefit of some change outweighs the detriment, then should we not still enact the change? Why would we not use a system in which participants assign cardinal variables to each ranking? We could then measure the total well-being for each possible ranking of outcomes, and define an algorithm to determine the exact net effect of changes. The use of ordinal variables significantly reduces our analytical capabilities, as our algorithm cannot account for opinions held by the population, though not unanimously.</p>

<p>Jean Carlé: Pareto efficiency isn’t how we determine our output; it is just a condition that it must fulfill. Our final result does not have to be unanimous, but it must reflect all unanimous preferences our population holds. As to your point of cardinal variables, you are correct that they would allow us to always find the exact result that perfectly optimizes social welfare. However, we run into two key issues. First, cardinal variables aren’t trustworthy for precise measurements of population opinions. If we are using a scale from 1-10, then an outcome that benefits two voters the same amount might be reported as an 8 for one voter and a 10 for another. This discrepancy would happen constantly and would be impossible to account for. Second, take two options: option X being very favorable and option Y being unfavorable. A voter might report 9 for X and 2 for Y, or some other set of similar numbers. Now, imagine a new scenario with the same voter, but options X, Y, and a new, extremely favorable option, Z. The presence of Z does not change the fact that X &gt; Y. This is why IIA holds with ordinal variables. However, if our voter wants to report a 9 for Z, they will most likely lower their number for X. This changes the relative ranking between X and Y. Because cardinal variables mean the introduction of a third outcome might change a voter’s evaluation of two original outcomes, we are forced to use ranked-choice voting. Let us now move on to Arrow’s proof itself.</p>

<p>Arrow’s Impossibility Theorem:</p>

<p>Imagine a scenario with a finite population of size N, and a finite set of outcomes. Among these are two arbitrary and distinct outcomes, X and Y. Assume that we have a unanimous ranking of these two; all participants rank X &gt; Y. We go one by one, and flip each voter’s ranking to Y &gt; X. When all voters rank Y &gt; X, society will as well. Therefore, at some point in this process, we will encounter a pivotal voter, such that flipping their vote will cause society to no longer rank X &gt; Y. We label this voter nxy. When nxy ranks Y &gt; X, we say society ranks Y ≥ X, not the stricter Y &gt; X. This is because it is possible for Y to be equal to X. Similarly, when changing society’s vote from Y &gt; X to X &gt; Y, we can find a voter nyx, who, when ranking X &gt; Y, dictates that society ranks X ≥ Y. With this in mind, we can imagine two possible ways a population N can vote. Label the first S1, and the second S2. We start with a unanimous ranking of X &gt; Y in both, then sequentially flip voters’ opinions to Y &gt; X. In the first group, S1, every voter up until nxy prefers Y &gt; X, while nxy and every subsequent voter prefers X &gt; Y. S2 is the same except for nxy, who prefers Y &gt; X. Therefore, societal ranking is X &gt; Y in S1, and Y ≥ X in S2.</p>

<p>Now we introduce a third outcome, Z, which does not change the relative ranking of X and Y in any way. In S1, all voters place Z directly below Y. By universality, S1 ranks Y &gt; Z. Because X &gt; Y, the societal ranking of S1 is X &gt; Y &gt; Z. Therefore, through transitive property, S1 ranks X &gt; Z even though the vote is not unanimous. In both societies, nxy ranks Z last, below both X and Y. In S2, all voters maintain their relative ranking of X and Z, meaning we keep the societal ranking of X &gt; Z due to IIA. The only change between S1 and S2 among the pairs (X, Y) and (X, Z) is nxy’s pivotal flip. Instead of X &gt; Y &gt; Z (S1), we have Y ≥ X &gt; Z (S2). However, in S2, each voter’s ranking of Y and Z is undefined. This means that each voter could place Z above, equal to, or below Y. This is problematic because, based on our conclusion of Y ≥ X &gt; Z, Y should be above Z. Even though the individual relationship between the outcomes is undefined for each voter, meaning society could rank Z &gt; Y, S2 can never do so. Therefore, with our current setup, Y must be ranked above Z. The only thing that is actually known about the way society ranks the variables is that nxy ranks Y &gt; Z. Next, because IIA says society ranks a pair of outcomes based solely on the way voters rank that pair, we are forced to conclude that when nxy ranks Y &gt; Z, society must rank Y &gt; Z. This is extremely unintuitive, but by the conditions posited, is true.</p>

<p>Let’s imagine all voters in S2 rank Y &gt; Z. We sequentially flip each voter’s ranking to Z &gt; Y in the same way we found nxy. Until we flip nxy, they rank Y &gt; Z, meaning society also ranks Y &gt; Z. Therefore, we must encounter the voter whose ranking of Z &gt; Y invalidates Y &gt; Z after we flip nxy, because while nxy still ranks Y &gt; Z, society can’t change. In other words, nyz ≥ nxy. Next, we repeat this process to find nzy. All voters in S2 rank Z &gt; Y, so society ranks Z &gt; Y. We know that once we flip nxy, society will no longer rank Z &gt; Y. Therefore, nzy must be before or equal to nxy, or nxy ≥ nzy. Putting our two inequalities together, we get: nyz ≥ nxy ≥ nzy. With this, we are almost complete. Remember that our variables are arbitrary and distinct outcomes. Because of this, when we determine a conclusion about them, we would be able to swap their initial positions and therefore know that what we concluded works with a different order of the variables. Think of this not as setting up a new system, but just changing the labels on the same outcomes. Repeating our recent conclusion with Y and Z swapped, we get nzy ≥ nxy ≥ nyz. Using these inequalities together:</p>

<p>nyz ≥ nxy ≥ nzy and nyz ≤ nxy ≤ nzy</p>

<p>We can conclude:</p>

<p>nyz = nxy = nzy</p>

<p>Tay Staviski: Fascinating. At this point, we have proven that a decisive voter for any two outcomes is also a decisive voter for either one of those outcomes and an arbitrary third variable. This clearly extends to a voting pool with any number of outcomes, as repeating this process is trivial. It really is interesting that the whole proof relies on the fact that S2 always conforms to nxy when they rank Y &gt; Z. Especially when that lemma seems to be based on logic, which requires a suspension of belief.</p>

<p>Jean Carlé: This point demonstrates how core Arrow’s axioms are to the proof. They seem so reasonable when introduced that one would assume most voting systems fulfill them. In reality, they are seldom realized entirely. Because this proof’s logic stretches its conditions to their absolute limit, any margin for error on a factor such as IIA would invalidate the proof entirely. In modern politics, people use this proof as an argument against ranked-choice voting systems, when there is little to no chance that even small elections could be done while ensuring IIA completely. Arrow’s theory reminds us to heavily prepare before attempting to prove something. It is a foundational piece of math, and its limited real-world application should not diminish its respect. It is a proof which has been re-written, taken out of context, and misused possibly more than any other in history, and will continue to be used in this way. Understanding its true mechanisms gives a slight shock, and at the very least, allows one to skepticize people quoting it in ways you now know to be invalid. Thank you for listening.</p>]]></content><author><name></name></author><summary type="html"><![CDATA[Jean Carlé: Kenneth Arrow’s Theorem of Impossibility tells us that if we wish to rank a set of outcomes based on a population’s preferences, ranked preference voting will not allow us to determine an optimal solution. In this lecture, I will explain his proof for this theorem and how we may uncover the paradox it outlines. Then I will discuss its context, specifically the implications of how we uncover our paradox, and why it causes many people to cite its conclusion misguidedly. Where shall we start?]]></summary></entry><entry><title type="html">The Library of Babel; Embedding Information</title><link href="https://lukakirigin.com/blog/2025/10/24/embeddings/" rel="alternate" type="text/html" title="The Library of Babel; Embedding Information" /><published>2025-10-24T00:00:00+00:00</published><updated>2025-10-24T00:00:00+00:00</updated><id>https://lukakirigin.com/blog/2025/10/24/embeddings</id><content type="html" xml:base="https://lukakirigin.com/blog/2025/10/24/embeddings/"><![CDATA[<p>The Library of Babel, imagined by writer and poet Jorge Luis Borges in his short story of the same name, is a theoretical, infinitely large library containing all possible information, constructed of symbols both familar and unknown. This library would contain the sum total of all human knowledge, as well as all future works which could be created. However, these peices would be lost in a sea of gibberish.</p>

<p>How could we define this library from a mathematical standpoint? The simlplest case, though we would lose information, is that of a standard alphabet and size. Let’s say we have 32 characters, each letter in the English alphabet, along with some basic punctuation symbols: (, . : ? ; space). We’re skipping quite a lot here, such as all of mathematics. However, with this set, it’s not unreasonable to claim that most information is conveyable.</p>

<p>In the original story, each book has a standard template: 410 pages, 40 lines on each page, and 80 characters in each line. We can imagine a library that contains all possible combinations of characters needed to fill this book. This library would be unimaginably massive, with the number of possible books having over 2 million digits, assuming there are no repetitions. Given, most of these would consist of nothing but random letters. However, it would also contain exact copies of all books written with less than 1.6 million characters, and variations of each with one-character differences. It contains all possible iterations of ASCII art, and therefore copies of every painting made with symbols.</p>

<p>Let’s abstract this a bit, stepping away from the story itself. It would be better not to have a limit on the number of characters we can put, so let’s do just that. We still have constraints, namely the fact that while the books themselves can be any size, they still must be finite. This new library would be truly infinite, but would still be countable, due to our defined number of characters. To create a correspondence between this library and the natural numbers, we could assign an order to our characters, then order our books as such:</p>

<p>1,1,1,1… 1</p>

<p>1,1,1,1… 2</p>

<p>…</p>

<p>1,1,1,1… 32</p>

<p>1,1,1,1… 2, 1</p>

<p>1,1,1,1… 2, 2</p>

<p>…</p>

<p>1,1,1,1… 2, 32</p>

<p>1,1,1,1… 3, 1</p>

<p>…</p>

<p>1,1,1,1… 32, 32</p>

<p>1,1,1,1… 2, 1, 1</p>

<p>…</p>

<p>32, 32, 32… 32</p>

<p>This is similar to counting in base 32. With this, we can establish a correspondence with the natural numbers, meaning the number of books is countably infinite.</p>

<p>So how would we organize the library when accepting books of any alphabet? We could use Paul Dancstep’s presentation of the library. For each combination of a number of characters and a number of spaces we can put them in, we create a collection of possible books.</p>

<p>We can organize these collections in a grid pattern, with the block closest to the origin using an alphabet with 1 character and having 1 space for it to go. One axis is the number of characters, and the other is the number of spaces. This grid is filled with blocks, which are collections of the possible books that can be in them. We can create a clear mapping between the structure of these categories and the x-y plane, and therefore, the number of categories is countable. Additionally, each collection will contain a finite number of books, so we can conclude that this new library is still countable.</p>

<p>One thing which we have overlooked is the order of characters, which is necessary for us to order the books within these categories. We can order each collection with the same process we used to count the books earlier, allowing us to call books directly. We could imagine a distinct vector that allows us to find the exact book we want with three distinct variables: alphabet length, book length, and the index within each collection. This is a useful way to index our library, though it is limited. To perform any analysis of a book, we would have to return to our model and find the structure of the book we are referencing. Also, we would be unable to perform operations between the vectors themselves, as they only act as an index. For the purpose of this essay, we want to try and find something generalizable to linear algebra, in which we have a very established framework for the interaction between vectors. While our current model does index books as vectors, we can’t perform actions such as matrix multiplication or addition. Let’s try another approach which uses linear algebra.</p>

<p>Before we do so, we have to slightly shift our focus. Exploring our infinite library has led us to a deeper task: quantifying langauge. We must break down the books into paragraphs, sentences, words, and even morphemes (pre, anti, etc.). Therefore, when moving forward, we will keep the same structure of the library in mind, but instead of thinking of books in their entirety, we will think of short words with much smaller amounts of characters. Another point: while we wish we could use every language format, there are different meanings for the same symbols across languages. Therefore, it will be extremely difficult to derive meaning unless we decide on a specific language. We can even generalize our format so it works with any language, but once we choose one, we can’t switch interchangeably. For now we may stick to english.</p>

<p>Instead of trying to fit an index of the books into a vector, we can build the system from the basis. We will avoid the generalization of languages for now and focus on defining a vector space for each possible alphabet. This way, we can represent words as vectors, and in this vector space perform interactions between them, allowing us to interpret them mathematically. For example, say we have a collection of books with 5 possible characters. We can establish a basis from these characters as such:</p>

<p>[1, 0, 0, 0, 0] for A</p>

<p>[0, 1, 0, 0 ,0] for B</p>

<p>[0, 0, 1, 0, 0] for C</p>

<p>[0, 0, 0, 1, 0] for D</p>

<p>[0, 0, 0, 0 ,1] for L</p>

<p>We can generalize this format to languages with any number of characters, simply adding more bases. Now, when making a book, sentence, or word, we can put together these rows in a matrix. For example, CABAL would look like such:</p>

<p>[0, 0, 1, 0, 0]</p>

<p>[1, 0, 0, 0, 0]</p>

<p>[0, 1, 0, 0, 0]</p>

<p>[1, 0, 0, 0, 0]</p>

<p>[0, 0, 0, 0, 1]</p>

<p>This is interesting, and now we have strings represented in a matrix, though when putting them together we might need to add a space into our set of characters. However, this is not that practical, especially if we want to represent something as long as a full book. As of now, we have little more than a simple cipher. But this structure is the building block of more complicated mathematical representations of language. Let’s discuss embedding.</p>

<p>One of the most prevalent cases of language-to-math conversion is AI and the study of natural language processing. However, they don’t observe texts using this one-hot vector form. This form allows us to represent each character as a vector, but it is rigid and mostly uninformative, not telling us anything about how one character relates to another. Therefore, an AI encodes words using an embedding matrix. An embedding matrix defines, for each of our characters, a specific number that determines the placement of that character. It might be easier to work with an example:</p>

<p><img src="/images/embd/embeddings_photo1.png" alt="Embeddings Example" /></p>

<p>General form:</p>

<p>[a1, a2, a3, a4, a5]</p>

<p>[b1, b2, b3, b4, b5]</p>

<p>[c1, c2, c3, c4, c5]</p>

<p>[d1, d2, d3, d4, d5]</p>

<p>[l1, l2, l3, l4, l5]</p>

<p>Our word cabal might look like:</p>

\[\begin{bmatrix}
c_1 &amp; c_2 &amp; c_3 &amp; c_4 &amp; c_5 \\
a_1 &amp; a_2 &amp; a_3 &amp; a_4 &amp; a_5 \\
b_1 &amp; b_2 &amp; b_3 &amp; b_4 &amp; b_5 \\
a_6 &amp; a_7 &amp; a_8 &amp; a_9 &amp; a_{10} \\
l_1 &amp; l_2 &amp; l_3 &amp; l_4 &amp; l_5
\end{bmatrix}\]

<p>Each row is a 5-dimensional embedding for the corresponding character. We can use this to express a word as a matrix of embedded vectors. One important caveat here is that while this form is usually discussed as a matrix, it is more akin to a list of vectors for each letter, which when combined act as a matrix. When referencing a word, AI needs not only the letters comprising it, but also the context surrounding their usage. In this form, the rows determine the letter order, and we still have the other information about weights.</p>

<p>These weights are not random, they are chosen by AI systems to represent the use of those characters in the position which each weight refers to. The weights are generally not readable by humans, or at least we have little to no intuition surrounding them. AI models are trained to create these weights. To understand this further, we need to break down the way AI models work, and how they are trained.</p>

<p>AI is trained by human input in the form of pre-labeled sentences which contain meaning. Sometimes, this meaning is made explicit, as would be the case with early stage development, using small sets of sorted data. However, it could also assume this implicitly, such as when training on huge sets of data such as the internet. Without this training, an AI has no way to distinguish real words from random combinations of letters.</p>

<p>This system is different from the way humans learn, but there are many parallels to be drawn, and without training of our own, humans are similarly incapable of deciphering meaning. If a human were to walk around the Library of Babel, and pick books to read at random, almost every time they would encounter nothing but random strings of letters. However, If a book you pick up happens to be a real story written in English, I, or someone similar, would recognize it as containing meaning. This decision would be based on prior knowledge and understanding of the English language. If there was a book written in a language you had no knowledge of, you would have a very difficult time differentiating it from something meaningless, or random letters. In the same way, a human who somehow had no knowledge of written language in any form would have no idea what anything in the library means, including books written in English.</p>

<p>Before training, AI acts similarly to someone like this, with no knowledge of how sentences are put together. When it is trained with human labeled data, it recalibrates its weights to fit the real examples of meaning. For example, an AI might receive a sentence such as</p>

<p>“The cat slept on the ___.”</p>

<p>It would then attempt to guess what the last word might be. To do this, it would use the weights and determine what it believes is most likely to conclude the sentence, based on past data it has seen. Let’s say the AI decided on “roof”, though the actual answer was “mat”. It would then compare the response it gave, calculating the dot product between its proposed answer and the correct answer. Then, it changes its weights, altering its prediction matrices to incorporate the error. Through this process, an AI develops its embedding matrix, which is used to discern the probabilities of words being used in the context of their surroundings. Advanced models are trained on unimaginable amounts of data, which is what allows it to be so accurate to human speech. Additionally, they can determine when specific kinds of responses are appropriate, for example when prompted to respond casually versus academically, they have different probability matrices for the way they decide their words.</p>

<p>Before we link AI back to our infinite library, let’s finish our explanation of how it works. When an AI is reading a sentence, as much as we would like it to be the case, it doesn’t just generate an embedded sequence of matrices for each word. First, it breaks down the sentence into tokens, which are normally words, but can be meaningful subsections of words as well. It then assigns positional matrices, which tell the AI the order in which the tokens are organized. Then, it sends the tokens through its constructed embedding matrix, and outputs them as embedded vectors. So the example I gave earlier is not entirely representative, most of the time a sentence will be represented as a list of embedded vectors corresponding to each word. The length of an embedded vector is dependent on the sophistication of the language model being used, as it is what the AI uses to encode meaning. This length can range from dozens to thousands. So the full process is as such: tokenization → embedding matrix → positional matrix.</p>

<p>Let’s explore how an AI might be able to help us organize our library of babel. If we are starting with a completely untrained AI, it would be even more lost than us, as even if it encountered sensible text it would not know what it meant. However, let’s say in addition to the AI we had a human interpreter. The AI could bring to this person books it believes to be containing real text, and have them assess its classification. Over time, by recognizing patterns in the books it brings which contain real information, the AI would gradually train itself. After enough time has passed, it will be able to correctly distinguish between sense and nonsense with some level of acceptable accuracy. We could then have this AI read as much of the library as possible, and return with the collection of books it deems to be sensible. This is a possible method of sorting the library, and probably one of the more efficient ones.</p>

<p>When Borges originally wrote his short story pertaining to this library, it was laced with futility, as he believed that the library was unexplorable to a meaningful extent. He begins the story with hope for the utility of the library, but as the characters explore the library, the amount of useless data becomes insurmountable. The number of books is truly infinite, and therefore, the number of illogical books is infinite as well, along with the number of logical ones. However, with the method described using AI, we can observe a very large number of books, arguably any finite number given enough time. Therefore, depending on how the library is organized, we could imagine sorting all books under a certain specific length. This throws the pessimistic view of the library from borges into question, as even if we can’t truly look through each and every book, we can find some extremely interesting things. There’s so many more questions to ask, and we haven’t even begun to delve into many of the interesting philosophical ramifications of having such a library, which was the main focus of the original story. However we have used it as an excuse to mathematically interpret language and explain an interesting technology, which I think is pretty cool. Thanks for reading!</p>]]></content><author><name></name></author><summary type="html"><![CDATA[The Library of Babel, imagined by writer and poet Jorge Luis Borges in his short story of the same name, is a theoretical, infinitely large library containing all possible information, constructed of symbols both familar and unknown. This library would contain the sum total of all human knowledge, as well as all future works which could be created. However, these peices would be lost in a sea of gibberish.]]></summary></entry></feed>