Näytetään tekstit, joissa on tunniste mathematics. Näytä kaikki tekstit
Näytetään tekstit, joissa on tunniste mathematics. Näytä kaikki tekstit

lauantai 29. heinäkuuta 2023

János Bolyai: The science absolute of space: independent of the truth or falsity of Euclid's axiom XI (which can never be decided a priori) (1832)

 

(1802–1860)

A good example of a remarkable coincidence of two persons having almost the same idea roughly simultaneously is the discovery of non-Euclidean geometries, and more precisely, the so-called hyperbolic geometry. We have already seen how Russian mathematician Lobachevsky approached the idea of not assuming Euclid’s parallel axiom, and we are now about to see the Hungarian János Bolyai do it in his own manner.

One might say that the interest in the parallel axiom ran in the Bolyai family, since János’s father, Farkas Bolyai had for a long time tried to deduce the axiom. Indeed, he had warned his son that the parallel axiom was something younger Bolyai should keep away from, since one could waste a lifetime thinking about it. To older Bolyai’s surprise, younger Bolyai sent his father a short paper dealing with the issue, which older Bolyai published as an appendix to his own textbook on mathematics.

The topic of Bolyai junior’s article is primarily the absolute geometry, that is, a geometry where neither Euclid’s parallel axiom nor its denial is assumed. Thus, just like Lobachevsky, Bolyai is interested firstly in the similarities between the Euclidean and the hyperbolic geometry. Again like Lobachevsky, Bolyai defines as a parallel line to a given line X as that precise line, which is a sort of limit of all the lines drawn through the same point y and not cutting the given line X, in the sense that any line falling from that point y more toward the given line X will cut this given line.

Bolyai also shows, like Lobachevsky before him, that parallelism, defined in this manner, is a transitive relation. Bolyai goes even further and shows that with the parallel lines, the congruence of line segments is also a transitive relation. He then notes that given a line segment AM, one can consider a collection of all such points B that if a line segment BN is parallel to AM, it is also congruent to it. Bolyai doesn’t really give any other name to this collection, but F. In addition to F, Bolyai considers the intersection of F with any plane containing AM, which he calls L, while AM he calls the axis of L. He also notes that F can be described by revolving L around AM. Because parallelism and congruence of line segments is transitive, every line segment BN starting from a point B in L and parallel and congruent to AM is also an axis of L.

In Euclidean geometry, Bolyai notes, this L is simply a line perpendicular to AM - and F, similarly, a plane perpendicular to AM. In hyperbolic geometry, on the other hand, L is not a straight line, but a curved line, and similarly F is a curved surface. Indeed, they are what Lobachevsky called respectively oricycle and orisphere. Like Lobachevsky, Bolyai notes that oricycles in an orisphere work like straight lines in a Euclidean plane.

Now, Bolyai pictures an axis AM move through its oricycle L, always staying at same angle to L. He saw that the other points of the axis AM described further oricycles, that is, taken C from AM would describe an oricycle, for which AC would be the axis. Furthermore, taking corresponding parts of two such oricycles, the relation X of their lengths is always a constant, which depends not on the length of the parts, but only of the distance x of the points on AM. Bolyai also notes that these oricycles are always congruent, although part of one appears to be multiple of the corresponding part of the other.

Bolyai further notes that whether one supposes Euclidean or hyperbolic geometry to hold, spherical trigonometry – that is, study of triangles, as it were, on the surface of a sphere – always follows the same rules. Then again, he adds, with the ordinary geometry the case is quite reversed. In Euclidean geometry, given the length of at least one side of a triangle and at least two other elements of the same triangle (whether angles or sides) are known, the other elements can be solved. In hyperbolic geometry, on the other hand, one has to also refer to some length x, of which the corresponding relation X is known, to make similar calculations. Bolyai suggests using length i, defined by having as the corresponding relation e - the basis of natural logarithms. This length i would then work as a sort of natural unit of length in the hyperbolic geometry.

The intriguing question is then, which of the two, Euclidean or hyperbolic geometry, is the one describing the world we live in. Bolyai notes that we cannot really determine this without any empirical facts to guide us, since neither of the two geometries has any intrinsic flaw in it. This statement goes against the common idea of Bolyai’s contemporaries that Euclidean geometry is somehow inherently intrinsic. As if to just spite such thinkers, Bolyai ends his short article by showing how one can construct in hyperbolic geometry a square equal in area to a circle – something that is impossible in Euclidean geometry.

sunnuntai 16. heinäkuuta 2023

George Peacock: Treatise on Algebra (1830)

 

(1791–1858)

On a surface level, Peacock’s Treatise on Algebra seems a rather humdrum textbook on algebra, with no particularly original innovations to be found in it. Yet, what is most important in the treatise are not any purely mathematical results, but the more philosophical considerations on the nature of algebra and its relation to other parts of mathematics and especially arithmetics.

What Peacock is particularly interested in is the question whether algebra should be somehow based on arithmetics, which is based on concrete groups of objects, like seven apples. Peacock’s answer is a resounding no: arithmetics is far more constrained, allowing e.g. no negative numbers, since you cannot produce a group of, say, minus seven apples, although negative numbers are a common occurrence in algebra.

Not restricted by characteristics of arithmetical numbers, what is to determine what rules are to be assumed in algebra? Peacock makes the bold suggestion that these rules are just assumed: algebra concerns only symbols, and we could in principle choose any rules to govern our calculations in it.

Still, Peacock does not yet reach the modern notion of algebra, where we can have many different systems of calculation with different rules. Instead, he thinks that algebra should be especially the most general system of calculation, where all sorts of calculation are possible. Thus, he accepts not just calculations leading to negative numbers, but also roots of negative numbers, which we nowadays call imaginary numbers. Peacock goes even so far as to suggest that such seemingly nonsensical symbols like 0/0 could be used in a meaningful manner (although he mentions this possibility only in passing, he is referring to certain readings of the infinitesimal calculus).

Peacock still wants that algebra would be of practical use. Here, he suggests that the rules of algebra should at least correspond to the rules of arithmetics at least in those cases where the calculations and their results make arithmetical sense. Then again, he continues, results of algebra could also be used in other mathematical sciences, where some of them might be of relevance. Thus, negative numbers do make sense, for instance, when speaking of debt or of movement to an opposite direction. Even imaginary numbers can have concrete meaning, as describing movement perpendicular to a line.

Algebra becomes then, for Peacock, like a general tool for mathematical problem solving. When a concrete problem is given, it is turned into algebraic symbols: what quantities are known and what are unknown that are supposed to be determined in terms of the known quantities? After doing the algebraic manipulation of the symbols, one must still interpret the results. Often the context of the problem restricts this interpretation, and although algebra would lead to a number of possible results, only some of them might make sense in the context. Sometimes no result suggested by algebra would make sense, and then the problem itself will be impossible.

keskiviikko 17. toukokuuta 2023

Évariste Galois: Dissertation on the conditions of solvability of equations by radicals (unpublished)

I’ve already discussed Galois as an example of a mathematical wunderkind. Like with many wunderkinds, his life was not long, as he was killed in a duel, when he was just twenty years old. Galois had not published much before his untimely death, and indeed, the work he is mostly remembered fpr nowadays, Mémoire sur les conditions de résolubilité des équations par radicaux, was rejected by the Academy of Sciences, due to being incomprehensible in its present form. Indeed, the evaluation of the Academy was not completely without grounds, since Galois often passes over required steps in his proofs, if he considers them to be sufficiently obvious. Still, it seems clear that at least part of the rejection was due to Galois’ method being so original and unseen.

The topic of Galois’ paper concerns the possibility of solving polynomial equations, that is, of equations of the form a1xn + a2xn-1 + …. an = 0. Methods for finding the possible solutions of the equation, using only the operations of addition, subtraction, multiplication, division, raising to power and taking a root of the coefficients, had been known for all cases where n is at most 4 (or as the mathematicians would say, where the degree of the equation was at most 4). In mathematical terms, such methods were known as solving the equation by radicals (radical being the symbol for taking the root). It had also been shown that already when the degree of the equation was 5, there are cases where the equation could not be solved by radicals. What Galois added was a criterion by which all equations solvable by radicals could be recognised.

Galois’ method of finding this criterion was preceded by the idea of Lagrange to study the solutions, or to put it in mathematical parlance, the roots of the equation, before they were actually known. Lagrange had especially considered what happens, when in combinations of these roots, using only addition, subtraction, multiplication and division, the places of these roots were changed - permuted, as they say in mathematics. Now, Lagrange had shown that the structure of the equations had something to do with these permutations, although he had not yet managed to find the complete tale.

Galois notes that if a polynomial equation with degree n has n different roots (the maximal number it could have), one could always form a combination of these roots, where every permutation of the roots would change the value of the combination. With a characteristic leap of thought, Galois just states that some combination of the form Aa + Bb + Cc + … would suffice (a, b, c… being roots of the polynomial and A, B, C … being appropriate integers), without proving this statement or even giving any method how to find such combination (the statement is true, but let’s just take it on faith and not go in full detail how to prove it). Furthermore, he points out that Aa + Bb + Cc + … can be assumed to be irreducible, that is, not expressible as a product of other polynomials.

Another point that Galois makes rather quickly is that all of the individual roots can then be expressed as roots of a polynomial in terms of V, which is the numerical result of the combination of the form above, where the root expressed is, say, the first in the combination. In other words, if V = Aa + Bb + Cc + …, then we can find another polynomial F, such that F(V, a) = 0. Even more, Galois points out, just by permuting the roots a, b, c … in a suitable fashion in the combination, one can find for each root (say, b) a numerical result of a suitable permutation of the combination (say, V’), which can then be used to express the root with the same polynomial F (e.g. F(V’,b) = 0). Note that it is quite arbitrary, what is the order of the roots in the first V: it is the permutation of them to make V’ out of V that is of great importance. Indeed, even the exact values V, V’, V’’ … corresponding to the roots a, b, c … are not as important, Galois notes, as the way in which the roots have to be permuted in order to construct the values V, V’, V’’ … . This group of permutations, or group of the equation, as Galois calls it, has the interesting characteristic that whenever some arbitrary function F(a, b, c …) has a value expressible in the by now familiar terms of addition, subtraction, multiplication and division of the roots, then this function will have the exact same value, if the places of the roots are changed with a permutation in this group.

Galois seems to have thought of his group of the equation as a sort of matrix formed of roots, where one line describes one permutation of the roots. Next, Galois makes the interesting suggestion that while the Vs determining the group of the equation are defined by integer coefficients A, B, C …, we could, as he says, adjoin a new quantity, beside integers, that could take the place of the coefficients (this quantity could be, for instance, an irrational root of some integer). With this addition or adjoinment in place, it might become possible to express the polynomial Aa + Bb + Cc + … as a product of further polynomials. If so happens, the group of the equation, expressed in terms of possible values of Aa + Bb + Cc + …, is, as Galois says, partitioned or decomposed into smaller groups. All of these smaller groups happen to be of the same size and also of the same form in the sense that one subgroup can be turned into another with a rulelike permutation.

Now, if just suitable quantities to be adjoined can be found, the procedure can in principle be continued further. In this case, we can move to polynomials of smaller and smaller degrees and eventually hit the rock bottom, when the corresponding group contains nothing but one row. If this can be done, the original equation will be solvable by radicals.

While the aim of the paper is to find the method for recognising polynomial equations solvable by radicals, Galois himself seems to become more and more interested in the supposed mere means for this goal, namely, the group of the equation and the permutations involved. Indeed, this is the road that the development of mathematics and especially algebra was to take: it was not anymore just a method for solving equations, but a more intricate study of such abstract structures like groups of permutations.

maanantai 21. helmikuuta 2022

Auguste Comte: Course of positive philosophy 1 - Mechanics

Comte’s first volume of his positive philosophy ends with a study of mechanics. Although mechanics is for Comte even more concrete than geometry, it still falls within mathematics. Indeed, Comte is against all readings of mechanics, where some elements of analysis are interpreted as real forces, although they would be nothing but means for making calculations. What forces are real can be decided only by observation and experience. Furthermore, Comte adds, mechanics does not investigate what is the nature of these forces, but merely the movements caused by them.

The abstract nature of mechanics, characteristic to mathematics in general, Comte notes, is seen in the fact that mechanical calculations are simplified by assuming bodies to be passive or inert, although in reality they are in many ways active. This idealisation of bodies as inert in mechanics, Comte explains, is not to be confused with inertia in the sense expressed in one of the basic laws of mechanics - the fact that bodies tend to move in straight lines and retain their state of movement. Together with the other two basic laws of mechanics - one being Newton’s law of action and reaction, the other being Galilei’s discovery that forces are independent of one another and can thus be composed with the parallelogram law - the law of inertia is, according to Comte, based on observation, not on any a priori deduction. Particularly, Comte adds, the law of inertia cannot be deduced from the law of sufficient reason.

Comte divides mechanics, expectedly, into statics dealing with instantaneous forces and uniform movement or equilibrium arising from them and dynamics dealing with continuous forces and varied movement arising from them. Within both statics and dynamics, he then differentiates a part examining solid bodies from a more complex part examining fluids. An important problem in this classification in Comte’s opinion concerns the relative status of statics and dynamics. Statics is clearly the older discipline, studied by ancient mathematicians long before dynamical questions. But as Comte has said earlier, the historical order of disciplines does not necessarily correspond to the order of the disciplines in a completed science. Indeed, when dynamics was finally introduced at the start of the Modern Age, statics was regarded as a mere abstract limit case of dynamics.

Yet, as Comte’s favourite mathematician, Lagrange, had argued, the notion of virtual displacement - an application of his calculus of variations to mechanics - could be used to reduce all of dynamics to statics. In effect, Comte is referring to the central idea of the so-called d’Alembert principle that all apparently dynamical systems can be regarded as being in equilibrium. Comte further links Lagrange’s idea with Poinsot’s notion of a force couple, which Comte thinks is a modification of a notion of force from translation to rotation.

Just like with other parts of mathematics, Comte is especially interested in the applications of mechanics. Thus, he points out that statics is used for finding mass centres of bodies, while in dynamics we are trying to calculate movement of a particle from forces affecting it, or the other way around, to find forces creating a known movement. We need not go in great detail to theorems that Comte lists as consequences of the three basic mechanical principles. I will just point out that Comte speaks against interpreting Maupertuis’ principle of least action in a theological or metaphysical manner suggesting that bodies would somehow choose to move in accordance with the principle.

lauantai 22. tammikuuta 2022

Auguste Comte: Course of positive philosophy 1 - Geometry

While many philosophers had considered it an important problem to put geometry on secure foundations, Comte finds such attempts mere unfounded metaphysics. For him, geometry is simply a natural science with an empirical basis. Of course, it is the most abstract natural science, dealing only with static spatial properties of things, in abstraction from any movement. Still, Comte feels no need to prove the basic axioms of geometry, since he can just assume them as bare facts. On the other hand, he also feels no need to consider the possibility of other geometries with other axioms, since experience appears to agree with the ordinary Euclidean geometry.

Comte supposes that this empirical science has a very practical purpose, namely, that of measuring spatial features of things. Of course, he adds, not all measuring is geometry, for instance, if we fill an oddly shaped container with water and then measure the volume of the water, this is still not geometry. Instead, geometry, like all mathematics, is an art of finding out quantities indirectly, through calculations.

When we measure spatial features of things, Comte continues, we can measure all the dimensions of it or then only some of them. When a geometer speaks of planes or lines, he adds, they are considering just such abstractions, that is, they are ignoring some of the dimensions the thing has. Thus, lines in reality always are wide and thick, we just concentrate on their length. We can even ignore all the dimensions of a thing and consider only its position in relation to other things - this is the origin of the notion of point.

While the essence of geometry is indirect measuring of spatial features of things, a geometer must assume some ways to directly measure these spatial features. This implicitly assumed form of measuring, Comte suggests, is the measurement of straight lines or lengths by comparing them with a length of some other thing, like a ruler. Other geometrical figures (curves, areas and volumes) are then to be measured with the help of these straight lines.

Comte admits that geometry is full of other things beyond mere measuring, that is, full of propositions about spatial properties of things. Still, he insists, even these properties are ultimately studied, because they could help with measuring (perhaps in some more concrete science). Because one cannot know beforehand what properties help with measurements, Comte advocates studying as many such properties as possible.

Comte mentions the old distinction between synthetic and analytic methods in geometry, but seems to have no idea what these terms meant originally and what their difference was supposed to be. Instead, Comte suggests that the distinction is just a roundabout way to distinguish ancient from modern geometry. Ancient geometry, Comte suggests, was mainly dealing with concrete and individual figures, taking one type of figure and finding all its characteristics. Because of the uniqueness of the chosen figure, none of these characteristics could be assumed to hold for other entities. Modern geometry, on the other hand, deals with abstract geometrical problems, which can then be applied to many different contexts.

Comte has little patience with ancient geometry, and he especially dislikes the common habit of starting to teach geometry from works like Euclid’s Elements - as he has noticed earlier, history of a discipline is usually not the most convenient way to teach it. Comte is especially critical of the proofs of early propositions in Elements, where Euclid simply places figures on top of each other, to show their similarity - this is just as futile as an attempt to prove the parallel axiom would be.

Although Comte is very critical of ancient geometry, he does admit that Greek mathematicians did make some progress. Particularly, they perfected the study of the simplest kind of figures, namely straight lines and polygons and polyhedras. Furthermore, while Comte ridicules the Greek use of diagrams as proofs of proposition, he does suggest it to be a sort of precursor for modern projective geometry. Another field developed by ancient geometers to perfection, Comte concludes, is trigonometry, where they used certain lines (sines,tangents etc.) to represent angles and their relations and so simplified calculations.

Comte associates the birth of modern geometry with Descartes. While it had been long known that geometric forms could be described in terms of spatial situations of their limiting points etc., it was the invention of Descartes, Comte says, to reduce the talk of situations to talk of lengths and other magnitudes through the notion of a coordinate system. In effect, Comte concludes, Descartes was able to transform at first sight qualitative properties like geometric figures into quantitative properties.

The outcome of the Cartesian transformation of geometry, Comte explains, was that now in two-dimensional geometry lines could be expressed by equations and equations by lines (in three-dimensional case, Comte adds, equations express surfaces, while lines are expressed by pairs of equations). These equations do not just characterise some random properties of the lines, Comte notes, but explicate how they could be generated. He adds that when a geometer is looking for an equation to describe a line, there is no need to choose any specific coordinate system - sometimes the searched for equation might be easier to describe e.g. in terms of polar coordinates, which are especially convenient when describing rotations. Then again, he admits, when finding lines to describe equations, it is best to pick the rectilinear coordinate system, which is the most natural for us to decipher.

Although Comte at first appears to say that all lines correspond to an equation and all equations correspond to a line, he is well aware that analytic geometry of his time has imperfections in this department. Firstly, he notes, discontinuous lines cannot be expressed so well in terms of a single equation. Furthermore, he continues, equations with more than three variables or equations with imaginary solutions have no proper geometric model.

Like in the case of abstract mathematics, Comte is mostly interested in the uses geometry could be put to and the problems that could be solved by its help. Solving some of these problems relies on simple algebraic means, such as when we try to find the number of points that are required for determining the course of a curve. Still, most of these problems, Comte notes, rely on the help of differential and integral calculus. Differentiation is useful not just for finding tangents, but also for describing e.g. the curvature of curves. Then again, Comte concludes, integration is the most useful tool, because it helps us to fulfill the true task of geometry, that of measuring lengths, areas and volumes.

lauantai 28. elokuuta 2021

Carl Gustav Jacob Jacobi: Fundaments of a new theory of elliptic functions (1829)

(1804-1851)
The development of mathematics has often been one of generalisation: concrete problems have demanded development of very abstract tools that can then be applied to various other fields. One clear example is integration. Originally developed as a tool for calculating lengths, areas and volumes, it could then be used in various other contexts where calculating limits of infinite sums with high precision was required.

Another aspect of the development of mathematics, then, has been that these very abstract tools themselves provide new problems and topics for discussion. For instance, going back to the example of integration, if you progress beyond Calculus 101, you soon learn that not all integrals can be solved through those neat formulas given in the textbook, but in the worst case scenario you have to go to the definition of integral and approximate it through various finite sums. Of course, mathematicians have found various new tools for simplifying this numerical process in individual cases. A particularly interesting case concerns the so-called elliptic integrals.

The very concept of an elliptic integral belies an origin in quite concrete geometric problems: measuring the length of pieces of a curve called ellipsis (picture an elongated version of a circle). In truth, this example is just one version of elliptic integrals, the unifying elements being certain simplicity in the formal characteristics of the base function integrated (to put it very briefly, they involve nothing more complex than fractions with denominator a square root of polynomial of third degree). These kinds of integrals are already harder than those in elementary textbooks and their exact values can often be just approximated. The question is how to simplify this process of approximation.

First step in this simplification was provided by Adrien-Marie Legendre, who showed that all the various elliptic integrals could be reduced to three paradigmatic cases - already a huge improvement. Another important step was to note that these elliptic functions could be expressed in terms of two parameters: an angle called an amplitude of the integral and a number called module. In the particular case of ellipsis, the amplitude describes the angle the x-axis makes with a line joining the origin and the tip of its particular arc, while the module describes how elongated the ellipsis is (0 being the case of a proper circle).

Jacobi’s Fundamenta nova theoriae functionum ellipticarum took the simplification a few steps further. The basic problem Jacobi set out to solve was to show that an elliptic integral with a seemingly more complex structure could be reduced to an elliptic integral of a less complex kind, provided some relation of them, expressible through relatively simple algebraic means, was shown to hold between them. In other words, by knowing that such a relation existed between the two integrals, one could calculate the value of the more complex on the basis of the simpler one. Jacobi calls this transformation of the elliptic integral.

Now, the question was how to determine this relatively simple transformation. Jacobi showed that this question was essentially the same as determining a certain type of relation between the modules of the two integrals (what he called modular equation). In principle, this modular equation could be calculated through algebraic means, but in practice, the more complex the equation changes, the more cumbersome this calculation becomes. Jacobi’s solution is to go a bit further in the level of abstraction and to construct a more general rule picking a series of suitable modules that can be linked with such transformations.

Jacobi’s derivation of this rule is based on the second parameter, the amplitude, and the so-called elliptic functions, which can be defined on the basis of the amplitudes. To give a rough idea of these elliptic functions, we can compare them with simpler trigonometric functions. It is a well-known fact that trigonometric functions can be described in terms of a unit circle and angles set up on its centre. Elliptic functions can also be described in terms of an angle set up on a centre of an ellipse.

It was just inevitable that this new tool - elliptic functions - became a topic interesting in itself. Thus, the second half of Jacobi’s work is dedicated to the study of elliptic functions, which, just like trigonometric functions, are a source of many beautiful equations. A particular question Jacobi dealt with was how to express these functions as infinite series. In effect, this was yet again a way to find more and more good approximations for elliptic functions. Finding these infinite series required the introduction of yet another tool: the so-called theta functions, which are a certain type of series of complex numbers - and the development of the mathematics continued.

perjantai 13. elokuuta 2021

Nikolai Lobachevsky: Geometrical researches on the theory of parallels (1840)

 

(1792-1856)

Euclid’s book on geometry has for ages been seen as an ideal of an axiomatic theory, in which everything is based on a solid basis of definitions and evidently certain axioms and proven through strict demonstrations, making the results presented appear indubitable. No wonder many works of philosophy tried to imitate Euclid’s style, to make their theories seem as indubitable and necessary, usually failing miserably to be as convincing as Euclid.

If you know your Euclid by heart, you know that he had not really achieved the ideal many want to see in his book. There are sometimes slight hidden assumptions in his proofs - and isn’t it a bit too empirical to carry around triangles and put them on top of one another (Euclid, I.4)?

The most glaring fault in Euclid’s work is, of course, the infamous parallel postulate. When compared with other postulates of Euclid, it appears complex and far from self-evident: “if a straight line falling on two straight lines makes the interior angles on the same side less than two right angles, the two straight lines, if produced indefinitely, meet on that side on which are the angles less than the two right angles”. The postulate can be made a bit clearer with some rewording and use of more modern phraseology - if line A cuts two other lines on the same plane, B and C, and the sum of interior angles on one side equals pi (or 180, if you are more into degrees), B and C eventually cut one another on that side and are therefore not parallel. Even with this rewording, it seems like a theorem we should prove, not a postulate to be just assumed.

Many professional geometers and even more geometry dilettantes were equally unimpressed by this postulate and tried to demonstrate it from other postulates. Their efforts led at most to finding other postulates that could replace Euclid’s. Most famous of them is the so-called Playfair’s axiom: given a line and a point, we can draw through the point, at the same plane as the line and the point, only one line that does not cut the first one. This does sound simpler, but still lacks the self-evidency of the other postulates.

Dissatisfaction with Euclid’s postulates and its alternatives continued, but no solution was forthcoming. All of this was changed by Lobachevsky’s seminal paper, Geometrische Untersuchungen zur Theorie der Parallellinien. Well, to be truthful, he had already written papers on the topic in his native language, Russian, in 1820s, but these did not circulate very widely (and in addition, I cannot read Russian).

Lobachevsky’s starting point is the Playfair’s axiom, but instead of just assuming it, he asks what would happen, if there were more than one line we could draw through the point - lines which would not cut the given line. He notes that even then we could find a single particularly interesting one among those non-cutting lines, namely, the limit between lines that do and those that do not, and suggests calling this the parallel line. Then, by tying Playfair’s axiom back to the original framing of Euclid’s postulate (C being the original line, B being parallel to it, A cutting them both and A and B meeting at the given point) and by making the assumption that A cuts C perpendicularly, he notes that in this peculiar setting B and A form an angle less than half the pi (or 90 degrees), thus contradicting Euclid’s postulate. The angle formed by A and B (and dependent on the distance of the given point from the line C), Lobachevsky calls the angle of parallelism.

Although Lobachevsky’s new geometry - later dubbed hyperbolic geometry, although he himself called it imaginary - has clearly different properties from Euclidean geometry, a significant portion of Lobachevsky’s paper is committed to show similarities to Euclidean geometry. The simplest similarity is that the new definition of parallel lines works similarly enough to the Euclidean notion, for instance, parallelism is a symmetrical and transitive relation.

A more intricate similarity Lobachevsky finds through notions of oricycle and orisphere. By oricycle Lobachevsky means such a curve in hyperbolic geometry, all perpendiculars or axes of which are parallel to each other. Furthermore, oricycle is also a sort of limit for circles - by enlarging the ray of the circle indefinitely, in hyperbolic geometry, the curve tends toward the oricycle. In a figurative way, we could say that an oricycle is an infinite circle. Interestingly, the same concept in Euclidean geometry means ordinary straight line.

The notion of oricycle taken into three dimensions forms, then, an orisphere. Technically, an orisphere can be formed from an oricycle by turning it around one of its axes. What is interesting is that oricycles on orisphere work like straight lines on a plane in Euclidean geometry, for instance, in a “triangle” formed of segments of three different oricycles, the sum of the angles equals pi.In effect, two-dimensional Euclidean geometry can be ingrained within three-dimensional hyperbolic geometry.

Although the aforementioned pseudotriangles in hyperbolic geometry do follow same rules as regular triangles in Euclidean geometry, regular triangles in hyperbolic geometry do not. Yet, Lobachevsky points out, when sides of the triangles in hyperbolic geometry decrease indefinitely, the more they start to resemble the triangles in Euclidean geometry, for instance, the sum of their angles approaches pi.An interesting consequence of this is that the bigger the triangles in question are, the more apparent the difference of the two geometries becomes. It becomes then an empirical problem to decide whether we live in a space with a Euclidean or a hyperbolic geometry - just make astronomical measurements of distances between stars and you might notice signs of non-Euclidean properties.

perjantai 15. tammikuuta 2021

Évariste Galois: Demonstration of a theorem about periodic continued fractions (1828)

(1811 - 1832)
Mathematics is a field where wunderkinds are a real possibility. It is a field with abstract enough objects that can be grasped even with very little experience, even if dealing with them and their interrelations might require some innate or acquired skills.

A good example of such mathematical wunderkind is Galois, who had managed to make significant strides in algebra before his death, when he was just a few months over twenty year old. Indeed, his first paper, ”Démonstration d'un théorème sur les fractions continues périodiques”, was published just a few years before his death, making his life as a mathematician remarkably short.

The paper itself contains just one beautiful theorem, with a proof simple enough to understand, once you just see how Galois does it. First, some terminological explanations are in order. The topic of the article is continued fractions. To understand what this means, one can start from a sum of a whole number and a fraction of the form a + 1/b. Then let us suppose that the denominator of the fraction is not just another whole number, but another sum of a whole number and a fraction: the whole is then something like a + 1/(b+1/c). We can clearly continue doing the same again, adding fractions within fractions. Like in many other cases we have witnessed, we can think of this development continuing indefinitely into ever smaller and smaller fractions and progressing closer and closer toward some definite number.

Now, the simplest kind of a continued fraction to deal with is such where the numbers a, b, c… follow some clear rule. A specifically simple case is such, where after a while, the numbers a, b, c…continue again from some previous number of the series, repeating after that the same series of numbers over and over again. In such a case, the continued fraction is called periodic.

Before Galois, mathematician Lagrange had showed that periodic continued fractions had a special relationship with quadratic equation, Ax2 + Bx + C = 0. The root of such an equation - that is, a number, which makes the equation correct when put instead of x - can be expressed as a periodic continued fraction, and conversely, periodic continued fraction is a root of such an equation.

Galois continued Lagrange’s work and demonstrated another relation, involving especially immediately periodic continued fractions. If the number where the repetition begins is same as the first number a - that is, if there is no additional sequence of numbers at the beginning of the series - the continued fraction is in the terminology used by Galois called immediately periodic.

Let’s say we have an immediately periodic continued fraction, where the repetition of the numbers begins after the fourth, d. Then the fraction is of the form a + 1/(b + 1/(c + 1/(d + 1/(a + … If you look very carefully, you can see that the denominator following the first d repeats the whole series representing the fraction. In other words, if we designate the fraction with x, then we can say that x = a + 1/(b + 1/(c + 1/(d + 1/x))).

Now, it is a simple task requiring simple operations to dig out the x on the right side of the equation. I’ll just give the first few steps to explain what I mean: a - x = - 1/(b + 1/(c + 1/(d + 1/x))); 1/(a - x) = - (b + 1/(c + 1/(d + 1/x))); b + 1/(a - x) = - 1/(c + 1/(d + 1/x)) and so on. The end result of this manipulation of the equation is this: -1/(d + 1/(c + 1/(b + 1/(a - 1/x))) = x. In the place of the inner x, we can again start writing the series from the beginning: -1/(d + 1/(c + 1/(b + 1/(a + 1/d

In other words, Galois has found by this simple method a new periodic continued fraction. This new fraction also works as a root of the same quadratic equation as the original periodic continued fraction. In addition, the numbers of the new fraction move in opposite direction from the original one. Furthermore, if we assume that all the numbers a, b, c and d are positive, the original fraction expresses a positive number larger than 1, while the new fraction is a negative number smaller than - 1.

lauantai 16. toukokuuta 2020

August Ferdinand Möbius - Barycentric calculus: new aid for analytical consideration of geometry (1827)

1790-1868

Concrete sciences have often been a spur on the development of new mathematical methods. Take as an example Archimedes, who used the idea of balancing a mass of parabola with a mass of a triangle and thus finding out the size of an area limited by the parabola. The exact method Archimedes used in determining this area was a rudimentary form of integration, with Archimedes, as it were, adding infinitely small pieces of the area together.


The idea of mass and weight was, coincidentally, a spur for Möbius’ text called Der barycentrische Calcül: ein neues Hilfsmittel zur analytischen Behandlung der Geometrie. Möbius starts with the notion of the center of mass, determined by mass of three objects situated in a triangle. He then quickly finds out that in fact any point within the triangle can be expressed as such a center, by just changing the masses assigned to the apex. Indeed, one need not be confined to points in this triangle, but allowing that in this case some of the masses could have a negative mass.

The end result is then the expression of points - and ultimately lines and figures - through three points of an arbitrary triangle. One can expand this meth to three-dimensional case by switching triangle to a pyramid. One can even express various simple functions in these terms - e.g.express curves like parabola, hyperbole and ellipse in terms of a three points of a triangle - and translate some concepts of infinitesimal calculus in this “barycentric” form.

A more interesting question is why Möbius uses such a peculiar form of expression instead of the more familiar method of expressing geometrical figures through two coordinate axes - that one can do seems a poor reason for actually doing so. The answer appears to be simply that some important relations are easier to express in this manner. Möbius is especially interested of what we nowadays might call equivalence relations between different figures. Two examples of such relations have been known since the beginning of mathematics - first is the equality of the sizes of the areas, no matter what their shape is, and second is the similarity of shapes, no matter what their size.

Beyond these two equivalence relations Möbius considers two others, first of which, affinity, had been introduced already by Euler. One might say that this affinity is a relaxed version of similarity, that is, a relation that could hold even between non-similar figures, although similar figures are always affine. Möbius introduced this relation by asking his reader to change the triangle used to determine figures with some completely different (not even similarly shaped) triangle. The figures defined by such change need not be similar - an example of such a pair is formed by a square and a rhombus. Still, they do have some resemblances, for instance, quantitative relations between respective parts (more precisely, parallel line segments) of these figures remain same.

A further form of equivalence relation Möbius discovers is what he calls collinear resemblance. Simply put, Möbius starts from affinity, takes away even further properties of the two figures and leaves only one point of resemblance - that all straight lines in one figure correspond to a straight line in the other figure. This definition loses even the identical ratios of lines, resulting in even more abstract resemblance. Some common quantitative properties still remain, namely, the so-called cross-ratios (a specific type of ratio for four points in the same line). Still, this is a very loose sort of relation - even circle and parabola have this sort of collinear resemblance.

maanantai 6. huhtikuuta 2020

Jean-Babtiste Joseph Fourier: Analysis of determined equations (1827)

One of the parts of mathematics most easily applicable to practice is the study of equations. You need to find out a certain quantity or certain quantities, and you know that it has or they have a certain relation to other quantities. Finding out what these unknown quantities are means solving the equation.

Although a layman might think that mathematics should always give exact solutions to such problems, it is quite obvious that whether and in what magnitude giving exact solutions is possible is a question that does not always have a clear cut answer. Take such a simple equation like x2 = 2. We know that the solution to this problem cannot be expressed as a ratio of whole numbers. Still, despite the objections Pythagoreans would have had, we are usually accustomed to say that solutions involving roots are precise - at least they tells us that the relationship that the searched for quantity has to known quantities, even if we can express the numeric value of the former only approximately.

It has been long known that for some, relatively simple equations, such relatively exact answers can be found, if there just is an answer to be found. Let’s take a case where we are searching for a single unknown quantity, with a relation to zero, describable in terms of such simple calculations as sum, multiplication and squaring:  x2 + ax + b = 0 (the so-called quadratic equation). There’s a simple enough formula for solving such equations, using again only very simple operations - addition, subtraction, multiplication, division, squaring and square roots.

We know that the formula for quadratic equation will give us two, one or none solutions - the last option occurs, when the formula would involve a square root of a negative number, something that is usually an impossibility, when applying mathematics in more concrete fields, although we can construct an abstract system with such square roots of negative numbers (the so-called imaginary numbers). We also know that the solutions revealed by the formula are all the solutions the equation could have, and we can even represent this geometrically: the equation describes a figure known as parabola, which can cut one of the axis of coordinate system twice, touch it at one point or then not cut or touch it at all.

The situation becomes somewhat more complicated when we allow exponents larger than 2 in the equations, that is, when we deal of general polynomial equations of the type xn + axn-1 + bxn-2  + … + rx + s = 0, where the highest exponent is called the degree of the polynomial. We do know something general about the solutions of such equations. If the n is odd, the figure described by the polynomial function xn + axn-1 + bxn-2  + … + rx + s is like a rising line: with very large, but negative values of x, the result of the polynomial is negative, while with very large, positive values of x, the result is positive. If n is even, the figure resembles parabola, where large values of x, whether positive or negative, produce positive results. The only difference is that with larger exponents, the figures might have more bends - the maximal number of bends in the figure is always n - 1, where n is the degree of the polynomial. This means that the maximal number of solutions for the equation is also the degree of polynomial - every new bend makes one more point of contact with the x-axis of coordinate system possible.

Although the maximum number of solutions of polynomial equation is known, we might not always be exactly sure what these solutions are. With polynomials of degree 3 or 4, a general solution of similar sort as with quadratic equations can be given. Then again, with polynomials of higher degree such a general solution does not - and even more, cannot - exist. We might be able to find the exact solutions sometimes, but there’s no guarantee we could do it always.

Even if a general method for finding exact solutions does not exist, we might still have a method for finding inexact solutions, that is, better and better approximations of the searched for solutions. Such a method of approximation can also be of mathematical nature, because we might have good mathematical reasons to say why a certain method works. A good example is the method invented by Isaac Newton. The basic idea behind Newton’s method is that at small distances a curve is similar to its tangent. Thus, if we have an estimate that is close to the final solution, we can use the tangent at the point of the estimate to get an even better estimate of the solution - just check where the tangent hits the x-axis and you get the new estimate.

The problem with this method is that if the first estimate is not close enough to the real solution, it might take many iterations to get even fairly good approximations. The problem thus becomes how to determine the regions where we should go looking for the A partial answer to this problem is provided by Fourier’s posthumously published work, Analyse des équations déterminées.

Fourier’s starting point is unexpected. He asks us to produce a derivative of the original polynomial, then a derivative of this derivative and so forth, until nothing else is left, but a constant function. The series beginning from the constant function and ending with the original function has n +1 members. What has this series of derivatives to do with the solutions of the original equation? Well, consider the results of the polynomial and the series of derivatives for very large negative numbers. The final constant function is always positive, the result of the next derivative in the series - a polynomial of degree 1 - is negative for very large negative values of x, while the next derivative - a polynomial of degree 2 - has with these values positive results. Generally, the polynomials of odd degree in this series have negative results for very large negative values of x, while polynomials of even degree have positive results. In other words, with these large negative values of x, a result of the function in the series is always of different sign than the result of its derivative, which means that the sign of the result changes throughout the series n times.

By itself, this result seems quite meager, but some further reflections show its importance. Firstly, checking what happens with very large positive values of x, we notice that the original polynomial and all the derivatives of the series have positive results, which means that the series has no sign variations. All the sign variations have vanished when moving from very large negative to very large positive numbers. Indeed, the only point when the number of sign variations may change is with those values of x, when the original polynomial or one of derivatives in the series produces zero - either the function producing zero cuts x-axis at that point and its sign changes when moving through, or then it just touches x-axis and the sign of its derivative changes.

A further important point is that the number of sign variations can never increase. If the function changes from positive to negative near a certain value of x, then the function is diminishing and its derivative must be negative near the same value of x, and if the function changes from negative to positive, then the function is growing and its derivative must be positive. Thus, supposing that the function changing the sign is also a derivative of another function of the series and thus in the middle of two other functions, the change of its sign can never increase the number of sign variations. For instance, if in one part of the series the signs are + for f’’(x), - for f’(x) and - for f(x) (with one sign variation between them), the change of the sign of the middle term changes the series into +, + and - (again with one sign variation between them). Then again, if the series was at first, +, - and + (with two sign variations), after the sign change it will be +, + and + (with no sign variation).

In effect, then, if we take two different values of x, at the smaller value the series of derivatives of the polynomial cannot have less sign variations than at the larger value. In fact, if we consider the difference between the sign variations at these two different places, this difference gives the maximum number of places between these values of x, at which the result of the polynomial will be 0 (roots of the polynomial, as they are called).

Of course, this method provides us only with a maximum number of possible roots of the polynomial between two values of x, and the interval might actually contain a lot less of roots. Still, with systematic division of such intervals - and few tricks Fourier uses to weed out intervals, which really contain no roots, despite the number of sign variations - it is possible to pick out certain intervals where the searched for solutions lie. The next step in Fourier’s method is then simply to use Newton’s method to approximate the solution found within a certain interval. The whole procedure is then strictly mathematical, although the result might never be truly exact - we can even count, Fourier notes, how close our approximations are to the real solution.

tiistai 18. helmikuuta 2020

Augustin-Louis Cauchy: Lessons on applications of infinitesimal calculus (1826)

While the previous book of Cauchy I discussed remained mostly on the level of pure mathematics, the very title of this book, Leçons sur les applications de calcul infinitésimal, promises to deal with applied mathematics. Of course, even applying can happen at different levels, and Cauchy is here dealing not with, say, application of mathematics to other sciences, but with application of one part of mathematics within another part mathematics, more precisely, in geometry.

What is applied in geometry is infinitesimal calculus, which consists of differential and integral calculus, two methods which in a sense are counterparts to one another. Yet, in a sense this is not enough, since Cauchy is actually using a variety of mathematical tools. For instance, the book begins with a long introductory section on trigonometric functions. These functions express various relations between lines and angles and can thus be used in simplifying the formulas dealt in calculus.

Another example of mathematical tools used by Cauchy is provided by polar coordinates. Unlike the regular xy-coordinates, polar coordinates express all positions through a distance between the position and the origin and the angle that the line expressing the distance forms with one of the axes. As Cauchy notes, some geometric shapes are easier to express with the polar coordinates. This is particularly true of spirals. Spirals circle in a regular fashion around a centre, which we can think as located at the origin. While the distance between the centre and a point in spiral grows, the direction of the point from the centre changes in a regular fashion.

Differential and integral calculus are still the primary methods used in the book. Originally both methods have been justified through the idea of infinitesimals or infinitely small quantities - hence, their common name, that is, infinitesimal calculus. In my previous discussion of Cauchy, I noted that by an infinitesimal calculus he meant actually a variable which was thought to be diminishing into nothing. Here, he returns to the more relaxed notion of true infinitesimal quantities, perhaps because such deep theoretical questions need not be addressed in a more applied context.

A good example of Cauchy’s tendency to ignore the theoretical questions is his rather free use of the notion of different levels of infinitesimals. It is undoubtedly difficult to understand how one infinitely small quantity can be of a different level from another infinitely small quantity, that is, larger or smaller than it. Cauchy’s theoretical account of infinitesimals is rather enlightening. Different levels of infinitesimals could be defined through different rates at which quantities approach nothing. For example, if a quantity is approaching zero, its square will also approach it, except that the square will approach the goal in a quicker fashion than the original quantity.

Of the two methods, integral calculus seems simpler in the sense that it has less areas of application - and indeed, this part of Cauchy’s book is just a fraction compared to the part dealing with differentiation. Integration in its original sense was thought to consist of dividing something (curve, area or body) into infinitely many infinitely small parts and then, as it were, adding them up. In more modern terms, integration means finding more and more fine grained divisions of curve, area and body and sums of these divisions, and examining whether these sums approach some definite limit. The result of integration is then, simply, the magnitude of the curve, area or body. In fact, the result is at worst only an assumption that the curve, area or body will have some magnitude, and for actually defining this magnitude, information on other relations between the geometrical entities is often needed.

Basic use of differential calculus, on the other hand, is to find out relations between differentials, or again in terms Cauchy used in the previous book, to find out whether relations between variable, diminishing quantities approach some definite limit. An obvious application for this is the relation between variable x- and y-coordinates of points of a curve. One begins by looking at cords connecting ever closer points of the curve and the relationship of their coordinates, which determines the direction or inclination of the cord. The limit of this approach is then a direction or inclination of a tangent moving through this one point of the curve - or, one might even say, the inclination of the curve at that point. That is, in case the tangent even exists.

The last sentence is important. Cauchy already understands that differentiation is not a universal method, but fails at some points. The curve might have a sudden change in its direction, or it might loop back and touch itself, making it impossible to say what is its inclination at this point. It is an important change in self-understanding of mathematicians to accept mathematics as not interested merely of general rules, but also of particular exceptions to these rules.

Inclination of a curve is only one of its characteristics. Indeed, a curve and its tangent have the same inclination at one point and still look quite different around that point. One would like to say that the line has less of a curve than the curve, but it needs a more precise definition to do this. Cauchy asks us to think of a circle - the bigger the circle the more the circle looks like a straight line from a given point and less curved it seems. Hence, since the circle grows with its ray, Cauchy defines the curvature of this circle as inverse of its ray.

Now, while inclinations are represented by tangents of a curve, curvature of a curve can be explained by circles touching the curve - or more precisely, the inverse of their ray - which are known as osculating circles. This is still not a proper definition of curvature, since it is not clear what circle touching the curve one is to take as the osculating circle.

Still, we can use the circle as a clue for finding out what curvature is. Picture a tangent moving through the circle, always remaining a tangent and changing its inclination as it advances through an arc of the circle. Taking smaller and smaller arcs around a point, the relation of the inclination of the tangent to length of the arc might approach a certain constant, which happens to be the inverse of the ray of the circle or its curvature.

This notion of curvature is at once applicable to other curves, since they also involve the variables of the arc and the inclination. True, the concept of curvature does not work with all curves, at least not in all points. In particular, it requires the curve in question to be twice differentiable.

Cauchy’s book deals with many other important concepts. Worth mentioning is what happens when we move from a setting with two coordinates to a case with three conditions. I won’t go into details, but note merely that then mere curvature defined in the sense above is not enough, because this first curvature deals only with a two-dimensional issue. We also need then to take into account another quantity defining the curvedness of the figure when looking at the third dimension - this curvedness Cauchy calls second curvature.

torstai 21. helmikuuta 2019

Jean-Babtiste Joseph Fourier: Analytical theory of heat (1822)

(1768-1830)

The history of the interaction of mathematics and physics has not just been one of unidirectional influence. Certainly the development of mathematics has been of great importance to physics, by providing it new and improved tools for modeling natural phenomena. Yet, physics has also offered inspiration and spur for development of new mathematical tools. The tale of Fourier’s Théorie analytique de la chaleur is of the latter sort.

Fourier’s starting point was the revolutionary use of mathematics in understanding nature, instigated by the works of Descartes and completed, in a sense, in the works of Isaac Newton. What they did was to extend the use of mathematics from mere tool for studying of figures into a tool for studying the motions of bodies. The success of Descartes and Newton inspired others to investigate whether mathematics could be useful in studying other natural phenomena.

One obvious candidate was the propagation of heat through a substance. Whatever heat was thought to be - often it was considered a distinct caloric substance that permeated all objects - it certainly appeared to “flow” through these objects, touching at first only one spot of the object in question and gradually spreading through the object and finding a point of equilibrium. The physical model was simple enough, all that was needed was to express the movement of the heat mathematically.

Fourier noted, firstly, that the movement of heat in an object was dependent on three things specific to the object, its constitution and its relation to its environment: the capacity of the object to assimilate heat (heat capacity), the capacity of its parts or molecules to transmit heat to one another (thermal conductivity) and the capacity of the environment to transmit heat to the object in question. In practice, we can limit our attention to the first two, because they form the basis of his theory of heat flow and the question of one object transmitting heat to another merely complexifies the basic theory. For simplicity’s sake, Fourier regarded these two quantities as simple constants, although the heating of an object might in reality affect them.

While physical objects are, of course, three-dimensional, we can first concentrate on the simple case of one-dimensional transfer of heat, e.g. within a barlike object. Clearly, the more distant a point in the object is from the source of the heat, the colder the point is and the opposite end from the source remains coldest. Now, Fourier supposed that the temperature of a point, at a given time, is in a sense proportional to the distance from the heat source. To put it more precisely, if at some time the temperature at the heat source is a and temperature at the opposite end of the bar is b and e is a given unit of temperature, then at a given point of the bar, with a distance z from the heat source, the temperature at that point can be calculated by subtracting z(a-b)/e from a.

Now, the temperatures a and b do not remain same, since heat is continuously flowing from one end of the bar to another and a keeps decreasing while b increases. To make the situation easier to handle, Fourier supposes that temperatures a and b are artificially kept constant, e.g. through an external heat source warming a. This means, he continues, that all the temperatures between the extremes of the bar also remain constant, that is, the temperature v always decreases while we move away from the heat source, at a rate (dv/dz) opposite to (a-b)/e. At the same time, heat is constantly flowing through the bar, and this flow, Fourier argues, is at least partially expressed by the formula (a-b)/e, that is, the greater the difference between the ends of the bar, the more heat flows from the warmer to the colder end.

If we forget the assumption of a and b being constant, we might say that (a-b)/e or −(dv/dz) partially represents the flow of heat characteristic of a bar of certain substance at a certain point of time. This expression cannot be the whole truth of the notion of flow of heat, because different substances have different capacities for conducting heat through them. Then again, he concludes, this expression together with the constant K describing the thermal conductivity of the substance describes completely the flow F of heat through a bar: F = K(a-b)/e or in terms of an infinitesimal change of temperature dv through an infinitesimally long length of bar dz, F = −K(dv/dz).

Next step in Fourier's argument is generalization of this formula to three-dimensional flow of heat. He begins by considering a prism, with one corner having the highest temperature A and heat flowing from this corner to all the other directions. Just like with the case of the bar, the further one goes from the source of hear, the less is the temperature, although now we have to account for three dimensions, when counting temperature at a given point at a given time - that is, the formula for counting the temperature of point (x, y, z) of prism looks something like A − ax − by − cz. Again, just like with the bar, by keeping the temperature constant at the limiting surfaces of the prism the temperature remains constant at all the points within the prism. By restricting then the investigation to heat flows in one dimension, heat flows in lines crossing planes perpendicular to x-, y- and z-axes will then be respectively −K(dv/dx), −K(dv/dy) and −K(dv/dz).

The final ingredient to be added to a general theory of heat is time. We again remove the assumption that the limits of a solid would have constant temperatures. Then the temperature of points within solid change as time goes on and heat flows in some manner through the solid, that is, temperature becomes a function of spatial coordinates and time: v = f(x, y, z, t). Fourier’s next move is to restrict the attention to flow through an infinitesimally small circle o at an infinitesimal instant dt, where due to extreme smallness the condition of the precious paragraph then apply. If we suppose the circle to be situated perpendicularly to z-axis, the heat flow going through it should be, Fourier says, −K(dv/dz)odt.

The notion of infinitesimals is, undoubtedly, unclear gibberish according to more modern understanding of differential and integral calculus, but even greater gibberish is to follow. Supposing the ring o to have an infinitesimal thickness, dz, we can distinguish between the heat flow coming within o and heat flow leaving o. The former is, expectedly, −K(dv/dz)odzdt, while the latter is almost the same, differing from this already infinitesimal quantity by “an even smaller infinitesimal”. The difference of the two quantities - that is, the amount of heat left within the ring after the instantaneous flow of heat - is just this type of “second-grade” infinitesimal, namely, K(d2v/dz2)odzdt. Here the expression (d2v/dz2) and its relation to (dv/dz) - rate of change of temperature, when moving through z-axis at moment dt - can be understood through an analogy with the relation of acceleration and velocity. In effect, (d2v/dz2) describes how the value of (dv/dz) itself changes, when moving through z-axis at moment dt.

Now, Fourier notes that the shape of o is not important and that we might as well take instead an infinitely small rectangle dxdy, making the flow through that rectangle, at instant dt, −Kdxdy(dz/dv)dt. Consider then an infinitely small cube of size dxdydz. The heat flow forming within that cube, at instant dt, is the sum of heat flows left within the cube, when heat flows coming in and going out from and to all three directions have been accounted for, namely K(d2v/dx2 + d2v/dy2 + d2v/dz2)dxdydzdt.

Now that the quantity of heat accumulating within an infinitesimal point of a solid at an infinitesimal instant of time has been, in a sense, determined, we can answer the question how does the temperature of that point develop over a period of time. It is not just a matter of dropping dt out from the formula, since heat and temperature are not completely same thing. Instead, we finally need the notion of heat capacity C of substance, which Fourier defines as the relation how much heat is required for increasing temperature of an object of certain weight. In order to get the required quantity of weight, we also need to take into account density D, that is, the relation how much certain volume of this substance weighs. By putting all these ingredients together, we find out that the rate of change of temperature over time, dv/dt, equals (K/(CD))(d2v/dx2 + d2v/dy2 + d2v/dz2).

What Fourier’s complex argument has provided is a position, where we can continue with purely mathematical methods. It is still unclear what the function f(x, y, z, t) determining the temperature of a point within solid at a certain time should be. We do know that the equation Fourier has found could correspond to infinitely many functions, but that certain additional conditions might be enough for determining the function. What really interests us is the method Fourier uses in solving the function from given conditions. To put it shortly, Fourier starts with an assumption that the function in question can be expressed in terms of simpler functions. To be more precise, he assumes that the function can be expressed as a sum of a possibly infinite series of trigonometric functions. The assumption happens to make sense in the context of heat transmission, because this physical process is not too erratic. In other words, the changes in the transmission can be approximately described with sums of cosines and sines. All Fourier then needs is a systematic method for determining these constituent functions, which is a simple enough task.

The idea of using sums of trigonometric functions as a way to determine heat functions was not completely novel. Yet, Fourier was the first person to assume that this method could be used in so extensive manner. While trying to solve a physical problem - how to describe movement of heat - he launched a completely new area of mathematical studies, the so-called Fourier series.

torstai 11. lokakuuta 2018

Augustin-Louis Cauchy: Course of analysis (1821)

1789-1857

Analyzing a rich text like Cauchy’s Cours d'Analyse is almost impossible. Thus, I shall instead concentrate, firstly, on general features of the book, and secondly, on some interesting peculiarities of the work.

The title of the book will probably not say much to anyone not experienced with what is nowadays called mathematical analysis - and indeed, it is not clear whether Cauchy’s intentions coincide exactly with this modern notion. Hence, instead of trying to define analysis at this context, I shall merely point out what appears to be the topic of this book, namely, functions.

When Cauchy speaks of functions, he seems to be generalising from individual mathematical operations, such as addition, multiplication, logarithm and sine - in a quite general manner we could say functions involve taking number or numbers and using them to calculate or determine other numbers. Unlike in the current notion of function, Cauchy allows that functions might have several alternative results. For instance, Cauchy notes that although we usually restrict the notion of square root to positive square roots, we might as well take also a negative square root as being one possible result of applying the function of square root.

We can point out two major questions involving functions that interest Cauchy. Firstly, Cauchy, like other mathematicians of the time, is interested of the general question of finding what could be called roots of a function. The function in question will take one number and produce another number as its result, and the root of this function is then a number that when applied to the function will produce zero as a result.

Secondly, Cauchy is interested of what happens when functions are applied to infinitely large or infinitely small quantities. These terms are a remnant from an earlier period of mathematics, when the unclear notion of an infinitely small number was used in making sense of results in differential calculus. With Cauchy, these notions are inevitably connected with the concept of a variable: when we speak of an infinitely great number, we mean a quantity that is meant to change its value by becoming larger and larger and eventually exceeding every finite number. Similarly, an infinitely small number is a quantity that is meant to become smaller and smaller, eventually diminishing beyond all positive numbers, without ever reaching zero.

Now, with the aid of the notion of infinitely small or infinitesimal number, Cauchy defines the concept of a continuity of function. His notion of continuity has something to do with a formula f(x + a) - f(x), where f refers to a function of some type, what is in () is the number applied to the function f, and a is an infinitesimal number. In effect, Cauchy wants us to think of the area where the results of a function lie, when the numbers applied to the function become less and less varied. If the function is continuous around certain number, Cauchy says, this means that the area of results will also be infinitesimal - in other words, the less variation we have with numbers applied to function, the less variation we have with the results.

Although Cauchy’s definition can be used with all numbers applied to functions, he is especially interested of cases where the numbers themselves are either infinitely small or infinitely great, that is, when we think of them as variables decreasing toward zero or increasing without any limit. This operation of finding a limit for a function appears to be yet another function, and just like with all the other functions, Cauchy accepts the possibility of a function having more than one result. For instance, the values of sine function vary constantly between -1 and 1, thus, when the numbers applied to the function increase indefinitely large, the limit of this function, Cauchy says, consists of all the numbers between -1 and 1.

While differential calculus had been the original spur for mathematician’s developing the notions of continuity and limit of a function, for Cauchy this is already just a one possible application of these notions, while other applications, such as the question of an unending series of sums of numbers, might be even more important.

An important application of the notions concerns the other interest mentioned earlier, that is, the question of roots of a function. Cauchy notes that if a function is continuous, whenever numbers are taken from some continuous part of number line, and the same function has a positive result with some number from that part and a negative result with another number from the same part, then the function has a root somewhere between those numbers. This sounds evident, but it requires some careful thinking to actually prove it. In effect, Cauchy notes that due to the continuity of the function in this area, we can find between the numbers described both numbers giving negative and numbers giving positive results, as close to one another as we like. These two series of numbers must have the same limit, which then can have neither positive nor negative result, that is, it must be the root of the function.

Another major theme Cauchy deals with in his book is imaginary - or as we would nowadays say, complex - numbers. Discussions thus far have been grounded in actual numeric operations that make some concrete sense. Thus, negative numbers can be understood as a simple way to speak about operations of subtraction (i.e. - a is just a summarised form of saying “subtract a”), raising number to a fractional expression 1/a is another way of saying that we are taking the ath root of that number, and any difficult calculation involving numbers expressible only as limits of certain number series (such as raising a number to pith power) refer to limits of functions, when applying them to numbers from that series.

Now, when Cauchy starts to speak of square roots of negative numbers, he doesn't really mention any means to make such a mathematical formula sensible (he has the means by which it could be done in his use, but that’s another story). Cauchy then accepts that what mathematicians are speaking about when dealing with square roots of negative numbers is just imaginary of fictional - the signs apparently referring to such roots mean nothing. Still, majority of the operations used in the context of real numbers work as well in the context of these imaginary numbers. Since these imaginary numbers can be used as tools for finding meaningful results to questions involving just real numbers, the interest for this part of the mathematical analysis is then also justified.