The derivative is nilpotent, and the shift is its exponential
Differentiate a quintic six times and it is gone. Exponentiate that same dying operator and you get the shift by h, invertible forever. The sentence connecting them is Taylor's theorem.
Differentiation lowers degree by one. Do it times to a polynomial of degree at most and there is nothing left. On the space the operator satisfies exactly, not approximately, and an operator with that property is called nilpotent.
In the monomial basis is a single stripe. Since , the matrix has and zeros everywhere else — strictly above the diagonal, nothing on it. Raising it to a power slides the stripe one step further out each time, and at the sixth power the stripe has slid off the edge of the matrix. That is all nilpotency is: a stripe with somewhere to fall off.
Turn the slider up and watch the left matrix empty itself. There is no drama, no eigenvalue crossing zero, no critical parameter. Every eigenvalue of was zero from the beginning — the diagonal is all zeros and is triangular, so its characteristic polynomial is . A nilpotent operator is one whose entire spectrum is and which is nonetheless not the zero operator, and it is the standard example of a matrix that cannot be diagonalized. If it could, it would be conjugate to the diagonal matrix of its eigenvalues, which is , which would make itself zero, and constants are not the only polynomials.
Now build a different operator out of the same stripe. Take its exponential:
Ordinarily the exponential of an operator is a convergence question. Here it is not, because after the sixth term every summand is zero. The series is a finite sum, exact, for every real , with no radius of convergence to worry about. The right-hand matrix in the demo is that sum, computed term by term, and its entries are — Pascal’s triangle scaled by powers of , which is precisely the matrix that the substitution article got out of the binomial theorem. The readout prints the largest disagreement between the two, and it is around , which is to say they are the same matrix and the difference is that the computer has to round.
So is the shift, . Written out, that identity is
which is Taylor’s theorem. For polynomials the Taylor series is not an approximation and there is no remainder term; the sum simply stops. The usual statement of Taylor’s theorem, with its awkward remainder and its hypotheses about continuous derivatives, is what you get when you try to write this same identity for functions on which is not nilpotent, and pay for the privilege.
The two matrices in the demo are both triangular, and the difference between them is entirely in the diagonal. has zeros there; it is nilpotent, it is not invertible, and it drops every polynomial one rung down the flag until they fall out the bottom. has ones there; it is unipotent — identity plus nilpotent — its determinant is for every , it is invertible with inverse , and it preserves degree exactly. The exponential map takes the first kind to the second kind. It takes an operator that destroys the grading to an operator that respects it, and it takes the additive structure to the multiplicative one, , which is just the observation that shifting by and then by shifts by .
This is a Lie algebra and its Lie group, at the smallest scale where the words mean anything. The nilpotent is an infinitesimal generator; the one-parameter family is the flow it generates; and the flow is the translation group of the line acting on functions. Differentiation is translation, infinitesimally, and the sentence “the derivative generates translations” — which in quantum mechanics gets written and called the momentum operator being the generator of spatial translation — is on a statement about a matrix with five entries in it.
There is a companion generator worth meeting. Let , the Euler operator, . It is diagonal, the monomials are its eigenvectors, and its eigenvalue on is the degree itself. Exponentiating it gives , so with the operator is exactly the scaling map from the substitution matrix — the diagonal factor , whose diagonal was . Degree is not merely an index into a basis, then. It is an eigenvalue: the eigenvalue of the Euler operator, the number that says how the polynomial responds to a rescaling of its input.
The two generators do not commute. Computing on gives , so
Two generators, one bracket relation, and that is the entire Lie algebra of the affine group of the line — the group of maps we have now met three times. Scaling and translation, degree and derivative. The relation says that conjugating a translation by a scaling gives you back a translation with its step size scaled, which you already knew from writing .
An operator that lowers degree, and its exponential, which cannot. Next, an operator that would like to raise the degree above and has no room: what happens when you divide.