The situation
I’ve always had a love-hate relationship with notation in mathematics.
On the one hand, I am an intuitive person who needs to understand the big picture of something before jumping into the details. In this sense, I prefer my math to be uncluttered and focusing on the key steps of a derivation.
On the other hand, I also need my understanding to be precise to feel that I really understand something. That means that the whole procedure must be logical, and everything somehow needs to come together in a meaningful structure. Because of that, I sometimes prefer a more cluttered — but more precise — notation.
Almost 10 years ago, I started learning about Lie derivatives, a beautiful concept in differential geometry that generalizes the notion of taking a small variation of a quantity along a path. It is the key behind concepts like Killing vectors which make the notion of symmetries in physical systems more precise. My first contact with Lie derivatives was through Wald’s General Relativity book and Stewart’s Advanced General Relativity.
Being much less enlightened than both Wald and Stewart, I always had some difficulties with two choices they took:
-
The definition of the Lie derivative is clear, but it takes quite a bit of mental effort to connect all dots and see how the definition makes sense; I would love for the definition to be more transparent, almost algorithmic;
-
The proof of the main result, namely that the Lie derivative of a vector along another equals their commutator, always involves choosing one of them as a coordinate basis vector; I would love a more “general”, albeit more involved version, to work without this requirement.
Hence, I am posting here for any other students of differential geometry or general relativity my own proof which is, cluttered as can be, very transparent (for me, at least!).
Set-up: pushforwards and pullbacks
Below, is a 1-parameter family of diffeomorphisms induced by a vector field , also called ‘s orbits; this means that is the solution to the following initial value problem, where is local coordinate chart centered around a point , which we will fix throughout the whole discussion:
where is the value of the components of the vector field at point .
Provided is smooth, the solution exists and is unique at least on a neighborhood of ; we will assume to be on such a region during the whole discussion below. Notice that what does is to map points on the manifold along the direction of the vector field ; in particular, maps the origin onto a point at a parameter distance along the orbits of .
First, we will present the rules that allow for the quick rule-of-thumb
where is some tensor.
It is crucial here that we deal with diffeomorphisms (which are, in particular, invertible); otherwise we could not take inverses and this argument wouldn’t work.
Pullbacks of functions
Given a map , define a function’s pullback as
This means that the pullback applied to point will be equivalent to applied to :
Pullback of function = function composition.
Pushforward of vectors
First, recall that in differential geometry one often thinks of vectors as differential operators: let be the tangent space to a point in the manifold. There exists a coordinate chart (where ranges from 1 to the dimension of the manifold) such that is the span of . Any vector is then of the form
and operations such as are well-defined, where is a function defined on the manifold.
Now, let . We define a new vector , its pushforward, via its action on functions:
That is, the pushforward of a vector, applied to a function, is the vector applied to that function’s pullback! (confusing, I know)
(Pushforward of vector)(function) = Vector(pullback of function).
Pullbacks of 1-forms
This line of thought can be generalized to covectors, and posteriorly to tensors. Pullback of 1-form, applied to vector = 1-form applied to the pushforward of that vector:
(Pullback of 1-form)(vector) = (1-form)(pushforward of vector).
Pushforwards of everything
Identifying
we can define pushforwards for any tensors. For example, let be a (1,1)-tensor in abstract index notation. I’ll write as its explicit action on a one-form and a vector. It follows that
All of the quantities in the RHS are well-defined (pushforward on vector, pullback on 1-form).
Lie derivatives
A big chunk of what we do below follows this reference.
First, we write for a vector field evaluated at a point : this notation will make explicit at which point where the tensor is defined.
Also, for pushforwards, we write down a subscript showing where the pushforward “acts on”. For example: if , then the pushforward will be written with a subscript as well:
In this notation, we also make explicit the pullback, which “departs” from :
Then, we define the Lie derivative once and for all as
This also fits Carroll’s formula B.4-B.5 (but he uses and an opposite convention for the asterisk position), and Isham’s formula 3.2.21. We also made explicit the origin of the pushforward: it starts from .
Does this definition make sense? The tensor is explicitly defined at point ; the other term, , is not, but it is “brought back” to via the pullback operation, which is “based” on and drags the tensor back by a parameter value of , effectively arriving at . All is well-defined.
Example of applying this formula
Let’s do this calculation for a function. As a tensor, it does not vary from point to point, even though its value when calculated at different points does. What I mean is that for any two points , but usually . As such, using the equation above, we can use the pullback-based expression , hence
To calculate the expression on the RHS, go to a local coordinate chart where has coordinates and where are the local coordinates of the vector field generating . Then, Taylor expanding, we get a RHS of which is just calculated at . It follows that
Calculating the Lie derivative of a vector by hand
Well then; what about acting on vectors? The spoiler is:
let’s see if this holds. Nothing like a nice manual calculation before we give the more general formula.
Let and consider a base point . Let
These are two vector fields. At every point of the manifold, they take values - for example, at ,
where the subscript doesn’t mean much in this particular case for the basis vectors (since we are in Euclidean space), but we keep it nonetheless.
Assume we want to compute . We need to first solve the equation for the flow of , that is, ; the differential equation for integral curves is
with solution
This means that the map based on which takes it downstream by a parameter value is
Notice that in this particular case is linear - this won’t always be the case. The pushforward is obtained by calculating the Jacobian w.r.t the base coordinates of (which justifies the often-used notation for the pushforward). Hence
and, for the opposite direction,
Notice how the sign has changed with the formal replacement of . Also, we changed the subscript from to for clarity.
Note: an important point here is the use of matrices. In the expression for , we used matrices as a compact notation of how coordinate systems changed. For the pushforward, however, the matrices actually act on vectors in one space and give vectors in another space. Explicitly,
We have all the ingredients we need. We want to calculate
First, we calculate . It is simply
Then
Finally, plugging back into the definition of the Lie derivative, we get
and we are done.
The intuition here is as follows: we are trying to calculate the derivative of the vector (field) along the integral curves of , which are circles. Notice how there is a difference exactly on the vertical direction between and its pullback, which is slightly inclined due to rotation.
By the way: since we already know the spoiler that , it is worth checking if it works. Recalling that
we have, for and , that the only non-zero component of the commutator is
yielding , as we obtained.
Proof of the commutator formula
First, notation. The field will be written as
where are the coordinates of point . Let us write the coordinates at as , generically, and those at as , ie. with a tilde.
The first important thing is to relate the coordinates of with those of , in the limit where is small and can be kept to first order. By definition,
Notice that we explicitly wrote as a functon of the coordinates at - this will come as important later in the calculation of the Jacobian for the pushforward.
With that being done, we follow the algorithm and calculate at point , in the limit where is small. For this, we use equation above:
where we Taylor-expanded in the third line, and cleaned up notation a bit in the fourth line.
To calculate the pushforward, we need to calculate the Jacobian of with respect to the coordinates . We can again use equation here; it is easy to see that
Inverting this yields, to first order,
We are ready: the Lie deriative can be calculated as
Thus, we have the proof!
Also check this reference here for a proof without using Taylor series.