- Tensors are invariant; their components transform — upper indices by B, lower by F (F, B inverse).
- Vectors (contravariant, upper) and covectors (covariant, lower) are the building blocks; a covector is a stack of lines and
α(v)counts piercings. - Combine p vectors and q covectors with
⊗to build any (p,q)-tensor; transform each factor and the rules fall out for free. - The metric
g_ij = e_i·e_jmeasures length/angle and rescues Pythagoras outside orthonormal bases. glowers (♭) andg⁻¹raises (♯) indices, linking vectors and covectors — butv^i ≠ v_iunless the basis is orthonormal.
What is a tensor: Imagine a car heading on a straight highway, towards a left turn that it wants to take. Imagine that the distance of the car is measured along the highway as x, starting from the left turn and increasing towards the car's present position at time 0. Imagine the y-axis as starting from the left turn's join with the highway and increasing along the highway's width towards the other end of the highway.
Now you could make predictions about the speed of this car knowing that it wants to take the left turn. You could say, as the car's x approaches 0, it's y will also approach 0. It's speed along x will go to 0 while the speed along y will increase in magnitude. You could maybe make a function that takes the coordinate of the car in x and y and spits out a number: the magnitude of its speed. That is a tensor (Let's assume that the car's speed/velocity varies linearly with x and y, since tensors must be multilinear)! It took two vectors, x and y, and returned a scalar: the speed. You could even ask it to spit out two vectors, velocity along x and y, or a single combined velocity vector. These are other tensors. In fact, a lot of relativity involves looking at how some quantity in different coordinate directions (including time) change with the values of some other quantities along different coordinate directions.
Tensors are written as (m,n) tensors, which is m+n-order tensor, which is also the total number of indices: m upper indices and n lower indices. Another way of looking at it is the tensor is the tensor product of m vectors and n covectors. Which means you can feed it n vectors and m covectors to get a scalar, or n vectors to get a m-vector space etc.
- Index placement rule: components from V go upstairs, components from V* go downstairs.
A (m,n) tensor's components are written as a row of row of row... m times, a column of column of column ... n times (the rows and columns might be mixed)
Two players run through everything:
- F — the forward matrix. Writes the new basis vectors as combinations of the old:
ẽ_j = F^i_j e_i. - B — the backward matrix. Writes the old basis vectors as combinations of the new:
e_i = B^j_i ẽ_j. - F and B are inverses:
F B = B F = δ(the Kronecker delta / identity,δ^i_j = 1ifi=j, else0). - A (m,n) tensor has m upper indices and n lower indices and transforms with m B's and n F's.
When you change basis, things split into two camps depending on which direction they transform:
- Contravariant = transforms opposite to the basis vectors → written with an UPPER index → uses B to go old→new.
- Covariant = transforms the same way as the basis vectors → written with a LOWER index → uses F to go old→new.
Intuition: if you double the length of your basis vectors, the same fixed vector now needs half as many of them → its components shrink (contravariant). If you double the basis, a covector's components double (covariant). Things that shrink when the ruler grows are contravariant; things that grow with the ruler are covariant.
Mnemonic: Upper index ⇒ uses B (backward i.e. backward of vector basis) ⇒ contravariant. Lower index ⇒ uses F (forward) ⇒ covariant. Going new→old just swaps F↔B.
A vector is a member of a vector space: a set of objects you can add and scale by scalars. The vector is invariant; only its component list [v^1, v^2, …] depends on the basis.
- A vector is built from basis vectors:
v = v^i e_i(Einstein sum). - Vector components are contravariant (upper index, transform with B).
Einstein summation convention (used everywhere): when an index appears once up and once down in a term, sum over it — drop the Σ. Real sums always pair one upper with one lower. The Kronecker delta cancels/renames a summed index: M^i_j δ^j_k = M^i_k.
A covector is a linear function that eats a vector and returns a scalar: α(v) ∈ ℝ, with α(v) = α_i v^i. Linearity: α(v+w) = α(v)+α(w) and α(nv) = n·α(v). Covectors live in their own vector space, the dual space V* (a row "vector" is just a covector's component list).
Picture a covector as a stack of lines:
Treat α = (2,1) as the function f(x,y) = 2x + 1y. Draw its level sets (f = 0, 1, 2, …): a family of equally spaced parallel lines, like a contour map, with an orientation arrow pointing the way the value increases.
- Measurement = counting pierced lines.
α(v)= the number of stack-lines the arrowvcrosses. Pierces 2 → value 2. Reaches halfway to the first line → 0.5. Runs parallel to the lines → 0. Points against the orientation → negative. - Denser lines = "bigger" covector. To compute
2α, make the lines twice as dense (NOT spread out):vthen crosses twice as many →(2α)(v) = 2·α(v). - Adding covectors: lay both stacks down; a vector pierces the combined lines, so
(α+β)(v) = α(v) + β(v).
Standard vectors (contravariant vectors) are written as columns, while covectors are written as rows.
Dual basis ε^i: defined by ε^i(e_j) = δ^i_j. The ε^i "read off" components: ε^i(v) = v^i. Any covector expands as α = α_i ε^i, where α_i = α(e_i). Basis covectors are contravariant (use B); covector components are covariant (use F) — opposite directions, which is the whole point of upper-vs-lower bookkeeping.
A linear map eats a vector and returns a vector, obeying L(v+w)=L(v)+L(w), L(av)=aL(v). Its component form is a matrix.
-
Geometrically (in words): a transformation that keeps grid lines parallel and evenly spaced and fixes the origin — stretches, rotations, shears. Translations are NOT linear (they move the origin).
-
Reading the matrix: column
jtells you where basis vectore_jis sent.L(e_j) = L^i_j e_i. -
A linear map transforms vectors, NOT the basis — input and output are measured in the same basis.
-
Action:
(Lv)^i = L^i_j v^j(ordinary matrix × column). -
Transformation rule:
L̃ = B L F, i.e.L̃^a_b = B^a_i L^i_j F^j_b. One B (upper index) + one F (lower index) ⇒ a (1,1)-tensor.Why two matrices, in words: to get a new-basis output from new-basis input, detour: F (new→old input), then L (old in→old out), then B (old→new output).
Problem: Pythagoras |v|² = (v^1)² + (v^2)² only works in an orthonormal basis (unit-length, perpendicular basis vectors). In a skewed/scaled basis it gives the wrong, non-invariant length.
Fix: length comes from the dot product, |v|² = v·v. Expanding v = v^i e_i:
|v|² = v^i v^j g_ij, whereg_ij = e_i · e_j= the metric tensor.
- It is the matrix of dot products of the basis vectors; as a sandwich,
|v|² = vᵀ g v. - Orthonormal basis ⇒
g_ij = δ_ij(identity) ⇒ Pythagoras, which is just the special case. - Symmetric:
g_ij = g_ji(dot product ignores order). - Angles too:
v·w = |v||w|cos θ, socos θ = (v^i w^j g_ij) / (|v||w|). (For unit basis vectors at angle θ,ẽ_1·ẽ_2 = cos θ.) - It's a (0,2)-tensor: both indices lower, transforms with two F's:
g̃_ij = F^k_i F^l_j g_kl. When the basis grows, theg_ijshrink to match, so|v|²stays invariant in every coordinate system.
A bilinear form B(v,w) eats two vectors, returns a scalar, and is linear in each slot separately:
- Scale one slot:
a·B(v,w) = B(av,w) = B(v,aw). (Scaling both givesa²— watch out.) - Distribute each slot:
B(v₁+v₂, w) = B(v₁,w)+B(v₂,w), same for the 2nd slot. Both slots summed ⇒ four terms. - Components
B_ij = B(e_i, e_j); outputB(v,w) = B_ij v^i w^j(row × matrix × column). - (0,2)-tensor, fully covariant:
B̃_kl = F^i_k F^j_l B_ij.
The metric is a bilinear form with two extra properties: (1) symmetric g_ij = g_ji, and (2) positive-definite g(v,v) ≥ 0 (lengths can't be negative). So: metric tensors ⊂ bilinear forms (a non-symmetric matrix, or a symmetric one that outputs a negative for some v, is a valid bilinear form but not a metric).
Stop defining tensors by their transformation rules. Build them from vectors and covectors using the tensor product ⊗ — then every rule falls out for free.
- Linear map = vector ⊗ covector:
L = L^i_j (e_i ⊗ ε^j). (A single paire_i ⊗ ε^jis a "pure" rank-1 map sending everything to one direction; interesting maps are sums of pure pairs.) - Bilinear form = covector ⊗ covector:
B = B_ij (ε^i ⊗ ε^j). (Two covectors because it eats two vectors — each covector consumes one.)
Why this is powerful — three freebies from transforming each factor on its own (vectors with B, covectors with F, then pull the matrices out front):
- Transformation rule appears automatically (one B + one F for a linear map; two F's for a bilinear form).
- The multiplication formula appears automatically — via
ε^j(e_k) = δ^j_kcollapsing indices: e.g.L(v) = L^i_j v^j e_i,B(v,w) = B_ij v^i w^j. - The array shape appears automatically (linear map = "row of columns"; bilinear form = "row of rows," which lets you honestly write both input vectors as columns).
- ⊗ has three meanings (same idea, different level): Kronecker product (combine arrays — distribute the left array into each entry of the right), tensor product (combine tensors), tensor product of spaces (combine
VandV\*into a bigger space likeV⊗V\*). The components from the tensor product are exactly the entries from the Kronecker product. - For big tensors, use index notation, not "Q(D)". Once there are several slots, "Q acting on D" is ambiguous — there are many valid contractions. Einstein notation states exactly which indices are summed, so it's the unambiguous language.
Deepest definition: a tensor is a multilinear map — feed it vectors/covectors and it's linear in each slot separately. A (1,1)-tensor, depending on what you feed it, can act as a map V→V, a map V*→V*, or a scalar-valued function — same object, different contractions. Any contraction is allowed as long as each summed pair is one-upper-with-one-lower. The number of leftover upper/lower free indices tells you what space the output lives in.
Question: is there a basis-independent partner covector for each vector? The naive pairing e_i ↦ ε^i fails: basis vectors grow with F, basis covectors shrink with B, so the pairing falls apart under a change of basis.
Correct partner: pair v with the covector v·( ) — "dot v with whatever comes in." This is genuinely linear, references no basis, and scales with v. Its components are gotten by feeding v into one slot of the metric:
- Lower (vector → covector), the flat
♭:v_j = g_ij v^i - Raise (covector → vector), the sharp
♯:v^k = g^{kj} v_j
where g^{ij} is the inverse metric, defined by g^{ik} g_kj = δ^i_j (a (2,0)-tensor).
- Mnemonic: musical flat ♭ lowers (pointy arrow → flat stack of lines); sharp ♯ raises (flat stack → pointy arrow). They're inverse operations.
- ⚠
v^i ≠ v_i. Upstairs and downstairs components are different objects; you must contract withg(org⁻¹) to convert. They coincide only in an orthonormal basis whereg_ij = δ_ij. - This works on any index of any tensor: contract a slot with
gto lower it (turns aVfactor into aV\*factor) or withg⁻¹to raise it. The metric stitches all the(p,q)spaces together into one family.
- Type signatures: grad ⟶ Scalar to Vector , div ⟶ V to S , curl ⟶ V to V , Laplacian ⟶ S to S
- The Gradient points in the direction of fastest increase; Divergence measures net expansion or compression (sources and sinks); and Curl measures local swirling or rotation.
-
$\text{curl}(\text{grad } f) = 0$ and$\text{div}(\text{curl } \mathbf{F}) = 0$ . Consequently, all gradient fields are perfectly irrotational (no swirl), and all curl fields are perfectly incompressible (no divergence). - Stokes' Theorem dictates that a gradient field produces exactly zero circulation around any closed loop
$C$ bounding a surface$S$ :$$\oint_C \nabla f \cdot d\mathbf{r} = \iint_S \nabla \times (\nabla f) \cdot \mathbf{n} , dS = 0$$ - Divergence Theorem dictates that a curl field produces exactly zero net flux through any closed surface
$S$ enclosing a volume$V$ :$$\iint_S (\nabla \times \mathbf{F}) \cdot \mathbf{n} , dS = \iiint_V \nabla \cdot (\nabla \times \mathbf{F}) , dV = 0$$ - You derive fundamental physics PDEs by counting a quantity, tracking its boundary flux, and applying Gauss's theorem to drop the integrals. This yields conservation laws like the continuity equation (
$\frac{\partial\rho}{\partial t} + \nabla \cdot (\rho \mathbf{u}) = 0$ ). Under ideal conditions (steady, incompressible, irrotational), this simplifies entirely to Laplace's equation ($\nabla^2\varphi = 0$ ).
A scalar field f(x, t) assigns a number to every point (e.g. temperature). A vector field u(x, t) assigns a vector to every point in space (e.g. fluid velocity in the Gulf of Mexico). The field is the solution of a PDE; dropping a particle into the field gives an induced ODE dx/dt = u(x, t) (e.g. tracking an oil spill). Taking the gradient of a scalar field gives a vector field. Taking the divergence of a vector field gives a scalar field. Taking the curl of a vector field gives a vector field.
∇ is a vector of partial derivatives — treat it like a real vector:
∇ = ( ∂/∂x , ∂/∂y , ∂/∂z )
Three products with it give the three core operations, plus a fourth built from them. Memorize the type signatures — they are the whole point:
| Operation | Written | Input → Output | Formula |
|---|---|---|---|
| Gradient | grad f = ∇f |
scalar → vector | (∂f/∂x, ∂f/∂y, ∂f/∂z) |
| Divergence | div F = ∇·F |
vector → scalar | ∂Fx/∂x + ∂Fy/∂y + ∂Fz/∂z |
| Curl | curl F = ∇×F |
vector → vector | <∂F₃/∂y − ∂F₂/∂z , ∂F₁/∂z − ∂F₃/∂x , ∂F₂/∂x − ∂F₁/∂y> (take the determinant) |
| Laplacian | ∇²f = ∇·(∇f) |
scalar → scalar | ∂²f/∂x² + ∂²f/∂y² + ∂²f/∂z² |
All four are linear operators i.e. ∇(af₁ + bf₂) = a∇f₁ + b∇f₂.
∇f at each points is oriented in the direction f (the scalar field) increases fastest; its magnitude is the rate of increase.
- Example:
f = x² + y²→∇f = (2x, 2y). Arrows point radially out, growing with radius — matches the paraboloid getting steeper outward. - Directional derivative (rate of change of
falong a chosen directionv):
NormalizeD_v f = (v / |v|) · ∇fv. Ifv ⊥ ∇f, the directional derivative is 0 (move that way andfdoesn't change).
∇·F > 0: field is sourcing/spreading out (faucet hitting the sink basin).
∇·F < 0: field is sinking/converging (the drain).
∇·F = 0: divergence-free / incompressible — volumes are neither created nor destroyed.
Intuition: drop a little blob of dye into the flow; positive divergence makes the blob grow, negative makes it shrink, zero keeps its area fixed (it might deform, but doesn't expand or contract).
F = (x, y)→∇·F = 1 + 1 = 2(diverging).F = (−x, −y)→∇·F = −2(converging).F = (−y, x)→∇·F = 0(pure rotation, divergence-free).
curl F measures how much the field swirls around a point. In 3D:
∇×F = ( ∂F₃/∂y − ∂F₂/∂z , ∂F₁/∂z − ∂F₃/∂x , ∂F₂/∂x − ∂F₁/∂y )
In 2D (F = (F₁, F₂, 0)) only the z-component survives — curl is effectively a scalar pointing out of the page:
(∇×F)_z = ∂F₂/∂x − ∂F₁/∂y
F = (x, y)→ curl= 0→ irrotational (radial, no swirl).F = (−y, x)→ curl= 2→ rotational (this is the divergence-free swirl from §3).- Sign by right-hand rule: thumb out of the page = positive curl.
Solid-body rotation: a body spinning with angular-velocity vector ω has velocity field v = ω × r. Then curl v = 2ω — curl is twice the angular velocity. A constant curl ⇒ the blob rotates without deforming.
curl(grad f) = 0 for every scalar f
div(curl F) = 0 for every vector field F
Consequences used constantly later:
- Any gradient field (
F = ∇f) is automatically curl-free / irrotational. → So a field with any curl cannot be the gradient of a scalar potential. - Any curl field (
F = ∇×G) is automatically divergence-free. Taking the curl of a flow extracts only its rotational part.
Scalars need 1 function; vector fields need 3 (F₁, F₂, F₃) — a strictly bigger space. So a generic field is not ∇ of anything.
- Conservative / gradient field:
F = ∇ffor some scalarf. These are exactly the curl-free (irrotational) fields (by §5).∂F₃/∂y − ∂F₂/∂z = 0(for a 2D plane) - Potential flow is the even more special case: a gradient field that is also incompressible → it satisfies Laplace's equation (§9). So: potential flow ⊂ gradient flows.
These convert between what happens inside a region and what happens on its boundary. They are the bridge from "rate of change in a blob" to a PDE.
Flux out through a closed surface = integral of divergence over the enclosed volume.
∯_S F · n dS = ∭_V (∇·F) dV
F·n = how much of the field pierces the surface normal-wise (tangential flow carries nothing out). Why it's true: chop the volume into tiny boxes; for continuous F, every internal wall's outflow cancels its neighbor's inflow, leaving only the outer perimeter = the surface flux. Requires F continuous — fails at shock waves / discontinuities.
Stokes: circulation around a closed loop = flux of curl through any surface it bounds:
∮_C F · dr = ∬_S (∇×F) · n dS
Green is the 2D special case:
∮_C (F₁ dx + F₂ dy) = ∬_R ( ∂F₂/∂x − ∂F₁/∂y ) dA
Pairing to remember: Gauss ↔ divergence ↔ flux through surface; Stokes ↔ curl ↔ circulation around loop.
This single pattern derives the continuity equation, Navier–Stokes, the heat equation, etc.
- Count the stuff. Mass in a volume
V:M = ∭_V ρ dV(ρ = density). - Accounting. Its rate of change = what flows out through the boundary (+ any source
Q):
(Minus sign: positive outward flux means mass decreasing inside.)d/dt ∭_V ρ dV = − ∯_S (ρ u) · n dS ( + ∭_V Q dV ) - Gauss's theorem turns the surface integral into a volume integral → everything under one
∭_V. - "True for all volumes." If the integral is zero for every volume, the integrand must be zero everywhere. Drop the integral → the PDE.
Result — the Mass Continuity Equation:
∂ρ/∂t + ∇·(ρ u) = 0
- Incompressible flow (ρ = constant ⇒
∂ρ/∂t = 0,∇ρ = 0) collapses this to:∇·u = 0 (velocity field must be divergence-free) - Source terms
Q(nuclear blast converting mass→energy, etc.) put a nonzero right-hand side → a forced PDE. - Caveat: Step 4 needs
ρandusmooth. Shock waves / discontinuities break it — derivatives blow up — and you must stay in the integral "control-volume" form. - Same recipe on momentum → Navier–Stokes; on energy → the energy/heat equation.
Potential flow = steady + incompressible + irrotational flow. Build it from a scalar potential φ via v = ∇φ:
- Irrotational is automatic (
curl ∇φ = 0, §5). - Steady is automatic if
φdoesn't depend on time. - Incompressible is the one real condition:
∇·v = ∇·(∇φ) = 0, i.e.
∇²φ = 0 ← LAPLACE'S EQUATION
Solutions are called harmonic functions. Laplace's equation is everywhere:
- Steady-state heat: the heat equation
∂T/∂t = α²∇²Twith∂T/∂t = 0becomes∇²T = 0. - Gravitational & electrostatic potentials away from masses/charges.
- Potential flows (early airfoil design).
Poisson's equation = Laplace with a source (force it and it becomes Poisson):
∇²φ = f (e.g. ∇²φ = −ρ/ε in electrostatics)
f is a source/forcing term (heat source, charge density). Solving Poisson is a core, expensive step inside Navier–Stokes solvers (pressure).
Boundary conditions are required for a unique solution: Dirichlet (value fixed on boundary, e.g. edge temperatures) or Neumann (normal-derivative fixed, e.g. insulation/flux).
Any analytic complex function F(z) = φ(x,y) + i ψ(x,y), z = x+iy (e.g. z², eᶻ, log z, sin z) hands you a potential flow for free:
- Real part
φ= velocity potential; imaginary partψ= stream function; both are harmonic (∇²φ = ∇²ψ = 0). - Cauchy–Riemann equations:
φ_x = ψ_y,φ_y = −ψ_x. - Velocity two ways:
v = (φ_x, φ_y) = (ψ_y, −ψ_x). ψ = consttraces streamlines; equipotential lines (φ = const) are orthogonal to streamlines.- Worked example
z² = (x²−y²) + i(2xy):φ = x²−y²,ψ = 2xy,v = (2x, −2y). Check:∇·v = 2−2 = 0✓, curl= 0✓. A saddle: stretch in x, squeeze in y, area preserved. - Superposition of elementary analytic flows (uniform stream + sources/sinks + vortices) builds flow over airfoils — the pre-computer design method.