Skip to content

Instantly share code, notes, and snippets.

@regstuff
Last active July 21, 2026 13:14
Show Gist options
  • Select an option

  • Save regstuff/33c509a4ad36999819f019e45ec6ed80 to your computer and use it in GitHub Desktop.

Select an option

Save regstuff/33c509a4ad36999819f019e45ec6ed80 to your computer and use it in GitHub Desktop.
Tensors and Vector Calculus

Tensors

The whole story in five lines

  1. Tensors are invariant; their components transform — upper indices by B, lower by F (F, B inverse).
  2. Vectors (contravariant, upper) and covectors (covariant, lower) are the building blocks; a covector is a stack of lines and α(v) counts piercings.
  3. Combine p vectors and q covectors with ⊗ to build any (p,q)-tensor; transform each factor and the rules fall out for free.
  4. The metric g_ij = e_i·e_j measures length/angle and rescues Pythagoras outside orthonormal bases.
  5. g lowers (♭) and g⁻¹ raises (♯) indices, linking vectors and covectors — but v^i ≠ v_i unless the basis is orthonormal.

What is a tensor: Imagine a car heading on a straight highway, towards a left turn that it wants to take. Imagine that the distance of the car is measured along the highway as x, starting from the left turn and increasing towards the car's present position at time 0. Imagine the y-axis as starting from the left turn's join with the highway and increasing along the highway's width towards the other end of the highway.

Now you could make predictions about the speed of this car knowing that it wants to take the left turn. You could say, as the car's x approaches 0, it's y will also approach 0. It's speed along x will go to 0 while the speed along y will increase in magnitude. You could maybe make a function that takes the coordinate of the car in x and y and spits out a number: the magnitude of its speed. That is a tensor (Let's assume that the car's speed/velocity varies linearly with x and y, since tensors must be multilinear)! It took two vectors, x and y, and returned a scalar: the speed. You could even ask it to spit out two vectors, velocity along x and y, or a single combined velocity vector. These are other tensors. In fact, a lot of relativity involves looking at how some quantity in different coordinate directions (including time) change with the values of some other quantities along different coordinate directions.

Tensor Notation

Tensors are written as (m,n) tensors, which is m+n-order tensor, which is also the total number of indices: m upper indices and n lower indices. Another way of looking at it is the tensor is the tensor product of m vectors and n covectors. Which means you can feed it n vectors and m covectors to get a scalar, or n vectors to get a m-vector space etc.

  • Index placement rule: components from V go upstairs, components from V* go downstairs.

A (m,n) tensor's components are written as a row of row of row... m times, a column of column of column ... n times (the rows and columns might be mixed)

Two players run through everything:

  • F — the forward matrix. Writes the new basis vectors as combinations of the old: ẽ_j = F^i_j e_i.
  • B — the backward matrix. Writes the old basis vectors as combinations of the new: e_i = B^j_i ẽ_j.
  • F and B are inverses: F B = B F = δ (the Kronecker delta / identity, δ^i_j = 1 if i=j, else 0).
  • A (m,n) tensor has m upper indices and n lower indices and transforms with m B's and n F's.

1. Contravariant vs covariant — the central distinction

When you change basis, things split into two camps depending on which direction they transform:

  • Contravariant = transforms opposite to the basis vectors → written with an UPPER index → uses B to go old→new.
  • Covariant = transforms the same way as the basis vectors → written with a LOWER index → uses F to go old→new.

Intuition: if you double the length of your basis vectors, the same fixed vector now needs half as many of them → its components shrink (contravariant). If you double the basis, a covector's components double (covariant). Things that shrink when the ruler grows are contravariant; things that grow with the ruler are covariant.

Mnemonic: Upper index ⇒ uses B (backward i.e. backward of vector basis) ⇒ contravariant. Lower index ⇒ uses F (forward) ⇒ covariant. Going new→old just swaps F↔B.


2. Vectors

A vector is a member of a vector space: a set of objects you can add and scale by scalars. The vector is invariant; only its component list [v^1, v^2, …] depends on the basis.

  • A vector is built from basis vectors: v = v^i e_i (Einstein sum).
  • Vector components are contravariant (upper index, transform with B).

Einstein summation convention (used everywhere): when an index appears once up and once down in a term, sum over it — drop the Σ. Real sums always pair one upper with one lower. The Kronecker delta cancels/renames a summed index: M^i_j δ^j_k = M^i_k.


3. Covectors

A covector is a linear function that eats a vector and returns a scalar: α(v) ∈ ℝ, with α(v) = α_i v^i. Linearity: α(v+w) = α(v)+α(w) and α(nv) = n·α(v). Covectors live in their own vector space, the dual space V* (a row "vector" is just a covector's component list).

Picture a covector as a stack of lines: Treat α = (2,1) as the function f(x,y) = 2x + 1y. Draw its level sets (f = 0, 1, 2, …): a family of equally spaced parallel lines, like a contour map, with an orientation arrow pointing the way the value increases.

  • Measurement = counting pierced lines. α(v) = the number of stack-lines the arrow v crosses. Pierces 2 → value 2. Reaches halfway to the first line → 0.5. Runs parallel to the lines → 0. Points against the orientation → negative.
  • Denser lines = "bigger" covector. To compute 2α, make the lines twice as dense (NOT spread out): v then crosses twice as many → (2α)(v) = 2·α(v).
  • Adding covectors: lay both stacks down; a vector pierces the combined lines, so (α+β)(v) = α(v) + β(v).

Standard vectors (contravariant vectors) are written as columns, while covectors are written as rows.

Dual basis ε^i: defined by ε^i(e_j) = δ^i_j. The ε^i "read off" components: ε^i(v) = v^i. Any covector expands as α = α_i ε^i, where α_i = α(e_i). Basis covectors are contravariant (use B); covector components are covariant (use F) — opposite directions, which is the whole point of upper-vs-lower bookkeeping.


5. Linear maps — a (1,1)-tensor

A linear map eats a vector and returns a vector, obeying L(v+w)=L(v)+L(w), L(av)=aL(v). Its component form is a matrix.

  • Geometrically (in words): a transformation that keeps grid lines parallel and evenly spaced and fixes the origin — stretches, rotations, shears. Translations are NOT linear (they move the origin).

  • Reading the matrix: column j tells you where basis vector e_j is sent. L(e_j) = L^i_j e_i.

  • A linear map transforms vectors, NOT the basis — input and output are measured in the same basis.

  • Action: (Lv)^i = L^i_j v^j (ordinary matrix × column).

  • Transformation rule: L̃ = B L F, i.e. L̃^a_b = B^a_i L^i_j F^j_b. One B (upper index) + one F (lower index) ⇒ a (1,1)-tensor.

    Why two matrices, in words: to get a new-basis output from new-basis input, detour: F (new→old input), then L (old in→old out), then B (old→new output).


6. The metric tensor — how you measure length and angle

Problem: Pythagoras |v|² = (v^1)² + (v^2)² only works in an orthonormal basis (unit-length, perpendicular basis vectors). In a skewed/scaled basis it gives the wrong, non-invariant length.

Fix: length comes from the dot product, |v|² = v·v. Expanding v = v^i e_i:

|v|² = v^i v^j g_ij, where g_ij = e_i · e_j = the metric tensor.

  • It is the matrix of dot products of the basis vectors; as a sandwich, |v|² = vᵀ g v.
  • Orthonormal basis ⇒ g_ij = δ_ij (identity) ⇒ Pythagoras, which is just the special case.
  • Symmetric: g_ij = g_ji (dot product ignores order).
  • Angles too: v·w = |v||w|cos θ, so cos θ = (v^i w^j g_ij) / (|v||w|). (For unit basis vectors at angle θ, ẽ_1·ẽ_2 = cos θ.)
  • It's a (0,2)-tensor: both indices lower, transforms with two F's: g̃_ij = F^k_i F^l_j g_kl. When the basis grows, the g_ij shrink to match, so |v|² stays invariant in every coordinate system.

7. Bilinear forms — the (0,2) family the metric belongs to

A bilinear form B(v,w) eats two vectors, returns a scalar, and is linear in each slot separately:

  • Scale one slot: a·B(v,w) = B(av,w) = B(v,aw). (Scaling both gives a² — watch out.)
  • Distribute each slot: B(v₁+v₂, w) = B(v₁,w)+B(v₂,w), same for the 2nd slot. Both slots summed ⇒ four terms.
  • Components B_ij = B(e_i, e_j); output B(v,w) = B_ij v^i w^j (row × matrix × column).
  • (0,2)-tensor, fully covariant: B̃_kl = F^i_k F^j_l B_ij.

The metric is a bilinear form with two extra properties: (1) symmetric g_ij = g_ji, and (2) positive-definite g(v,v) ≥ 0 (lengths can't be negative). So: metric tensors ⊂ bilinear forms (a non-symmetric matrix, or a symmetric one that outputs a negative for some v, is a valid bilinear form but not a metric).


8. The big reframe — tensors are built from vectors and covectors

Stop defining tensors by their transformation rules. Build them from vectors and covectors using the tensor product ⊗ — then every rule falls out for free.

  • Linear map = vector ⊗ covector: L = L^i_j (e_i ⊗ ε^j). (A single pair e_i ⊗ ε^j is a "pure" rank-1 map sending everything to one direction; interesting maps are sums of pure pairs.)
  • Bilinear form = covector ⊗ covector: B = B_ij (ε^i ⊗ ε^j). (Two covectors because it eats two vectors — each covector consumes one.)

Why this is powerful — three freebies from transforming each factor on its own (vectors with B, covectors with F, then pull the matrices out front):

  1. Transformation rule appears automatically (one B + one F for a linear map; two F's for a bilinear form).
  2. The multiplication formula appears automatically — via ε^j(e_k) = δ^j_k collapsing indices: e.g. L(v) = L^i_j v^j e_i, B(v,w) = B_ij v^i w^j.
  3. The array shape appears automatically (linear map = "row of columns"; bilinear form = "row of rows," which lets you honestly write both input vectors as columns).

9. The (p, q) classification — the general rule

  • ⊗ has three meanings (same idea, different level): Kronecker product (combine arrays — distribute the left array into each entry of the right), tensor product (combine tensors), tensor product of spaces (combine V and V\* into a bigger space like V⊗V\*). The components from the tensor product are exactly the entries from the Kronecker product.
  • For big tensors, use index notation, not "Q(D)". Once there are several slots, "Q acting on D" is ambiguous — there are many valid contractions. Einstein notation states exactly which indices are summed, so it's the unambiguous language.

Deepest definition: a tensor is a multilinear map — feed it vectors/covectors and it's linear in each slot separately. A (1,1)-tensor, depending on what you feed it, can act as a map V→V, a map V*→V*, or a scalar-valued function — same object, different contractions. Any contraction is allowed as long as each summed pair is one-upper-with-one-lower. The number of leftover upper/lower free indices tells you what space the output lives in.


10. Raising and lowering indices — the metric as a bridge

Question: is there a basis-independent partner covector for each vector? The naive pairing e_i ↦ ε^i fails: basis vectors grow with F, basis covectors shrink with B, so the pairing falls apart under a change of basis.

Correct partner: pair v with the covector v·( ) — "dot v with whatever comes in." This is genuinely linear, references no basis, and scales with v. Its components are gotten by feeding v into one slot of the metric:

  • Lower (vector → covector), the flat ♭: v_j = g_ij v^i
  • Raise (covector → vector), the sharp ♯: v^k = g^{kj} v_j

where g^{ij} is the inverse metric, defined by g^{ik} g_kj = δ^i_j (a (2,0)-tensor).

  • Mnemonic: musical flat ♭ lowers (pointy arrow → flat stack of lines); sharp ♯ raises (flat stack → pointy arrow). They're inverse operations.
  • ⚠ v^i ≠ v_i. Upstairs and downstairs components are different objects; you must contract with g (or g⁻¹) to convert. They coincide only in an orthonormal basis where g_ij = δ_ij.
  • This works on any index of any tensor: contract a slot with g to lower it (turns a V factor into a V\* factor) or with g⁻¹ to raise it. The metric stitches all the (p,q) spaces together into one family.

Vector Calculus

  1. Type signatures: grad ⟶ Scalar to Vector , div ⟶ V to S , curl ⟶ V to V , Laplacian ⟶ S to S
  2. The Gradient points in the direction of fastest increase; Divergence measures net expansion or compression (sources and sinks); and Curl measures local swirling or rotation.
  3. $\text{curl}(\text{grad } f) = 0$ and $\text{div}(\text{curl } \mathbf{F}) = 0$. Consequently, all gradient fields are perfectly irrotational (no swirl), and all curl fields are perfectly incompressible (no divergence).
  4. Stokes' Theorem dictates that a gradient field produces exactly zero circulation around any closed loop $C$ bounding a surface $S$:$$\oint_C \nabla f \cdot d\mathbf{r} = \iint_S \nabla \times (\nabla f) \cdot \mathbf{n} , dS = 0$$
  5. Divergence Theorem dictates that a curl field produces exactly zero net flux through any closed surface $S$ enclosing a volume $V$: $$\iint_S (\nabla \times \mathbf{F}) \cdot \mathbf{n} , dS = \iiint_V \nabla \cdot (\nabla \times \mathbf{F}) , dV = 0$$
  6. You derive fundamental physics PDEs by counting a quantity, tracking its boundary flux, and applying Gauss's theorem to drop the integrals. This yields conservation laws like the continuity equation ($\frac{\partial\rho}{\partial t} + \nabla \cdot (\rho \mathbf{u}) = 0$). Under ideal conditions (steady, incompressible, irrotational), this simplifies entirely to Laplace's equation ($\nabla^2\varphi = 0$).

A scalar field f(x, t) assigns a number to every point (e.g. temperature). A vector field u(x, t) assigns a vector to every point in space (e.g. fluid velocity in the Gulf of Mexico). The field is the solution of a PDE; dropping a particle into the field gives an induced ODE dx/dt = u(x, t) (e.g. tracking an oil spill). Taking the gradient of a scalar field gives a vector field. Taking the divergence of a vector field gives a scalar field. Taking the curl of a vector field gives a vector field.

1. The operator: ∇ (del / nabla)

∇ is a vector of partial derivatives — treat it like a real vector:

∇ = ( ∂/∂x , ∂/∂y , ∂/∂z )

Three products with it give the three core operations, plus a fourth built from them. Memorize the type signatures — they are the whole point:

Operation Written Input → Output Formula
Gradient grad f = ∇f scalar → vector (∂f/∂x, ∂f/∂y, ∂f/∂z)
Divergence div F = ∇·F vector → scalar ∂Fx/∂x + ∂Fy/∂y + ∂Fz/∂z
Curl curl F = ∇×F vector → vector <∂F₃/∂y − ∂F₂/∂z , ∂F₁/∂z − ∂F₃/∂x , ∂F₂/∂x − ∂F₁/∂y> (take the determinant)
Laplacian ∇²f = ∇·(∇f) scalar → scalar ∂²f/∂x² + ∂²f/∂y² + ∂²f/∂z²

All four are linear operators i.e. ∇(af₁ + bf₂) = a∇f₁ + b∇f₂.


2. Gradient — direction of fastest increase

∇f at each points is oriented in the direction f (the scalar field) increases fastest; its magnitude is the rate of increase.

  • Example: f = x² + y² → ∇f = (2x, 2y). Arrows point radially out, growing with radius — matches the paraboloid getting steeper outward.
  • Directional derivative (rate of change of f along a chosen direction v):
    D_v f = (v / |v|) · ∇f
    
    Normalize v. If v ⊥ ∇f, the directional derivative is 0 (move that way and f doesn't change).

3. Divergence — sources & sinks (net flux per unit volume)

∇·F > 0: field is sourcing/spreading out (faucet hitting the sink basin). ∇·F < 0: field is sinking/converging (the drain). ∇·F = 0: divergence-free / incompressible — volumes are neither created nor destroyed.

Intuition: drop a little blob of dye into the flow; positive divergence makes the blob grow, negative makes it shrink, zero keeps its area fixed (it might deform, but doesn't expand or contract).

  • F = (x, y) → ∇·F = 1 + 1 = 2 (diverging).
  • F = (−x, −y) → ∇·F = −2 (converging).
  • F = (−y, x) → ∇·F = 0 (pure rotation, divergence-free).

4. Curl — circulation / rotation

curl F measures how much the field swirls around a point. In 3D:

∇×F = ( ∂F₃/∂y − ∂F₂/∂z ,  ∂F₁/∂z − ∂F₃/∂x ,  ∂F₂/∂x − ∂F₁/∂y )

In 2D (F = (F₁, F₂, 0)) only the z-component survives — curl is effectively a scalar pointing out of the page:

(∇×F)_z = ∂F₂/∂x − ∂F₁/∂y
  • F = (x, y) → curl = 0 → irrotational (radial, no swirl).
  • F = (−y, x) → curl = 2 → rotational (this is the divergence-free swirl from §3).
  • Sign by right-hand rule: thumb out of the page = positive curl.

Solid-body rotation: a body spinning with angular-velocity vector ω has velocity field v = ω × r. Then curl v = 2ω — curl is twice the angular velocity. A constant curl ⇒ the blob rotates without deforming.


5. Two identities you must know cold

curl(grad f) = 0      for every scalar f
div(curl F)  = 0      for every vector field F

Consequences used constantly later:

  • Any gradient field (F = ∇f) is automatically curl-free / irrotational. → So a field with any curl cannot be the gradient of a scalar potential.
  • Any curl field (F = ∇×G) is automatically divergence-free. Taking the curl of a flow extracts only its rotational part.

6. Are all vector fields gradients of a scalar? — No.

Scalars need 1 function; vector fields need 3 (F₁, F₂, F₃) — a strictly bigger space. So a generic field is not ∇ of anything.

  • Conservative / gradient field: F = ∇f for some scalar f. These are exactly the curl-free (irrotational) fields (by §5). ∂F₃/∂y − ∂F₂/∂z = 0 (for a 2D plane)
  • Potential flow is the even more special case: a gradient field that is also incompressible → it satisfies Laplace's equation (§9). So: potential flow ⊂ gradient flows.

7. The integral theorems — local ↔ global

These convert between what happens inside a region and what happens on its boundary. They are the bridge from "rate of change in a blob" to a PDE.

Gauss's Divergence Theorem

Flux out through a closed surface = integral of divergence over the enclosed volume.

∯_S  F · n  dS   =   ∭_V  (∇·F)  dV

F·n = how much of the field pierces the surface normal-wise (tangential flow carries nothing out). Why it's true: chop the volume into tiny boxes; for continuous F, every internal wall's outflow cancels its neighbor's inflow, leaving only the outer perimeter = the surface flux. Requires F continuous — fails at shock waves / discontinuities.

Stokes's Theorem & Green's Theorem (circulation)

Stokes: circulation around a closed loop = flux of curl through any surface it bounds:

∮_C  F · dr   =   ∬_S  (∇×F) · n  dS

Green is the 2D special case:

∮_C (F₁ dx + F₂ dy)  =  ∬_R ( ∂F₂/∂x − ∂F₁/∂y ) dA

Pairing to remember: Gauss ↔ divergence ↔ flux through surface; Stokes ↔ curl ↔ circulation around loop.


8. From a conservation law to a PDE (the master recipe)

This single pattern derives the continuity equation, Navier–Stokes, the heat equation, etc.

  1. Count the stuff. Mass in a volume V: M = ∭_V ρ dV (ρ = density).
  2. Accounting. Its rate of change = what flows out through the boundary (+ any source Q):
    d/dt ∭_V ρ dV  =  − ∯_S (ρ u) · n dS  ( + ∭_V Q dV )
    
    (Minus sign: positive outward flux means mass decreasing inside.)
  3. Gauss's theorem turns the surface integral into a volume integral → everything under one ∭_V.
  4. "True for all volumes." If the integral is zero for every volume, the integrand must be zero everywhere. Drop the integral → the PDE.

Result — the Mass Continuity Equation:

∂ρ/∂t  +  ∇·(ρ u)  =  0
  • Incompressible flow (ρ = constant ⇒ ∂ρ/∂t = 0, ∇ρ = 0) collapses this to:
    ∇·u = 0          (velocity field must be divergence-free)
    
  • Source terms Q (nuclear blast converting mass→energy, etc.) put a nonzero right-hand side → a forced PDE.
  • Caveat: Step 4 needs ρ and u smooth. Shock waves / discontinuities break it — derivatives blow up — and you must stay in the integral "control-volume" form.
  • Same recipe on momentum → Navier–Stokes; on energy → the energy/heat equation.

9. Laplace's & Poisson's Equations + Potential Flow

Potential flow = steady + incompressible + irrotational flow. Build it from a scalar potential φ via v = ∇φ:

  • Irrotational is automatic (curl ∇φ = 0, §5).
  • Steady is automatic if φ doesn't depend on time.
  • Incompressible is the one real condition: ∇·v = ∇·(∇φ) = 0, i.e.
∇²φ = 0          ← LAPLACE'S EQUATION

Solutions are called harmonic functions. Laplace's equation is everywhere:

  • Steady-state heat: the heat equation ∂T/∂t = α²∇²T with ∂T/∂t = 0 becomes ∇²T = 0.
  • Gravitational & electrostatic potentials away from masses/charges.
  • Potential flows (early airfoil design).

Poisson's equation = Laplace with a source (force it and it becomes Poisson):

∇²φ = f          (e.g. ∇²φ = −ρ/ε in electrostatics)

f is a source/forcing term (heat source, charge density). Solving Poisson is a core, expensive step inside Navier–Stokes solvers (pressure).

Boundary conditions are required for a unique solution: Dirichlet (value fixed on boundary, e.g. edge temperatures) or Neumann (normal-derivative fixed, e.g. insulation/flux).

The complex-analysis shortcut (2D)

Any analytic complex function F(z) = φ(x,y) + i ψ(x,y), z = x+iy (e.g. z², eᶻ, log z, sin z) hands you a potential flow for free:

  • Real part φ = velocity potential; imaginary part ψ = stream function; both are harmonic (∇²φ = ∇²ψ = 0).
  • Cauchy–Riemann equations: φ_x = ψ_y, φ_y = −ψ_x.
  • Velocity two ways: v = (φ_x, φ_y) = (ψ_y, −ψ_x).
  • ψ = const traces streamlines; equipotential lines (φ = const) are orthogonal to streamlines.
  • Worked example z² = (x²−y²) + i(2xy): φ = x²−y², ψ = 2xy, v = (2x, −2y). Check: ∇·v = 2−2 = 0 ✓, curl = 0 ✓. A saddle: stretch in x, squeeze in y, area preserved.
  • Superposition of elementary analytic flows (uniform stream + sources/sinks + vortices) builds flow over airfoils — the pre-computer design method.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment